Prompt Evaluation
[Preview] -- The prompt evaluation framework is in active development.
Prompt evaluation provides a structured way to test and measure agent quality. Define test cases with expected outcomes, run them against your agents, and track quality over time.
Prerequisites: Agent Declarations, Tool Declarations. What you'll learn: How to evaluate agent behavior with structured test cases.
Why Evaluate Prompts?
Agents are non-deterministic -- the same input can produce different outputs. Prompt evaluation addresses this by:
- Defining test scenarios with expected behaviors
- Running scenarios repeatedly to measure consistency
- Tracking quality metrics across prompt iterations
- Catching regressions when instructions or tools change
Evaluation Structure
eval customer_support_quality {
agent customer_support
scenario "order_lookup" {
input "Where is my order #12345?"
expect tool_called "lookup_order"
expect output contains "order"
expect no_tool_called "process_return"
}
scenario "return_request" {
input "I want to return the shoes I bought last week"
expect tool_called "lookup_order"
expect tool_called "process_return"
expect output contains "return"
}
scenario "out_of_scope" {
input "What's the weather today?"
expect no_tool_called
expect output contains "can't help"
}
}
Assertion Types
| Assertion | Description |
|---|---|
expect tool_called "name" | The agent must invoke this tool |
expect no_tool_called "name" | The agent must NOT invoke this tool |
expect no_tool_called | The agent must not call any tool |
expect output contains "text" | The response must contain this text |
expect output matches "regex" | The response must match this pattern |
expect error | The scenario should produce an error |
Running Evaluations
Evaluations run through grove-cli:
manzano eval <module-dir>
Results report pass/fail per scenario with details on which assertions failed.
Best Practices
- Cover edge cases -- Test out-of-scope requests, ambiguous inputs, and error conditions.
- Test tool selection -- Verify the agent picks the right tool for each scenario.
- Avoid brittle assertions -- Use
containsover exact match for natural language output. - Version your evaluations -- Track eval results over time to catch regressions.
- Test sub-agent routing -- For orchestrator patterns, verify correct delegation.
See Also
- Agent Declarations -- Define agents to evaluate
- Agentic Patterns -- Patterns to test against
- CLI Reference -- The
evalcommand