Prompt Evaluation

[Preview] -- The prompt evaluation framework is in active development.

Prompt evaluation provides a structured way to test and measure agent quality. Define test cases with expected outcomes, run them against your agents, and track quality over time.

Prerequisites: Agent Declarations, Tool Declarations. What you'll learn: How to evaluate agent behavior with structured test cases.

Why Evaluate Prompts?

Agents are non-deterministic -- the same input can produce different outputs. Prompt evaluation addresses this by:

  1. Defining test scenarios with expected behaviors
  2. Running scenarios repeatedly to measure consistency
  3. Tracking quality metrics across prompt iterations
  4. Catching regressions when instructions or tools change

Evaluation Structure

eval customer_support_quality {
  agent customer_support

  scenario "order_lookup" {
    input "Where is my order #12345?"
    expect tool_called "lookup_order"
    expect output contains "order"
    expect no_tool_called "process_return"
  }

  scenario "return_request" {
    input "I want to return the shoes I bought last week"
    expect tool_called "lookup_order"
    expect tool_called "process_return"
    expect output contains "return"
  }

  scenario "out_of_scope" {
    input "What's the weather today?"
    expect no_tool_called
    expect output contains "can't help"
  }
}

Assertion Types

AssertionDescription
expect tool_called "name"The agent must invoke this tool
expect no_tool_called "name"The agent must NOT invoke this tool
expect no_tool_calledThe agent must not call any tool
expect output contains "text"The response must contain this text
expect output matches "regex"The response must match this pattern
expect errorThe scenario should produce an error

Running Evaluations

Evaluations run through grove-cli:

manzano eval <module-dir>

Results report pass/fail per scenario with details on which assertions failed.

Best Practices

  1. Cover edge cases -- Test out-of-scope requests, ambiguous inputs, and error conditions.
  2. Test tool selection -- Verify the agent picks the right tool for each scenario.
  3. Avoid brittle assertions -- Use contains over exact match for natural language output.
  4. Version your evaluations -- Track eval results over time to catch regressions.
  5. Test sub-agent routing -- For orchestrator patterns, verify correct delegation.

See Also