An AI agent evaluation should measure whether the system completes the intended task safely, not whether its answer sounds confident. Test the agent with representative cases, known expected outcomes, tool failures and situations where it should hand off to a person. Then measure the cost of a completed, acceptable task. This makes a launch decision possible instead of relying on a compelling demo.
Key takeaways
- Define the task and acceptable outcome before collecting test cases.
- Score factual accuracy, tool actions, escalation and recovery separately.
- Include uncertain and adversarial inputs, not only clean examples.
- Compare quality and cost per accepted task after each meaningful system change.
Write the job description as a contract
“Help sales” is too broad to evaluate. “Read a new lead, identify the requested service, set the CRM owner and draft a first reply without making a pricing promise” is testable. List the allowed systems, fields and actions. Define what the agent must never do, such as deleting a record, sending a final quote or changing an order after dispatch. Specify when a human must approve a draft. The evaluation should reflect this contract, not a generic benchmark score.
Choose cases from the real distribution of work, with sensitive information removed or appropriately protected. Include straightforward examples, missing data, conflicting information, duplicates and unusual requests. Record an expected outcome for each. For some cases the correct outcome is “escalate,” not a completed automated action. A system that guesses an answer where it should ask for help may appear productive while increasing risk.
Score more than the final text
OpenAI's agent evaluation guidance describes assessing agent behaviour, including traces. This matters because two identical-looking replies can come from very different tool paths. One agent may check the current CRM record; another may fabricate a status. Evaluate the sequence of tool calls, permissions used and changes written, not only the final message. Keep a record of the source facts the agent had at the time.
| Dimension | Pass condition | Example failure |
|---|---|---|
| Task accuracy | Outcome matches the case facts | Wrong customer or order selected |
| Tool use | Correct system and allowed action used | Writes to an unrelated record |
| Grounding | Claims trace to available source | Invents a policy or price |
| Escalation | Hands off when boundary is reached | Guesses during ambiguity |
| Recovery | Fails safely and reports the problem | Retries a write and duplicates it |
| Cost | Accepted task stays within budget | Long retry chain for one result |
Score these separately. An agent can have excellent language quality and poor tool safety. Conversely, a cautious agent may escalate more often but still create value by preparing a clear case summary. Decide which trade-off the business can tolerate before looking at results. Avoid combining everything into one opaque percentage that hides a serious failure mode.
Build an evaluation set that catches reality
Start with a small set of high-frequency cases and add cases from actual failures. Include at least one example for each policy boundary and each integration. If the agent uses search or retrieval, include stale and contradictory documents. If it sends messages, include a customer who opts out or asks for something outside scope. If it changes records, include duplicate records and a timed-out write. Re-run the same set after updating the model, prompt, tools or knowledge source.
Also review the evaluator itself. A perfect score from cases that only test the happy path is not useful. Have a domain owner inspect a sample of passes and failures. Keep expected answers concise and observable: “assign owner A; do not send a message; flag missing consent” is better than “respond professionally.” OpenAI's agent evaluation guidance can inform repeatable testing, but the business still has to define what correct work looks like.
Test handoff and recovery deliberately
When the agent is uncertain, where does the case go? A handoff should carry the input, attempted steps, source facts, uncertainty and proposed next action. The human should not have to reconstruct the whole journey. Test both a correct escalation and a false escalation. Too many avoidable handoffs can erase the economics; too few can put the wrong decision into production. Measure the rate and reason, then improve the workflow rather than hiding the queue.
For recovery, interrupt a tool call or return an invalid API response. The agent should not quietly invent success. If an update may have been written before a timeout, a blind retry could duplicate it. The workflow needs a way to check state or use an idempotent operation. Log failures in a place an owner actually reviews. Our AI agent consultant guide explains why operational boundaries matter more than a flashy demonstration.
Measure cost per accepted task
Model calls, retrieval, external tools, retries and human review all contribute to cost. Divide the total by tasks that met the acceptance standard, not by calls attempted. Record median and high-end cost because a small number of difficult cases may dominate spend. Compare the result with the current human process, including rework. Do not assume the agent saves time merely because it creates a draft quickly; measure review and correction time as well.
Provider rates can change, so check live API pricing when updating the model. For a broader budgeting method, see the AI automation cost guide. A system can be accurate but too expensive for a low-value task; another can be inexpensive but need so much review that it is not operationally useful. Both are valid reasons to redesign the scope.
Make the launch decision explicit
Before live write access, agree on minimum pass criteria for high-risk cases, permitted escalation rate, owner response time, acceptable cost and a rollback trigger. Start with limited permissions and a small cohort. Monitor real cases for new failure patterns and add them to the evaluation set. A passing prelaunch test is a starting point, not a permanent certificate.
If your team needs to turn a prototype into a measurable operating workflow, see AI workflow automation. You can contact Raghav with a task definition and anonymised examples for an evaluation design that covers accuracy, escalation and cost.
AI Workflow Automation
I design workflows that use Claude, DeepSeek, OpenAI, spreadsheets, slides, docs, and business data to save real team hours.