Raghav Mittal
Menu
Approach →
Services
Automation Replace manual follow-ups, spreadsheet relays, and status chasing with trigger-based systems that move work automatically. CRM Systems Turn your CRM from a contact database into an operating system for pipeline, ownership, follow-up, and reporting. SEO & UX Find the crawl, speed, mobile, UX, and conversion issues that stop good pages from ranking or turning traffic into leads. Shopify & Web Build or clean up websites that are fast, trackable, SEO-ready, and designed around how buyers actually decide. Shopify Automation Connect Shopify orders, customers, inventory, fulfilment, support, subscriptions, and reporting so your team spends less time moving the same information. AI Workflows Use AI practically inside business workflows: Excel analysis, PPT reporting, SOPs, dashboards, content, and team documentation. Lead Systems Capture, qualify, assign, and follow up with leads across forms, CRM, WhatsApp, email, and sales teams. GST Automation Reduce repetitive GST data work by connecting invoices, sales channels, approvals, reconciliation checks, and reporting into a workflow your team can actually operate. Tally Automation Make Tally and TallyPrime part of a reliable operating workflow for invoices, ledgers, receivables, reports, approvals, and data handoffs. GST + Tally Connect the work between sales, invoices, GST review, Tally, reconciliation, and management reporting so finance operations have fewer blind spots. View all servicesBrowse the complete delivery bench →
Blueprints → Work → Blog → Free Audit →
AI Agents· Sep 28, 2026· 5 min read 3 recorded views

AI Agent Evaluation Checklist: Accuracy, Escalation and Cost per Run

By Raghav Mittal · Consultant & Solutions Architect

Test AI agents against task accuracy, tool use, human escalation, recovery and cost before expanding live permissions.

AI Agent Evaluation Checklist: Accuracy, Escalation and Cost per Run

An AI agent evaluation should measure whether the system completes the intended task safely, not whether its answer sounds confident. Test the agent with representative cases, known expected outcomes, tool failures and situations where it should hand off to a person. Then measure the cost of a completed, acceptable task. This makes a launch decision possible instead of relying on a compelling demo.

Key takeaways

  • Define the task and acceptable outcome before collecting test cases.
  • Score factual accuracy, tool actions, escalation and recovery separately.
  • Include uncertain and adversarial inputs, not only clean examples.
  • Compare quality and cost per accepted task after each meaningful system change.

Write the job description as a contract

“Help sales” is too broad to evaluate. “Read a new lead, identify the requested service, set the CRM owner and draft a first reply without making a pricing promise” is testable. List the allowed systems, fields and actions. Define what the agent must never do, such as deleting a record, sending a final quote or changing an order after dispatch. Specify when a human must approve a draft. The evaluation should reflect this contract, not a generic benchmark score.

Choose cases from the real distribution of work, with sensitive information removed or appropriately protected. Include straightforward examples, missing data, conflicting information, duplicates and unusual requests. Record an expected outcome for each. For some cases the correct outcome is “escalate,” not a completed automated action. A system that guesses an answer where it should ask for help may appear productive while increasing risk.

Score more than the final text

OpenAI's agent evaluation guidance describes assessing agent behaviour, including traces. This matters because two identical-looking replies can come from very different tool paths. One agent may check the current CRM record; another may fabricate a status. Evaluate the sequence of tool calls, permissions used and changes written, not only the final message. Keep a record of the source facts the agent had at the time.

Dimension Pass condition Example failure
Task accuracy Outcome matches the case facts Wrong customer or order selected
Tool use Correct system and allowed action used Writes to an unrelated record
Grounding Claims trace to available source Invents a policy or price
Escalation Hands off when boundary is reached Guesses during ambiguity
Recovery Fails safely and reports the problem Retries a write and duplicates it
Cost Accepted task stays within budget Long retry chain for one result

Score these separately. An agent can have excellent language quality and poor tool safety. Conversely, a cautious agent may escalate more often but still create value by preparing a clear case summary. Decide which trade-off the business can tolerate before looking at results. Avoid combining everything into one opaque percentage that hides a serious failure mode.

Build an evaluation set that catches reality

Start with a small set of high-frequency cases and add cases from actual failures. Include at least one example for each policy boundary and each integration. If the agent uses search or retrieval, include stale and contradictory documents. If it sends messages, include a customer who opts out or asks for something outside scope. If it changes records, include duplicate records and a timed-out write. Re-run the same set after updating the model, prompt, tools or knowledge source.

Also review the evaluator itself. A perfect score from cases that only test the happy path is not useful. Have a domain owner inspect a sample of passes and failures. Keep expected answers concise and observable: “assign owner A; do not send a message; flag missing consent” is better than “respond professionally.” OpenAI's agent evaluation guidance can inform repeatable testing, but the business still has to define what correct work looks like.

Test handoff and recovery deliberately

When the agent is uncertain, where does the case go? A handoff should carry the input, attempted steps, source facts, uncertainty and proposed next action. The human should not have to reconstruct the whole journey. Test both a correct escalation and a false escalation. Too many avoidable handoffs can erase the economics; too few can put the wrong decision into production. Measure the rate and reason, then improve the workflow rather than hiding the queue.

For recovery, interrupt a tool call or return an invalid API response. The agent should not quietly invent success. If an update may have been written before a timeout, a blind retry could duplicate it. The workflow needs a way to check state or use an idempotent operation. Log failures in a place an owner actually reviews. Our AI agent consultant guide explains why operational boundaries matter more than a flashy demonstration.

Measure cost per accepted task

Model calls, retrieval, external tools, retries and human review all contribute to cost. Divide the total by tasks that met the acceptance standard, not by calls attempted. Record median and high-end cost because a small number of difficult cases may dominate spend. Compare the result with the current human process, including rework. Do not assume the agent saves time merely because it creates a draft quickly; measure review and correction time as well.

Provider rates can change, so check live API pricing when updating the model. For a broader budgeting method, see the AI automation cost guide. A system can be accurate but too expensive for a low-value task; another can be inexpensive but need so much review that it is not operationally useful. Both are valid reasons to redesign the scope.

Make the launch decision explicit

Before live write access, agree on minimum pass criteria for high-risk cases, permitted escalation rate, owner response time, acceptable cost and a rollback trigger. Start with limited permissions and a small cohort. Monitor real cases for new failure patterns and add them to the evaluation set. A passing prelaunch test is a starting point, not a permanent certificate.

If your team needs to turn a prototype into a measurable operating workflow, see AI workflow automation. You can contact Raghav with a task definition and anonymised examples for an evaluation design that covers accuracy, escalation and cost.

AI Workflow Automation

I design workflows that use Claude, DeepSeek, OpenAI, spreadsheets, slides, docs, and business data to save real team hours.

Explore the service →
Turn the idea into a working system

Have a bottleneck that needs an accountable owner?

Send me the problem, where it is getting stuck, and what a useful outcome looks like. I will reply with the clearest next step.

Prefer a conversation? Book a call →
Keep reading
View the full archive →