
An agent that demos well can still fail in quiet ways once real users reach it: calling a tool wrong, losing track mid-task, or returning an answer that reads fine but misses the point. Agent evaluation is how a team catches those failures on purpose, testing the steps an agent takes and the results it returns against clear criteria. This article covers what agent evaluation measures, the metrics and methods teams use, and how to build a check that runs before a team puts an agent into production.
The short answer
Agent evaluation tests an agent against clear criteria before it reaches production, checking both the steps it takes and the results it returns. Whether an agent holds up is something a team measures on cases it assembles, since one that reads well in a demo can still act wrong once real use reaches it. A team sets the criteria from its own work, so what counts as a passing result on one task differs from another, and this article is a way to run that check by capability, not a fixed list of methods to copy.
What a team gets from a check, then, is evidence for this round, and not a promise about every later run. The sections below set out what agent evaluation measures, the metrics and methods a team uses, the criteria to set up front, and where the check fits against building and running an agent, so a team can test an agent on its own work before it goes live.
What agent evaluation measures
Agent evaluation looks at two layers of the same run: the steps an agent takes and the results it returns. The step layer covers how the agent moved through the task, such as which tools it called and whether it stayed on the task or drifted. The result layer covers what the agent returned at the end, read against the criteria a team set for the task.
The two layers answer different questions, and a team weighs both against its own work. A result can read as correct while the steps behind it went wrong in a way that shows up later, and steps can look sound while the result misses what the task needed. What each layer counts as right is set by the team for the task at hand, and it is not fixed in advance for every agent.
Metrics and methods to test an agent before production
The metrics and methods below are angles a team can check an agent on, each against the work it means to run. Two of them cover the layers named above, the steps and the results, and the third covers how a team runs the check itself. None is a fixed authority; a team picks the angles that fit its task and sets what each one counts as passing.
Whether the agent's steps and tool use hold up
One angle is the steps an agent takes: which tools it called, in what order, and whether it stayed on the task. Checking the steps shows where a run went wrong even when the final answer reads fine, since a wrong tool call or a lost thread mid-task can sit behind a result that looks right. A team reads the steps against what the task needed, and what counts as a sound step is set for that task, since a tool call that fits one task can be wrong for another.
Whether the results meet the criteria a team set
Another angle is the result the agent returned, read against the criteria a team set for the task. A result can be checked for whether it did what the task asked, kept within cost and time bounds, and stayed inside the limits a team set. The criteria carry the weight here, since a result means something against a standard the team wrote down, and a team sets that standard from its own work and not from a general score.
How a check is run: fixed cases, grading, and member review
A check is run on cases a team assembles, with a way to grade each result and a point for a member to review. Fixed cases give the agent the same inputs each run, so a team can see whether a change helped or hurt. Grading can be by a rule that checks the result against a set answer, by a model that scores it, or by a member who reads it, and each has a place depending on what the task needs. A member review can supplement automated grades where the task calls for it, since a score from a rule or a model gives evidence for this round and does not settle on its own whether the result fits.
Criteria to set before you test
Before a team tests an agent, it sets the criteria the check will read against, starting with what counts as the task done. The criteria cover the result a task should reach, the tool use a team accepts, the cost and time a run should stay within, and the safety limits the work carries. Each of these is set by the team from its own work, so the criteria for one task differ from another, and the check means something against them.
Writing the criteria down up front keeps the check honest, since a criterion added after seeing a result can bend to fit it. A team that sets the standard before the run has a fixed line to read the result against, and a member can confirm the criteria hold for the work before the agent is tested against them.
Where agent evaluation fits against building and running an agent
Agent evaluation sits next to building an agent workflow and next to the run-time loop the agent runs, and it is a different job from either. Naming where it differs keeps the check from turning into a rebuild or a rerun.
Building an agent workflow and making it reliable is the work of putting the agent together and hardening it, and that is covered in the piece on AI agent workflows. Agent evaluation does not rebuild or harden the workflow; it tests and measures what the built agent does against the criteria a team set. The build practices belong with the AI agent workflow piece, and this piece stays on the test.
The run-time loop a single agent runs, where it plans, acts, and checks its own work as it goes, is covered in the piece on agentic workflows. That in-loop check is the agent's own step at run time, and agent evaluation is a separate test a team runs from outside, mostly before production, against criteria set in advance. Run-time monitoring and observability of a live agent stay with the platform that runs it, and this piece stays on the pre-production check.
Running the evaluation before production
Running the evaluation is a set of steps a team takes before an agent goes live: assemble cases from real use, set the criteria, run the agent on the cases, read the results by the grades and a member review, and decide from there. The cases are most useful when they come from the work the agent will meet, so the check sees what real use looks like and not a tidy sample.
For a comparison across changes, a team may run the agent on the same cases each time, so a later change can be read against an earlier run. Reading the results covers both the step layer and the result layer, weighed against the criteria set up front, and a member confirms what the grades show. What the team decides from there, including whether the agent goes live, sits with the team and a member it names, since the check gives evidence and does not make the call.
Where an evaluation can fall short
An evaluation can fall short of what a team reads into it in a few ways, and naming them keeps the check in proportion. Each is a gap to look for, and none shows up from a passing score alone.
The cases may not cover what real use brings, so an agent passes the check and still meets inputs it was not tested on. The grading criteria may miss what matters, so a result scores well and still falls short on the task. And a pass on this round is evidence for this round, and it does not prove the agent will hold across later runs or a broader body of work. A team reads a passing check as a step, and it keeps the go-live decision with a member who weighs the gaps against the risk the work carries.
Keeping the criteria, results, and sign-off visible
- Records stay visible, and review follows the run. Selected work, handoff, and review records stay visible in a shared channel to members and agents at the same time, so a member reviews against the run as it happened, with the record in hand.
- Status and ownership stay clear. On a task, each piece of work carries a status and an owner, so who is working, where it stands, and who picks it up next stay clear to the team.
- Outputs are judged against a standard. The result of a step can be represented as a deliverable, which a member opens and assesses against the standard the team wrote down.
- Key actions get a human gate. An action a team marks in advance as high-risk or outward-facing can be prepared as an Action Card, which an authorized member reviews and submits under their own identity.
- Context carries across sessions. Across conversation turns and separate sessions the relevant context stays continuous, so a handoff does not start from scratch.
An evaluation produces numbers, and numbers outlive the reasoning behind them. Six months later the team has a pass rate and no account of what it was measuring or who accepted the result.
Running a check before production is worth little if the criteria were settled afterwards. The team that wrote the standard, the check that produced the result, and the member who signed off are three separate things. Syfo keeps all three visible together, and it covers the coordination an evaluation exercise brings with it, namely how criteria are recorded, who owns the results, and who signs off on the evaluation.
- The criteria are kept where the results are. The standard a team wrote and the check that produced the result sit in one channel, so the two can be read against each other.
- A result has an owner. The task carries who ran the check and where it stands, which matters when an evaluation is repeated across several rounds.
- Sign-off is recorded as an act. A member's acceptance attaches to the work as a Deliverable, so a later reader sees who took responsibility.
- A consequential action waits for that member. Where a team marked a step in advance as high-risk or outward-facing, it can be prepared for an authorized member to review and submit in their own identity.
The same record carries context across sessions, which is what lets a second round of testing build on the first rather than repeat it. Running and hosting the evaluation itself stays with whichever tool the team uses.
How it relates to nearby terms
Several terms sit close to agent evaluation, and they are easy to blur. For this article:
- AI agent workflow is building an agent workflow and making it reliable, covered in its own article, where the build and the hardening sit.
- Agentic workflow is the run-time loop a single agent runs and the design patterns behind it, covered separately, where the in-loop check sits.
- Run-time monitoring of a live agent, its observability and governance, stays with the platform that runs the agent, and it is a different job from the pre-production check this article covers.
Questions people ask
What does agent evaluation measure? It measures two layers of a run: the steps an agent takes, such as its tool calls and whether it stayed on the task, and the results it returns, read against the criteria a team set. A team weighs both against its own work, since a result can read as correct while the steps went wrong, and the criteria are set for the task at hand.
How do I set criteria for an agent evaluation? Set them before the run, starting with what counts as the task done, and cover the result, the tool use a team accepts, the cost and time bounds, and the safety limits. The criteria are set by the team from its own work, so they differ from one task to another, and writing them down up front keeps a later result from bending them.
What types of evaluation are there? The count depends on how a source splits them, so a fixed number is less useful than the angles a team actually checks. This article groups the check into the steps an agent takes, the results it returns, and how the check is run, and a team can add angles its work needs.
How do I get started with agent evaluation? At a general level, a team runs a small check before production: a few cases from real use, clear criteria set up front, a way to grade the results, and a member to review. Watching how the agent moves through the cases and what it returns shows more than a demo does.
Where to start
Start with one task where the criteria are easy to set: a clear definition of the task done, a few cases from real use, a way to grade the result, and a point where a member reviews. Run the agent on those cases before it goes live, read both the steps and the results against the criteria, and see where it holds and where it falls short. A small check like that shows what the agent does, and it does more than a demo on the page.
Once the check holds and the records support a review, a team has a basis for weighing the agent against a broader body of work. Whether the agent is ready is a call the team and a named member make from the evidence, since a passing check is a step toward that call and does not make it on its own.
Start with the work your team needs to move.
Begin with one real workflow, keep ownership visible, and review the result before expanding the setup.