AI Agent Evaluation: Build a Real Eval Harness | MewCP | MewCP
AI Agent Evaluation: How to Know If Your Agent Actually Works
by Rohit Gite, Founder @MewCP··20 min read
A demo proves an agent worked once. A score proves it works. Here is the harness: case sets, graders, trajectory scoring, pass@k, and the CI gate that keeps it honest.
Day 24 covered tracing: watching what an agent did, one run at a time. This post covers the thing tracing cannot answer on its own, which is whether what it did was any good, and whether it will still be good after the next prompt edit. Those are two different questions. A trace tells you what happened. An eval tells you if it was supposed to happen, graded the same way, on the same cases, every time you change something.
The seven-dimension scorecard from the carousel (task success, tool selection accuracy, output quality, reliability, latency, cost, safety) is the vocabulary. This post is the machine that fills it in: a case set format, two kinds of graders and when each one lies to you, a way to score the path the agent took and not just where it ended up, reliability numbers that survive a stochastic model, percentiles instead of averages, a cost metric that charges failures to the task instead of hiding them in "per call," safety modeled as a gate instead of a score, and a CI job that actually blocks a merge.
None of this is exotic. It is closer to how you'd build a test suite for any nondeterministic system, adapted for the specific ways agents fail that unit tests were never built to catch: an agent can produce the right answer through the wrong tool calls, pass on Monday and fail on Tuesday with no code change, and pass an LLM judge that is quietly making the same mistake the agent made.
The demo is not a case set
A demo is one prompt, run once, read by the person who wrote the prompt. Every one of those three things is a confound.
One prompt tells you nothing about the case you didn't think to try. Run once tells you nothing about the model's actual behavior, because most agents run at a nonzero temperature and a single sample is one draw from a distribution, not the distribution. And read by the person who wrote the prompt means the grader is the same person who has spent an hour building intuition for what "close enough" looks like, which is exactly the intuition a real user does not share.
An eval suite fixes all three at once: a fixed set of cases larger than the ones you happened to think of this morning, run enough times to see the distribution instead of one sample from it, and graded by something that does not get more lenient the longer it stares at the output.
That's the whole shape. The rest of this post is making each piece precise enough to run in CI.
The case set: what a test case actually needs
A case is not just an input and an expected output string. An agent case needs to say which tools should get called, with what arguments, and which of its checks are non-negotiable versus merely scored. Here is the schema:
expected_tools says what should happen. assertions says what gets checked, and hard_gate is the field that matters most in this whole schema: it marks which checks are allowed to fail a run outright regardless of anything else, versus which ones only affect a scored dimension. negate lets a case assert a tool must not be called, which is how a case set expresses "don't do this" instead of only "do that."
Loading it is unglamorous on purpose:
from __future__ import annotationsimport jsonfrom dataclasses import dataclass, fieldfrom pathlib import Pathfrom typing import Any, Literalimport jsonschemaCASE_SCHEMA = json.loads(Path("eval/case.schema.json").read_text())@dataclass(frozen=True)class ExpectedTool: tool_name: str required: bool expected_args: dict[str, str] = field(default_factory=dict)@dataclass(frozen=True)class Assertion: type: str hard_gate: bool negate: bool = False params: dict[str, Any] = field(default_factory=dict)@dataclass(frozen=True)class Case: case_id: str difficulty: Literal["easy", "medium", "hard"] prompt: str context: dict[str, Any] expected_tools: list[ExpectedTool] assertions: list[Assertion] max_tool_calls: int timeout_s: float tags: tuple[str, ...] = ()def load_case_set(path: str | Path) -> list[Case]: raw = json.loads(Path(path).read_text()) cases: list[Case] = [] for entry in raw["cases"]: jsonschema.validate(entry, CASE_SCHEMA) # fail loudly on a malformed case, not at run time expected = [ ExpectedTool(t["tool_name"], t["required"], t.get("expected_args", {})) for t in entry.get("expected_tools", []) ] cases.append(Case( case_id=entry["case_id"], difficulty=entry["difficulty"], prompt=entry["input"]["prompt"], context=entry["input"].get("context", {}), expected_tools=expected, assertions=[Assertion(a["type"], a["hard_gate"], a.get("negate", False), a.get("params", {})) for a in entry["assertions"]], max_tool_calls=entry.get("max_tool_calls", len(expected)), timeout_s=entry.get("timeout_s", 60.0), tags=tuple(entry.get("tags", [])), )) ids = [c.case_id for c in cases] if len(ids) != len(set(ids)): raise ValueError("duplicate case_id in case set") return cases
jsonschema.validate at load time, not at grading time, is deliberate. A malformed case should break the CI job with a schema error on the line that's wrong, not silently grade as a pass because a typo'd field name made an assertion never fire.
Deterministic graders versus LLM-as-judge
An agent's tool calls are checkable in code. Its final answer, when the task allows more than one correct phrasing, usually is not. That split is the whole decision, and most eval suites get it wrong in one specific direction: they reach for an LLM judge because it feels more thorough, and use it on things a five-line function could check for free, deterministically, forever, at zero marginal cost.
Deterministic grading:
import redef _match_arg(actual: Any, matcher: str) -> bool: kind, _, value = matcher.partition(":") if kind == "exact": return str(actual) == value if kind == "regex": return bool(re.fullmatch(value, str(actual))) if kind == "present": return actual is not None raise ValueError(f"unknown arg matcher: {matcher}")def run_deterministic_assertion(assertion: Assertion, trajectory: "Trajectory") -> bool: p = assertion.params if assertion.type == "exact_match": return trajectory.final_output.strip() == p["value"].strip() if assertion.type == "regex": return bool(re.search(p["pattern"], trajectory.final_output)) if assertion.type == "tool_called": for tc in trajectory.tool_calls: if tc.tool_name != p["tool_name"]: continue arg = p.get("arg") if arg is None: return True if _match_arg(_dig(tc.args, arg), p["matcher"]): return True return False # "json_schema" follows the same shape: json.loads the output, then # jsonschema.validate it against assertion.params["schema"], catching # both JSONDecodeError and ValidationError as a plain False. raise ValueError(f"{assertion.type} is not a deterministic assertion type")def _dig(args: dict, dotted: str) -> Any: cur: Any = args for part in dotted.split("."): cur = cur.get(part) if isinstance(cur, dict) else None return cur
That covers "did it call create_issue with repo=acme/api-service" for exactly the cost of running Python. It never drifts, never costs a token, and gives the same answer on run 1 and run 10,000.
LLM-as-judge earns its cost on the case it's actually for: judging whether a free-form answer is correct and complete when there is no single string that counts as right. Here is a real judge prompt, not a placeholder:
import anthropicJUDGE_MODEL = "claude-sonnet-4-5" # pin an exact model, never "latest"JUDGE_SYSTEM_PROMPT = """You are grading whether an AI agent's response completes a task.You will be given the task, the agent's tool call trace, and its final answer.Score strictly. Do not reward confident wording over correctness.Return only JSON in this exact shape, nothing else:{"pass": true|false, "score": 1-5, "rationale": "one sentence, cite specific evidence"}Rubric:5 - Fully correct, complete, and directly answers what was asked.4 - Correct but missing a minor, non-critical detail.3 - Partially correct or ambiguous; a careful reader would ask a follow-up.2 - Addresses the topic but the core claim is wrong or unsupported.1 - Off-task, refuses without cause, or contradicts the tool results it gathered.pass = true only for a 4 or 5. Never mark pass = true because the tonesounds confident, or because the answer is close."""def judge_with_llm(client: anthropic.Anthropic, case: Case, trajectory: "Trajectory") -> dict: trace = "\n".join( f"- called {tc.tool_name}({tc.args}) -> {'error: ' + tc.error if tc.error else 'ok'}" for tc in trajectory.tool_calls ) or "(no tool calls)" response = client.messages.create( model=JUDGE_MODEL, max_tokens=300, temperature=0, system=JUDGE_SYSTEM_PROMPT, messages=[{"role": "user", "content": f"Task:\n{case.prompt}\n\nTool trace:\n{trace}\n\nFinal answer:\n{trajectory.final_output}"}], ) return json.loads(response.content[0].text)
temperature=0 and a pinned model version are not stylistic choices. A judge sampled at temperature 0.7 makes your CI gate itself nondeterministic, which is the one place nondeterminism cannot be tolerated, and an unpinned model means the definition of "pass" can change on a date you don't control. Both come back in the failure modes section below.
Each is the wrong choice in specific, predictable situations:
Deterministic grading is wrong when a task has more than one legitimately correct output and you didn't enumerate them. Exact-matching a paraphrased-but-correct answer produces a false failure, and the usual fix people reach for, loosening the match with fuzzier and fuzzier regexes, quietly turns the assertion into one that stops checking anything. If the space of correct answers is genuinely open, that's the signal to use a judge, not a more forgiving string comparison.
LLM-as-judge is wrong in five situations that come up constantly: when the answer is checkable in code (a judge here adds cost, adds latency, and adds its own sampling variance for zero benefit over a function call); for anything that gates safety (a judge can be manipulated by the same injected content the agent read, so a hard "must never" cannot rest on a model's opinion); when the judge is the same model family as the agent under test, because they share blind spots and a model will not reliably catch the exact class of mistake it would itself make; when you need a CI-blocking decision and haven't pinned the judge's model version and temperature; and at the scale of hundreds of cases times tens of runs times every pull request, where judge calls become the largest line item in your eval bill for a check that a schema validator would have done for free.
Scoring the trajectory, not just the answer
The first comment on the Day 25 carousel post makes the point that matters most here: an agent can reach a correct final answer after calling the wrong tool three times, and a grader that only looks at the final answer gives that run a perfect score. Output quality and tool selection accuracy have to be scored as separate dimensions, or the second one's failure gets erased by the first one's success.
@dataclass(frozen=True)class ToolCall: tool_name: str args: dict[str, Any] latency_ms: float error: str | None = None@dataclass(frozen=True)class Trajectory: tool_calls: list[ToolCall] final_output: str total_latency_ms: float cost_usd: float@dataclass(frozen=True)class TrajectoryScore: tool_precision: float tool_recall: float arg_precision: float redundant_calls: int score: floatREDUNDANCY_PENALTY = 0.15 # per call beyond the case's max_tool_callsdef score_trajectory(case: Case, trajectory: Trajectory) -> TrajectoryScore: expected_by_name = {t.tool_name: t for t in case.expected_tools} required = {t.tool_name for t in case.expected_tools if t.required} correct_calls = 0 arg_checks = arg_hits = 0 seen_required: set[str] = set() for tc in trajectory.tool_calls: expected = expected_by_name.get(tc.tool_name) if expected is None: continue # called something not on the expected list at all correct_calls += 1 if expected.required: seen_required.add(tc.tool_name) for arg_name, matcher in expected.expected_args.items(): arg_checks += 1 if _match_arg(_dig(tc.args, arg_name), matcher): arg_hits += 1 total_calls = len(trajectory.tool_calls) tool_precision = correct_calls / total_calls if total_calls else 1.0 tool_recall = len(seen_required) / len(required) if required else 1.0 arg_precision = arg_hits / arg_checks if arg_checks else 1.0 redundant_calls = max(0, total_calls - case.max_tool_calls) penalty = min(1.0, redundant_calls * REDUNDANCY_PENALTY) composite = max(0.0, (tool_precision + tool_recall + arg_precision) / 3 - penalty) return TrajectoryScore(tool_precision, tool_recall, arg_precision, redundant_calls, composite)
Four numbers, each catching a different failure. tool_precision drops when the agent calls tools that were never expected, which is what a guessed or hallucinated tool name looks like in this metric. tool_recall drops when a required tool never got called at all, which is a different failure than calling the wrong one. arg_precision catches the case that precision and recall both miss entirely: the agent called the right tool and still passed the wrong repo, the wrong ID, or a stale value it didn't refresh. And redundant_calls charges for the agent that calls the same read-only tool five times "to be sure," which produces a correct trajectory by every other measure while quietly burning latency and cost the other three numbers don't see. If your agent runs behind an MCP gateway with a fixed four-tool surface, this is also where search and get_schema calls belong in expected_tools: an agent that skips straight to guessing a tool name and its arguments is a tool-selection failure even when it happens to guess right.
Reliability: pass rate, pass@k, and why variance beats the mean
Run the same case against the same agent ten times at a nonzero temperature and you get a distribution, not an answer. Reliability is the discipline of reporting that distribution instead of collapsing it into one number too early.
Pass rate is the simplest form: the fraction of N independent runs of a case that passed.
pass@k answers a different question: if the agent gets k independent attempts and any one of them succeeding counts as a win, what's the probability of success? This matters whenever your system actually does retry, or presents the agent with a few candidate completions and picks one. Estimating it naively, by literally sampling k attempts once and checking if any passed, is a high-variance coin flip. The fix, from the HumanEval paper that introduced pass@k for code generation, is to oversample n independent runs (n well above k), count how many passed as c, and use the unbiased estimator:
from math import combdef pass_at_k(n: int, c: int, k: int) -> float: """Unbiased pass@k estimator (Chen et al., 2021).""" if k > n: raise ValueError("k cannot exceed n") if n - c < k: return 1.0 return 1.0 - comb(n - c, k) / comb(n, k)
comb(n - c, k) / comb(n, k) is the probability that all k of a random k-subset drawn from your n runs land among the n - c failures, so one minus that is the probability at least one success is in the subset. Run n = 20 samples of a case, see c = 14 pass, and pass_at_k(20, 14, 1) gives you a lower-variance estimate of single-attempt pass probability than trusting any one of those 20 runs alone would.
Here's why the mean is the wrong headline number even when it's correctly computed. Take two cases that both land at an 85% aggregate pass rate across a full suite:
Case A passes 10/10 runs, every time, for one case, and fails 0/10, every time, for another. Averaged together: 85% doesn't even apply here, this is two separate 100%-and-0% cases, which is a coverage gap, not flakiness. It needs a capability fix: the agent has genuinely never solved that case.
Case B passes 17/20 runs with no pattern to which ones fail. That's a stability problem, not a coverage gap. It needs a determinism fix: lower the temperature, add a retry with a schema-validating parse, or make a tool response less order-sensitive.
A single "85% pass rate" line collapses these into a number that suggests the same fix twice, when the two cases need different engineering work entirely. Report per-case pass rates, not just the suite mean, and flag anything that lands strictly between roughly 10% and 90% over enough runs as unstable rather than partially passing. A case sitting at 0% or 100% is a known quantity. A case sitting at 60% is the one that will pass in your demo and fail for a user an hour later, and it is exactly the failure mode the whole carousel post exists to catch.
def per_case_pass_rates(results: list["RunResult"]) -> dict[str, float]: by_case: dict[str, list[bool]] = {} for r in results: by_case.setdefault(r.case_id, []).append(r.passed) return {cid: sum(v) / len(v) for cid, v in by_case.items()}
Latency and cost: percentiles, not averages, and per completed task, not per call
An average latency is dragged down by every cache hit and every trivially short case, and it hides exactly the number a user actually experiences: the slow tail. Report p50 as the typical case and p95 as the number you'd actually put in an SLA, and gate CI on the p95, never the mean.
import mathdef percentile(values: list[float], p: float) -> float: if not values: return 0.0 data = sorted(values) idx = min(len(data) - 1, math.ceil(p * len(data)) - 1) return data[idx]p95_latency_ms = percentile([r.latency_ms for r in results], 0.95)
Cost has the same problem in a different shape. "Cost per call" makes a flaky agent look cheap: an agent that fails 40% of the time and gets retried looks fine on a per-call basis, because each individual call is inexpensive, and the real cost, the extra calls burned retrying a task that eventually succeeded or never did, is invisible in that number. Cost per completed task charges it correctly:
def cost_per_completed_task(results: list["RunResult"]) -> float: total_cost = sum(r.cost_usd for r in results) completed = sum(1 for r in results if r.passed) return total_cost / completed if completed else float("inf")
float("inf") when nothing completed is the correct answer, not zero. A case that costs money on every attempt and never once succeeds has an undefined, not a low, cost per completed task, and a dashboard that silently reports 0.00 for that case is a dashboard that will hide a total capability failure behind a good-looking number.
Safety assertions are gates, not scored dimensions
Every other dimension in this post produces a number between 0 and 1 that gets averaged into a scorecard. Safety should not, because averaging is exactly the operation that lets a catastrophic failure hide behind good performance everywhere else. An agent that scores 96% on output quality and once, in 50 runs, sends an email nobody approved or leaks a token into its own output does not have a 96%-ish safety story. It has a bug that will happen again, and the average told you the opposite of what you needed to know.
Model it as a boolean check that can block a run outright, independent of every other score:
def hard_gate_violations(case: Case, trajectory: Trajectory, judge_fn=None) -> list[str]: violations = [] for a in case.assertions: if not a.hard_gate: continue condition = ( judge_fn(a, trajectory)["pass"] if a.type == "llm_judge" else run_deterministic_assertion(a, trajectory) ) if condition == a.negate: # negate=True means the condition must NOT hold violations.append(f"{case.case_id}:{a.type}") return violations
Then the aggregation rule is simple and deliberately not statistical: if any run of any case in the suite produces a hard-gate violation, that run fails regardless of its quality score, and the suite as a whole is marked blocked regardless of the aggregate pass rate. A safety gate at 99.98% is not "basically safe." It is a violation waiting for the run that hits it, and the whole point of a hard gate is that it does not get to be outvoted by everything the agent got right.
Wiring it together: the executor and the scorecard
The pieces above compose into an N-run executor and an aggregator. agent_fn here is whatever tool-use loop you already have; the harness doesn't care whether it's built on the Claude SDK's mcp_servers connector or a hand-rolled loop, only that it returns a Trajectory.
import timefrom typing import CallableAgentFn = Callable[[Case], Trajectory]@dataclass(frozen=True)class RunResult: case_id: str passed: bool hard_gate_violations: list[str] quality_score: float trajectory: TrajectoryScore latency_ms: float cost_usd: floatdef run_case_n_times(agent_fn: AgentFn, case: Case, n: int, judge_fn=None) -> list[RunResult]: results = [] for _ in range(n): started = time.monotonic() trajectory = agent_fn(case) elapsed_ms = (time.monotonic() - started) * 1000 violations = hard_gate_violations(case, trajectory, judge_fn) scored = [a for a in case.assertions if not a.hard_gate] quality = sum( (judge_fn(a, trajectory)["pass"] if a.type == "llm_judge" else run_deterministic_assertion(a, trajectory)) != a.negate for a in scored ) / len(scored) if scored else 1.0 results.append(RunResult( case_id=case.case_id, passed=not violations and quality == 1.0, hard_gate_violations=violations, quality_score=quality, trajectory=score_trajectory(case, trajectory), latency_ms=elapsed_ms, cost_usd=trajectory.cost_usd, )) return resultsdef run_suite(agent_fn: AgentFn, cases: list[Case], n: int, judge_fn=None) -> list[RunResult]: results: list[RunResult] = [] for case in cases: results.extend(run_case_n_times(agent_fn, case, n, judge_fn)) return resultsdef build_scorecard(results: list[RunResult]) -> dict: all_violations = [v for r in results for v in r.hard_gate_violations] return { "pass_rate": pass_rate([r.passed for r in results]), "per_case_pass_rate": per_case_pass_rates(results), "p50_latency_ms": percentile([r.latency_ms for r in results], 0.50), "p95_latency_ms": percentile([r.latency_ms for r in results], 0.95), "cost_per_completed_task_usd": cost_per_completed_task(results), "avg_tool_precision": sum(r.trajectory.tool_precision for r in results) / len(results), "avg_redundant_calls": sum(r.trajectory.redundant_calls for r in results) / len(results), "safety_violations": all_violations, "blocked": bool(all_violations), }
This is the whole loop from the carousel's mechanism slide made literal: case set in, N runs through the agent, graded by deterministic checks and a judge where one is actually needed, aggregated into a scorecard. Run it against version 1 of a prompt, save the scorecard, change the prompt, run it again, and now you have something to compare instead of two demos and an opinion.
Wiring evals into CI: thresholds that actually block a merge
An eval suite that only runs manually before a release catches regressions a week after they happened. The suite earns its keep when it runs on every pull request against a committed baseline scorecard and fails the build on a real regression, not on ordinary sampling noise.
The trap here is the phrase "any decrease fails." With a stochastic agent, a case that truly passes 90% of the time will occasionally show 85% or 95% across two different runs of 20 samples each, from binomial sampling noise alone, with nothing in the code having changed. A CI gate that fails on any decrease will be red half the time for no reason, and a team will disable it within a month. The threshold has to be wider than the noise floor, and the noise floor is computable, not guessed:
def wilson_lower_bound(successes: int, n: int, z: float = 1.96) -> float: """95% Wilson score lower bound for a binomial proportion. Use this, not the raw pass rate, as the number that has to clear a threshold.""" if n == 0: return 0.0 p = successes / n denom = 1 + z * z / n center = p + z * z / (2 * n) margin = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) return (center - margin) / denomdef check_regression_gate(current: dict, baseline: dict, n_per_case: int) -> list[str]: failures = [] if current["safety_violations"]: failures.append(f"safety hard gate violated: {current['safety_violations']}") current_lb = wilson_lower_bound(round(current["pass_rate"] * n_per_case), n_per_case) if current_lb < baseline["pass_rate"] - 0.03: failures.append(f"pass rate lower bound {current_lb:.3f} below baseline " f"{baseline['pass_rate']:.3f} minus margin") if current["p95_latency_ms"] > baseline["p95_latency_ms"] * 1.20: failures.append(f"p95 latency regressed: {current['p95_latency_ms']:.0f}ms " f"vs baseline {baseline['p95_latency_ms']:.0f}ms") if current["cost_per_completed_task_usd"] > baseline["cost_per_completed_task_usd"] * 1.25: failures.append("cost per completed task regressed more than 25%") return failures
Four rules, in priority order that matches how much they should matter: any safety violation blocks unconditionally, with no threshold at all. The pass rate check uses a confidence lower bound instead of the raw number, so a run that got unlucky by sampling variance doesn't fail the build, while a genuine regression still gets caught because the lower bound of a truly worse agent stays below the line across most runs, not just one. Latency and cost use a percentage margin, because a small, expected wobble from provider-side latency variance shouldn't block a merge, and a real regression usually blows past 20 to 25% by a wide margin, not by a hair.
Wire it into the pipeline as a real gate, not an informational report:
Baseline promotion should be a deliberate step, not automatic: merging to main updates eval/baseline.json only after a human confirms the new number is actually better, not just different. An automatically ratcheting baseline lets a slow regression creep in one small, individually-approved step at a time.
Seven failure modes that quietly corrupt an eval suite
An eval suite doesn't fail loudly. It keeps reporting numbers, and the numbers keep being wrong in a way nobody notices until a case that should be failing has been passing for months.
1. Judge drift. You pin claude-sonnet-4-5 today. The provider doesn't change model weights under a fixed version string, but if your judge config ever points at an alias like "latest" instead of a dated snapshot, the definition of "pass" moves without a single line of your code changing, and every historical scorecard becomes incomparable to every new one. Pin the exact model string, note it in the baseline file, and treat a judge model upgrade as a breaking change that requires re-baselining, not a free upgrade.
2. Case-set rot. A case asserts the agent's answer matches "the current version of the library is 3.2," and eleven months later the library is on 4.0. The case now fails for a reason that has nothing to do with the agent, and if nobody investigates, the fix people reach for is deleting or loosening the case, which quietly shrinks your coverage. Any case whose expected output depends on external, mutable state needs either a mocked, frozen fixture or a periodic human review pass, not a live dependency on the real world.
3. Flaky tool dependencies. A case calls a real third-party API in CI because mocking felt like extra work. The API rate-limits you during a busy afternoon, the case fails, and the failure gets attributed to the agent instead of the network. Mock or fixture every tool call in the eval harness itself; the agent should not be able to tell the difference between a fixture and a live call, and your suite should never be down because a vendor is.
4. Grader-agent collusion. Using the same model family as both the agent under test and the LLM judge means they share training data, share biases, and share blind spots. A judge built on the same base model as the agent is disproportionately likely to accept the exact category of subtly wrong answer that model tends to produce, because it doesn't recognize the error as an error either. Where the stakes justify it, use a different model family for judging than for the agent, or at minimum audit judge decisions against a human-labeled sample often enough to catch the pattern.
5. Case-set leakage. Cases end up copy-pasted into a few-shot prompt, a system prompt example, or a fine-tuning set, and the agent starts "passing" the eval by having effectively memorized it rather than by generalizing. This is the eval equivalent of training on the test set, and it happens by accident constantly, usually because someone used a real failing case as a helpful example while debugging a prompt. Keep the case set out of anything that becomes training or prompting material, and hold a rotating private subset that never appears anywhere but the harness.
6. Single-run false signal. Someone runs a case once, it passes, and that single boolean gets treated as ground truth about whether the fix worked. At a nonzero temperature, one run is a coin flip with an unknown weight, not an answer. Never accept N=1 as evidence for a case that has ever shown any variance, and for a case you've never sampled more than once, don't assume the result generalizes until you have.
7. Schema drift in tool contracts. An upstream tool's parameter shape changes, an integration gets updated to match, and the eval case's expected_args matchers keep checking the old shape, silently passing because they never look at the field that actually changed. The grading logic and the tool contract it's checking against need to be versioned together, and a good practice is calling get_schema for the tools a case exercises as part of the eval run itself, so a real contract change fails the case loudly instead of the case quietly checking nothing that matters anymore.
The ship checklist
Every case has a fixed case_id, a difficulty tier, and at least one assertion, validated against a committed JSON schema at load time
expected_tools names which tools should be called and which are merely optional, with argument matchers on the ones that matter
Deterministic assertions are used everywhere the correct output is checkable in code; the LLM judge is reserved for genuinely open-ended grading
The judge prompt pins an exact model version and temperature=0, and both are recorded next to every baseline
Tool selection is scored on precision, recall and argument correctness separately from output quality, so a right answer via the wrong path is visible
Redundant tool calls beyond a case's max_tool_calls reduce the trajectory score instead of disappearing into a "it worked" pass
Every case runs N ≥ 10 times before its pass rate means anything, and per-case pass rates are reported, not just the suite-wide mean
pass@k is computed with the unbiased estimator when retries or multiple candidates are part of the real system, not estimated by sampling k once
Any case sitting strictly between roughly 10% and 90% pass rate is flagged unstable and investigated as a determinism problem, not averaged away
Latency is reported as p50 and p95; nothing in the pipeline gates on an average
Cost is reported per completed task, and a case that never completes reports infinite cost, not zero
Safety assertions carry hard_gate: true and are aggregated as a blocking boolean, never averaged into a 0-100 score
CI runs the suite on every pull request against a committed baseline, with a Wilson-bound margin on pass rate and a percentage margin on latency and cost, so the gate blocks real regressions without flaking on sampling noise
Baseline promotion is a deliberate, reviewed step, never an automatic ratchet on every green run
All tool calls in the eval harness hit mocked or fixtured dependencies, never live third-party APIs
The case set lives outside any prompt, few-shot example or fine-tuning data the agent could have seen
Someone has manually reviewed a sample of judge decisions against human judgment in the last quarter
Where this leaves you
Seven dimensions give you a vocabulary. This harness gives you a number you can actually trust twice: once today, and again after the change you're about to make. The pieces are unglamorous on purpose, a schema, a couple of grading functions, a percentile and a binomial confidence bound, because the discipline here isn't clever math, it's refusing to let an average hide a real regression or a single lucky run stand in for reliability.
If your agent already resolves credentials per request the way Day 19 described, the same request context is where an eval run's case_id and attempt number belong, so a failing case traces straight back to the exact run that produced it instead of a grep through logs.
Next in the series: once a score can fail a build, the next step is making the agent enforce its own limits at run time instead of finding out about them after the fact.