companion piece to the Bagman case study. unit tests tell me the code does what i wrote. these tell me what the model does with it.
what an eval has to answer
the caddy is one agent with 33 tools. when a player says "par", the right answer is not a
sentence. it is one call to submit_hole_scores, with the chatting player's roundPlayerId,
three strokes on a par 3, and no other tool calls first. most of what can go wrong is in that
trajectory: the wrong tool, the right tool with the wrong arguments, an id the model made up, or
three reads before the write.
so the evals score the trajectory first and the prose second. they live in
packages/domains/src/packages/ai/evals/ and run against real models through the same
gateway the app uses. a run bills, so it only runs when someone asks for it or on a schedule.
one file per behaviour
a scenario is one expected behaviour, written in a small fluent DSL. this is the whole file for "say par on the current hole":
// evals/scenario/scenarios/score-fast-par.scenario.ts
export default defineAiScenario("score-fast-par")
.given({ surface: "round", round: EVAL_ROUND, roster: ROSTER })
.whenUserSays("par")
.expectIntent("submit_score")
.expectSkills([AI_SKILL_IDS.scoreParsing, AI_SKILL_IDS.scoreSubmission])
.noSkill(AI_SKILL_IDS.statsCoaching)
.noSkill(AI_SKILL_IDS.social)
.expectTools(["submit_hole_scores"])
.withArgs({ "scores[*].strokes": 3 }, "current eval hole is a par 3")
.noTool("read_stats")
.noTool("read_round_context")
.expectStepsAtMost(1)
.golden();
given sets the screen and the fixture round. whenUserSays is the message. every clause
after that maps to one scorer, and each scenario only pays for the clauses it declares. a few
more options cover the cases that needed them:
.orExpectTools(...)for an ask with two defensible answers. "how much am i up on mike" is right throughread_money_breakdownor throughread_statswith the money lens. the score is the best of the two..givenPriorTurns(...)for multi-turn cases. the earlier turns are scripted, not generated, so a run is comparing the same conversation every time..expectAnswer(...)for when the wording matters, judged by a model.
the registry is an explicit list of imports, not a folder scan. a glob would hide the scenarios from the type checker, and a typo in a tool name should fail at compile time. a separate lint test, with no model calls, checks that every tool a scenario mentions is one the agent actually registers. it caught two dangling references when a feature was deleted.
there are 120 scenario files today: 92 in the main set, 13 safety cases, 3 prompt-injection cases, and 12 frontier cases that are expected to fail. the main suite also carries 51 utterance variants migrated from an older JSON corpus, compiled through the same DSL.
eight scorers
| scorer | checks | how |
|---|---|---|
| tool selection | the agent called the expected tools, without loops or blacklisted tools | deterministic |
| tool args | the arguments match, e.g. scores[*].strokes is 3 | deterministic |
| chain optimality | at most N steps, and none of the forbidden tools | deterministic |
| context grounding | ids from the preloaded round appear in the tool calls or the answer | deterministic |
| output shape | each read's result still matches the schema the UI renders | deterministic |
| skill selection | the right prompt skills were loaded for the screen, and the wrong ones weren't | deterministic |
| intent recognition | the model understood what kind of ask it was | model judge |
| answer correctness | the reply is right and grounded, when a scenario asks for it | model judge |
six of the eight need no model. that keeps a run cheap and, more usefully, keeps most of the
score stable from run to run: a deterministic scorer gives the same answer on the same
trajectory. the two judges run on the verifier role, which is a different model from the one
being tested.
the canonical suite's thresholds say how much each dimension is allowed to slip:
| scorer | overall threshold |
|---|---|
| output shape | 1.0 |
| skill selection | 1.0 |
| intent recognition | 0.9 |
| context grounding | 0.9 |
| chain optimality | 0.9 |
| tool selection | 0.85 (and 0.7 within each intent group) |
| tool args | 0.8 |
| answer correctness | 0.8 |
output shape is 1.0 on purpose. if a read's result drifts from the schema the app renders, a card breaks on a phone, and that is not a matter of degree.
the suites
a suite is a selection of items plus thresholds. the caddy's eval folder has ten, and they ask different questions:
- golf-scenarios is the canonical suite: every scenario in the main registry plus the migrated regressions.
- golf-smoke is a 22-item subset of the highest-risk contracts (score writes, banker state, money reads, clarification, guardrail blocks), sized for a pull request.
- golf-game-coverage and golf-rules are focused slices for game setup and per-hole input, and for rulings.
- golf-safety covers money and privacy damage cases.
- golf-injection puts prompt-injection payloads in the round data (a player name, a game name) instead of the user's message. it gates on tool selection at 1.0, because the question is whether a payload ever reached a write tool. it runs on both the live and the reasoning role, since the router sends short asks to the cheap one.
- golf-frontier is the capability ceiling. every threshold is 0, so it never fails a run.
- golf-write-gate, golf-suggest-precision and golf-option-order don't run the caddy at all. they test a classifier used by group chat against labeled corpora: whether it separates ambiguous write requests from clear ones, how precise its "does this need the caddy" call is, and whether its answer changes when the same options are listed in a different order.
a structured artifact suite for round recaps lives with the artifacts domain and uses the same runner.
on a pull request that touches the AI code, the CI workflow is opt-in: adding a ci label
spends a hosted run. a weekly scheduled run executes the full set against the shipped models.
the frontier suite, and why it never fails
a suite the incumbent model already passes can't show you a better model. a stronger candidate can only tie. the frontier suite holds cases no current model passes reliably: resolving "he" through a score that was reassigned two turns earlier, recalling something a player said several holes back, combining several reads into one answer, telling a house rule from the official one, carrying a ruling through to the standings, and compiling a game nobody has described before. scores are recorded per model per week.
a scenario leaves the frontier only by graduating into the canonical suite, once the shipped model passes it reliably. deleting it because models fail it would defeat the point.
the model sweep
runModelSweep runs one suite against a list of targets, either literal models or
role:<name> resolved through the role table, and returns one row per model. a few rules
in it came from mistakes:
- sequential, not parallel. concurrent runs share a rate limit and slow each other down.
- average per scorer, then across scorers. if results were pooled, a scorer that happens to run on more items would dominate the headline number.
- cached input still counts as tokens. treating it as free would flatter whichever model caches most.
- a missing price is a blank, not NaN and not a failed run. pricing is looked up outside the run, so a model with no price row still gets scored.
a weekly Inngest job uses the same function. it trials up to four new candidate models against the smoke suite, then sweeps the frontier suite over the same targets. it never changes a role assignment. the output is a ranking and a suggestion, and a person makes the switch. a 22-item smoke suite is a proxy for quality, not a definition of it, and a role swap changes every round at once.
the first frontier sweep, on five scenarios with GLM 5.2 judging, came out the opposite way from what i expected. Haiku 4.5, the live-round model, scored 0.874 overall. Sonnet 4.6, then the reasoning model, scored 0.849, mostly because it lost on answer correctness (0.800 against 1.000). on tool arguments (0.600), tool selection (0.792) and intent (0.600) the two scored exactly the same. five items is a small sample, but the identical failures said something specific: those cases were hard for the task, not for the model size. that is part of why the turn router keeps round traffic on the cheap role by default.
the regression that guards the context bug
in april 2026 the caddy's round context was being sent with each request and the framework was
dropping it without an error. the model re-read the round with read_round_context on every
turn. the fix moved the context into
the agent's instructions. the evals now make sure the extra read can't come back quietly:
- the tool-selection scorer has a blacklist for preloaded rounds:
read_round_context,read_hole_context, and the framework's skill-loading tools. one call to any of them sets that item's score to 0, regardless of what else was right. - 49 scenarios declare
.noTool("read_round_context")with a step cap, so chain optimality fails as well. - context grounding checks that the ids from the preloaded round are the ids in the tool calls, which only works if the model actually saw them.
a change that puts the context back on the request, or a framework upgrade that resolves instructions differently, shows up as a red tool-selection score on the next run instead of as a slow drift in token counts.
what it doesn't measure yet
- latency. the runner records tokens, not wall-clock time. the frontier suite was meant to include a latency-bounded voice case, and that part is deferred until there is a per-item timer.
- live traffic. in production, a 5% sample of turns gets a relevancy grade and a faithfulness check, but the trajectory scorers need an expected trajectory, so they only run on the dataset.
the useful habit has been writing the scenario before the fix. a bug report becomes one
*.scenario.ts file that fails, then the change that makes it pass, and the file stays.