keeping the caddy's evals honest

October 1, 2026 · 4 min read

the rules around bagman's eval suite: lint before a model call, treat the ci workflow as a contract, keep safety cheap and deterministic, fix the fixture instead of the threshold, and pin every bug the evals find.

companion to the evals behind bagman's caddy, which covers what the suite is. this one is the rules around it: the ones that came from it being wrong.

what came first

before there was a harness, prompt and scorer changes shipped after i looked at traces and decided they seemed fine. that doesn't scale, and a few changes went out that should have been diffed against a regression set first.

the first set was a json file with 53 rows. each row had a message, the intent it should be classified as, the tools the agent should call (strict, relaxed or unordered) and some of the arguments, scored on intent, tool choice and arguments.

two things went wrong with it.

  • a row couldn't say most of what mattered. there was no way to write "don't call this tool", "finish in at most n steps", "use the ids from this round", "keep this output shape" or "the answer is right". the set stayed useful as cheap smoke coverage and stopped there.
  • tool names in data rot quietly. when i deleted the golf bag feature, two rows still named read_club_recommendation and log_shot. nothing failed. they just scored against tools that no longer existed.

the replacement is one typed scenario per behaviour, described in the other post. the 51 rows still worth keeping moved into the same format as regressions, and the json file was deleted.

lint before any model call

every tool a scenario names has to be a real tool on the agent, checked by a test that makes no model calls. it found the two golf-bag leftovers the day it was added.

the lint walked a list of scenario registries, and that list was the next gap: it covered the capability and frontier registries but not safety or injection, so those could name tools that didn't exist and nobody would know. now a test compares the lint's list with the scenario directories on disk, so a new registry can't be added without being linted.

the ci workflow is a contract

which suites run where is decided by a workflow file, and a workflow file is easy to edit without noticing what you dropped. a test reads it and asserts the suite names on each lane.

  • on a pull request (only when it carries the ci label): smoke, safety and injection.
  • weekly, and on demand: the full scenario suite, artifacts, safety, injection and the precision, option-order and write-gate suites. the last three run weekly only, because one has 200 items and another calls the model three times per item.
  • the frontier suite isn't on either lane. the model sweep runs it.

the whole run used to be daily. that was more than the signal needed, the failing runs were mostly noise, and stale copies of the workflow in other repos kept failing on secrets. weekly catches the same regressions.

safety stays cheap and deterministic

the safety suite has four rules:

  • only deterministic gates decide it. answer correctness is measured but can't fail it.
  • it runs on the cheap live-round model. a safety property that only holds on the expensive tier isn't a property.
  • it reuses existing scenarios by id instead of copying them, so a fix lands in one place.
  • it asserts the public tool names, the same names a visitor's client sees.

relax one scenario, fix the fixture

when a scenario misses its bar, the threshold is the last thing to change.

  • the recap scenario allows the writer to retry once, which is correct behaviour, and that retry dipped tool selection to about 0.88. its threshold went to 0.85. the suite's didn't.
  • output shape sat at 0.5 on another scenario. the answer was fine; a test stub wasn't echoing an id back. the stub got fixed and the score went to 1.0 without touching the bar.

pin what you fix

a fixed bug gets an assertion so it stays fixed. when the caddy leaked internal player ids into its answers, the fix came with expectAnswer checks on the standings and projection scenarios.

it works the other way too. a set of privacy adversarial cases, added to probe what one player can learn about another, found two real leaks out of eight. the safety suite stayed red until they were fixed, which is the point of it being a gate.

the judge is on a budget

two scorers use a model as the judge. it gets one step and a capped output, and the policy carries a version so a change in scores can be traced to a change in the judge. the judge's model comes from the same role table as everything else, so there are no model names or price tables in the eval code.

what i'd set up first next time

the scenarios get the attention, but most of what went wrong here was around them: names that pointed at nothing, a registry nobody linted, a lane that could drift. next time the lint, the registry check and the workflow test come before the second scenario.

⌘Kterminal⌘Lask