companion piece to the Bagman case study. the case study says the caddy uses a four-layer prompt with a model per role. this is how that works and why.
two decisions per turn
every message to the caddy needs two answers before any model runs: which model should answer it, and what goes in the prompt. Bagman answers them separately. a model is chosen by role, and the prompt is assembled by layer. neither depends on the other, which means i can change the model for a job without touching its instructions, and change the instructions without fragmenting the cache.
models by role
a role is a job, not a model. the code asks for live_round or background, and a table in
the database says which model fills that role today. the constants in
models.constants.ts are only the fallback when the database isn't seeded. everything goes
through OpenRouter, so a model swap is a row change and the billing still resolves.
| role | used for | default model |
|---|---|---|
live_round | score entry, standings, quick in-round asks | Claude Haiku 4.5 |
game_configuration | planning game setups and templates | Claude Haiku 4.5 |
general_chat | chat outside a round | Gemini 2.5 Flash |
background | summaries, recaps, caddy notes nobody is waiting on | Qwen3 235B |
structured | compilers and compact structured output | Gemini 2.5 Flash Lite |
verifier | second-pass checks and eval judging | GLM 5.2 |
reasoning | rulings, disputes, "explain the press" | GLM 5.2 |
text_guardrail | prompt-injection and safety classification | GPT-4.1 nano |
vision_guardrail | moderating marks users draw | Gemini 2.5 Flash |
transcription | speech to text | Whisper large v3 turbo |
embedding | semantic memory | text-embedding-3-small |
the rule in the file header is "use the cheapest model that reliably handles each task". the eval history is what decides "reliably", and in at least one case it pointed at the cheaper model: on the hardest scenarios, Haiku 4.5 scored higher than the larger model then in the reasoning role. that is in the evals post.
which role answers a round message
inside a round, a small router picks between live_round and reasoning for each turn
(classifyTurnRole in turn-router.utils.ts). it is a few regexes, in this order:
- not in a round:
general_chat. - a voice turn: always
live_round. voice answers are one sentence, spoken between shots. - a rules question ("is that OB?", "can i take a drop"):
reasoning, even when it is short, because a wrong ruling costs the hole. - anything under 60 characters:
live_round. that is score entry, confirmations and quick reads. - an analytical ask ("explain", "walk me through", "what if", "who's right"):
reasoning. - everything else:
live_round.
when in doubt it stays cheap. a wrong upgrade costs time in the middle of a round; a wrong stay
costs one weaker answer. behind a flag, a calibrated classifier can override the regex result,
under a short fixed timeout. on a timeout, an error or a low-confidence answer it falls back to the
regex, so a failure costs one regex-routed turn. the regexes have their own labeled corpus test,
and one invariant in it matters: no text from a gating eval scenario may upgrade, because those
suites run pinned to live_round and would otherwise score a model the user never gets.
agents: one caddy, a few specialists
there is one conversational agent, the caddy, with a flat set of 33 tools: 20 reads that run on their own and 13 writes, 11 of which stop for a confirm tap. it used to be an orchestrator over four sub-agents. each sub-agent only saw the prompt string it was handed, not the round, so i collapsed them in april 2026.
the specialists that remain are self-contained. the dispute mediator and the rules compiler
are called directly by the AI service, not as tools of the caddy. the artifact writer is the
one delegate, and it gets its whole context in the delegation prompt because the caddy gathers
it with read tools first. a game configuration specialist plans setups as structured output on
the game_configuration role. none of them can lose the round, because none of them depends on it being passed
implicitly.
four prompt layers
the caddy's instructions are a function that Mastra calls once per generation. it returns an array of system messages, and each of the first three carries its own Anthropic cache breakpoint:
// golf-assistant.utils.ts (trimmed)
const blocks: SystemMessage = [
{ role: "system", content: GLOBAL_BASE, providerOptions: SYSTEM_CACHE_PROVIDER_OPTIONS },
];
const surfaceLayer = instructionLayerForSurface(surface);
if (surfaceLayer) {
blocks.push({ role: "system", content: surfaceLayer, providerOptions: SYSTEM_CACHE_PROVIDER_OPTIONS });
}
const { stableBlock, volatileContent } = await buildDynamicLayers();
if (stableBlock) {
blocks.push({ role: "system", content: stableBlock, providerOptions: SYSTEM_CACHE_PROVIDER_OPTIONS });
}
if (volatile) blocks.push({ role: "system", content: volatile }); // never cached
| layer | what's in it | changes when | cached |
|---|---|---|---|
| global base | identity, tool-use rules, response style, golf rules knowledge | a deploy | 1 hour, shared by every user |
| surface | guidance for the screen: round, stats or friends | a deploy | 1 hour, per screen |
| round | course, roster with roundPlayerIds, every hole's par and yardage, who is chatting | the round starts | 1 hour, per round |
| turn | current hole, weather, scorecard, game standings, the voice flag, the user's chosen persona | every turn | no |
the instruction text in the first two layers comes from typed skill modules (score parsing, score submission, banker bookkeeping, stats coaching and so on). each screen lists the skills it needs, and they are composed into the layer at build time. Mastra can also attach skills as tools the model calls to load instructions, and i deliberately don't do that on the caddy: it would let the model spend a step reading instructions in the middle of a score entry.
why the cache splits there
Anthropic caches the prompt prefix up to each breakpoint, in order: tools, then system blocks. three things follow from that.
- every screen gets every tool. tool definitions sit at the very start of the prefix. if the stats screen had a different tool subset from the round screen, the caches would diverge from the first byte and nothing would be shared. a cache read costs about a tenth of normal input, so sending all 33 tools everywhere is cheaper than trimming them.
- the order is least-changing first. the base never changes between deploys, the surface layer changes only with the screen, the round layer only with the round. anything that changes per turn goes last, after the last breakpoint, so it can't invalidate the layers above it.
- the TTL is an hour, not five minutes. holes are several minutes apart, so a five-minute cache would often expire between turns in a real round. a one-hour write costs more than a five-minute one, but these blocks are written rarely and read on every turn.
the stable/volatile split
the round layer and the turn layer come from the same round context. the round context service returns it as two strings instead of one:
- stable (
formatStableRoundContext): course, round id, the roster withroundPlayerIds and handicaps, all the course holes, and who is chatting. none of that changes during a round. - volatile (
formatVolatileRoundContext): the current hole and status, the current hole's par and yardage, the weather, each player's score against par, the chatting player's stats, the scorecard so far, and the active games with their standings.
a completed round has no live state, so the whole thing goes in the stable block and the volatile block is empty. the voice flag, the persona and any recent group-chat messages are added to the volatile layer only, so a user changing their caddy's voice never busts a cache anyone else is using.
user-written strings (player names, guest names, game names, course names) are interpolated
into the stable block, which the model is told to trust. so each one goes through untrusted()
first, which folds newlines, strips control characters and caps the length. without that, a
guest named Walk-On followed by a blank line and a ### SYSTEM OVERRIDE heading would put an
instruction into every other player's prompt.
what keeps it from drifting
prompt caching fails silently. if a prompt edit leaks a per-turn value into a cached layer, every turn pays full input price and nothing errors. two tests guard it:
prompt-tier-stability.test.tssnapshots each cacheable layer and asserts it is byte-stable. it also checks that the global base contains none of the round's markers (roundPlayerId, the current hole, the scorecard, the weather), and another test checks that the voice persona lands in the uncached block.- a real-model cache-hit eval runs a two-turn round conversation through the gateway and asserts that turn two reports cached input tokens above zero. it bills, so it only runs when asked, not in the unit suite.
the first one catches the mistake at review time. the second one catches the case where the code is right and the provider isn't caching anyway.
what i'd keep
- keep the model choice out of the agent graph. a role table made model swaps a data change, and it made the eval sweep possible: the same suite, run against each role's candidates.
- design the prompt around what changes and how often. the four layers are just the round's timeline (deploy, screen, round, turn) written down in cache order.
- test the property you care about, not the text. the snapshot catches a changed layer; the cache-hit eval catches a broken cache.