how the caddy's prompt is layered

August 24, 2026 · 7 min read · Updated Oct 1, 2026

bagman's caddy picks a model per role, then builds its prompt in four layers so the parts that don't change during a round are cached for an hour. why the layers split where they do, and the tests that keep the cache from breaking.

companion piece to the Bagman case study. the case study says the caddy uses a four-layer prompt with a model per role. this is how that works and why.

two decisions per turn

every message to the caddy needs two answers before any model runs: which model should answer it, and what goes in the prompt. Bagman answers them separately. a model is chosen by role, and the prompt is assembled by layer. neither depends on the other, which means i can change the model for a job without touching its instructions, and change the instructions without fragmenting the cache.

models by role

a role is a job, not a model. the code asks for live_round or background, and a table in the database says which model fills that role today. the constants in models.constants.ts are only the fallback when the database isn't seeded. everything goes through OpenRouter, so a model swap is a row change and the billing still resolves.

roleused fordefault model
live_roundscore entry, standings, quick in-round asksClaude Haiku 4.5
game_configurationplanning game setups and templatesClaude Haiku 4.5
general_chatchat outside a roundGemini 2.5 Flash
backgroundsummaries, recaps, caddy notes nobody is waiting onQwen3 235B
structuredcompilers and compact structured outputGemini 2.5 Flash Lite
verifiersecond-pass checks and eval judgingGLM 5.2
reasoningrulings, disputes, "explain the press"GLM 5.2
text_guardrailprompt-injection and safety classificationGPT-4.1 nano
vision_guardrailmoderating marks users drawGemini 2.5 Flash
transcriptionspeech to textWhisper large v3 turbo
embeddingsemantic memorytext-embedding-3-small

the rule in the file header is "use the cheapest model that reliably handles each task". the eval history is what decides "reliably", and in at least one case it pointed at the cheaper model: on the hardest scenarios, Haiku 4.5 scored higher than the larger model then in the reasoning role. that is in the evals post.

which role answers a round message

inside a round, a small router picks between live_round and reasoning for each turn (classifyTurnRole in turn-router.utils.ts). it is a few regexes, in this order:

  1. not in a round: general_chat.
  2. a voice turn: always live_round. voice answers are one sentence, spoken between shots.
  3. a rules question ("is that OB?", "can i take a drop"): reasoning, even when it is short, because a wrong ruling costs the hole.
  4. anything under 60 characters: live_round. that is score entry, confirmations and quick reads.
  5. an analytical ask ("explain", "walk me through", "what if", "who's right"): reasoning.
  6. everything else: live_round.

when in doubt it stays cheap. a wrong upgrade costs time in the middle of a round; a wrong stay costs one weaker answer. behind a flag, a calibrated classifier can override the regex result, under a short fixed timeout. on a timeout, an error or a low-confidence answer it falls back to the regex, so a failure costs one regex-routed turn. the regexes have their own labeled corpus test, and one invariant in it matters: no text from a gating eval scenario may upgrade, because those suites run pinned to live_round and would otherwise score a model the user never gets.

agents: one caddy, a few specialists

there is one conversational agent, the caddy, with a flat set of 33 tools: 20 reads that run on their own and 13 writes, 11 of which stop for a confirm tap. it used to be an orchestrator over four sub-agents. each sub-agent only saw the prompt string it was handed, not the round, so i collapsed them in april 2026.

the specialists that remain are self-contained. the dispute mediator and the rules compiler are called directly by the AI service, not as tools of the caddy. the artifact writer is the one delegate, and it gets its whole context in the delegation prompt because the caddy gathers it with read tools first. a game configuration specialist plans setups as structured output on the game_configuration role. none of them can lose the round, because none of them depends on it being passed implicitly.

four prompt layers

the caddy's instructions are a function that Mastra calls once per generation. it returns an array of system messages, and each of the first three carries its own Anthropic cache breakpoint:

// golf-assistant.utils.ts (trimmed)
const blocks: SystemMessage = [
	{ role: "system", content: GLOBAL_BASE, providerOptions: SYSTEM_CACHE_PROVIDER_OPTIONS },
];
const surfaceLayer = instructionLayerForSurface(surface);
if (surfaceLayer) {
	blocks.push({ role: "system", content: surfaceLayer, providerOptions: SYSTEM_CACHE_PROVIDER_OPTIONS });
}
const { stableBlock, volatileContent } = await buildDynamicLayers();
if (stableBlock) {
	blocks.push({ role: "system", content: stableBlock, providerOptions: SYSTEM_CACHE_PROVIDER_OPTIONS });
}
if (volatile) blocks.push({ role: "system", content: volatile }); // never cached
layerwhat's in itchanges whencached
global baseidentity, tool-use rules, response style, golf rules knowledgea deploy1 hour, shared by every user
surfaceguidance for the screen: round, stats or friendsa deploy1 hour, per screen
roundcourse, roster with roundPlayerIds, every hole's par and yardage, who is chattingthe round starts1 hour, per round
turncurrent hole, weather, scorecard, game standings, the voice flag, the user's chosen personaevery turnno

the instruction text in the first two layers comes from typed skill modules (score parsing, score submission, banker bookkeeping, stats coaching and so on). each screen lists the skills it needs, and they are composed into the layer at build time. Mastra can also attach skills as tools the model calls to load instructions, and i deliberately don't do that on the caddy: it would let the model spend a step reading instructions in the middle of a score entry.

why the cache splits there

Anthropic caches the prompt prefix up to each breakpoint, in order: tools, then system blocks. three things follow from that.

  • every screen gets every tool. tool definitions sit at the very start of the prefix. if the stats screen had a different tool subset from the round screen, the caches would diverge from the first byte and nothing would be shared. a cache read costs about a tenth of normal input, so sending all 33 tools everywhere is cheaper than trimming them.
  • the order is least-changing first. the base never changes between deploys, the surface layer changes only with the screen, the round layer only with the round. anything that changes per turn goes last, after the last breakpoint, so it can't invalidate the layers above it.
  • the TTL is an hour, not five minutes. holes are several minutes apart, so a five-minute cache would often expire between turns in a real round. a one-hour write costs more than a five-minute one, but these blocks are written rarely and read on every turn.

the stable/volatile split

the round layer and the turn layer come from the same round context. the round context service returns it as two strings instead of one:

  • stable (formatStableRoundContext): course, round id, the roster with roundPlayerIds and handicaps, all the course holes, and who is chatting. none of that changes during a round.
  • volatile (formatVolatileRoundContext): the current hole and status, the current hole's par and yardage, the weather, each player's score against par, the chatting player's stats, the scorecard so far, and the active games with their standings.

a completed round has no live state, so the whole thing goes in the stable block and the volatile block is empty. the voice flag, the persona and any recent group-chat messages are added to the volatile layer only, so a user changing their caddy's voice never busts a cache anyone else is using.

user-written strings (player names, guest names, game names, course names) are interpolated into the stable block, which the model is told to trust. so each one goes through untrusted() first, which folds newlines, strips control characters and caps the length. without that, a guest named Walk-On followed by a blank line and a ### SYSTEM OVERRIDE heading would put an instruction into every other player's prompt.

what keeps it from drifting

prompt caching fails silently. if a prompt edit leaks a per-turn value into a cached layer, every turn pays full input price and nothing errors. two tests guard it:

  • prompt-tier-stability.test.ts snapshots each cacheable layer and asserts it is byte-stable. it also checks that the global base contains none of the round's markers (roundPlayerId, the current hole, the scorecard, the weather), and another test checks that the voice persona lands in the uncached block.
  • a real-model cache-hit eval runs a two-turn round conversation through the gateway and asserts that turn two reports cached input tokens above zero. it bills, so it only runs when asked, not in the unit suite.

the first one catches the mistake at review time. the second one catches the case where the code is right and the provider isn't caching anyway.

what i'd keep

  • keep the model choice out of the agent graph. a role table made model swaps a data change, and it made the eval sweep possible: the same suite, run against each role's candidates.
  • design the prompt around what changes and how often. the four layers are just the round's timeline (deploy, screen, round, turn) written down in cache order.
  • test the property you care about, not the text. the snapshot catches a changed layer; the cache-hit eval catches a broken cache.
⌘Kterminal⌘Lask