custom architecture checks as guardrails for coding agents

September 30, 2026 · 8 min read · Updated Oct 1, 2026

coding agents copy whatever pattern is nearest. i wrote 67 architecture rules into my cli so the nearest pattern stays correct, and a verification gate that fails when it checked nothing.

companion piece to the swarm case study. the case study mentions the architecture checks in a paragraph. this is the long version: what they enforce, how they run, and the gates that make an agent prove its work.

The setup

Most of the code in my repos is now written by coding agents, a lot of it in cloud sessions I'm not watching. The agents are good. They are also very consistent about one thing: they copy the nearest example. If the file next door declares its types inline, the new file will too. If one router reaches into a repository directly, the next router will.

I used to fix that in review, one comment at a time, and I kept writing the same comments. A review comment that I have written five times should be a check. A check doesn't get tired, it runs on every change, and an agent reads its error message the same way it reads mine.

There are two kinds of guardrail here. The first reads the code: swarm check arch. The second makes the agent prove its work: skills that end in a command, and a gate whose green actually means something.

What the rules enforce

swarm check arch runs two things. dependency-cruiser checks import boundaries. Then 67 semantic rules, one file each in swarm's check package, look for patterns that an import graph can't see. Several started as a local review script in the Bagman repo and moved into swarm so every repo gets them.

The frontend rules are mostly about where things live:

  • one-component-per-file: one primary component per file; supporting components get extracted
  • inline-declaration-no-inline-types, inline-declaration-no-inline-constants and inline-declaration-no-inline-utils: a component, hook or route file doesn't declare reusable types, constants or helpers inline. They go in a sibling types, constants or utils file
  • co-located-own-folder: once a file has grown siblings like .types.ts, .utils.ts or a test, it moves into its own folder, so it's obvious which siblings belong to which file
  • no-types-folder: one types.ts next to the code, not a types/ folder full of small files

Then the ones about behaviour and boundaries:

  • no-emoji: no emoji in source. Emoji in code tends to leak into UI strings, log lines and commit messages
  • no-process-env-outside-config: every app has one zod-validated env module, and nothing else reads process.env
  • layer-purity-router-no-direct-repository-or-db: an api router calls services, never a repository or the database
  • service-mutation-requires-policy: a service method that changes data has to show an authorization check
  • missing-test: services, repositories, routers and util modules ship with a test file next to them
  • no-silent-catch: a catch that swallows the error and returns nothing is a finding
  • no-slop-copy: in user-facing strings, no mid-sentence em dash, no exclamation marks, no apology filler, no placeholder names. It's opt-in per repo, because it's house voice, not architecture

None of these are deep. That's the point. They are the things an agent gets wrong because the codebase around it got them wrong once, and each one is cheap to check mechanically.

Exceptions are written down

The rules have escape hatches, and an escape hatch has to say why. Most rules take an opt-out comment on the line above the code, with a reason:

// arch-allow no-process-env: read before the config module exists
const configPath = process.env.APP_CONFIG_PATH;

That keeps exceptions searchable. When I want to know where a rule is being bent, I grep for its name, and every hit comes with a sentence explaining it. An agent that wants to skip a rule has to write that sentence too, which is usually enough to make it fix the code instead.

Wider exceptions go in the repo's arch config, by glob, with severity set per rule: error, warn or off, and overrides that apply to part of the tree. Some rules are opt-in because they only make sense once a repo is ready for them.

How it runs

The check runs in each repo's git hooks, on commit and again on push, with --affected. That diffs the branch against its base and only scans the packages that changed. If that diff can't be computed, it tries the last commit instead, and if that fails too, it scans everything, which is slower but never narrower.

New rules don't land as errors on a codebase that already breaks them. They land as a baseline of warnings. In swarm's own repo, a handful of the strict rules are set to warn everywhere, with a comment calling it a transitional baseline: the rules stay visible while I convert one package at a time. Bagman used a ratchet for one rule: error in the areas that were already clean, warn in the roughly 170 call sites that weren't. Each time an area was migrated, its glob came off the warn list, so the rule could never regress what was clean or block someone halfway through what wasn't. When the list was empty, I deleted it and the rule became an error everywhere.

rendering diagram…
where the checks run

A check that finds nothing has not passed

The failure I worry about most isn't a rule that's too strict. It's a rule that quietly matches nothing.

That has happened three ways. A batch of swarm's domain rules looked for domain code under a path with the app name in it, and Bagman's domain package doesn't have one. A planted vendor SDK import in a service and a repository call in a background job both passed with zero findings. The fix went into swarm, and a later release made several rules read Effect v4 code, which they had also been silently skipping.

The third was in Bagman's mobile app. The layer rules are scoped by folder name: one only looks inside containers/, another only inside components/. A package that invented its own top-level folder made everything inside it invisible to all of them. When one of those files moved to containers/, where it belonged, it turned red on the first run. The fix was to split every one-off folder into the layers it belonged to rather than allowlist them.

A planted violation is what found the first one. That's the test worth running on any new rule: watch it fail before you trust its green.

Skills that end in a command

The second kind of guardrail is the skills. Bagman has 35 repo skills, one per kind of change: adding a domain service, an api endpoint, a database migration, a mobile screen, an email template, a background job. Each is a short how-to that loads when an agent works in that part of the repo, and each one ends in the command that proves the change is done. The domain-change skill says the structure is machine-enforced and names ./swarm check arch --affected as a step, not a suggestion.

Skills drift, so they have evals. There are 31 eval cases across 26 of the skills. Each case gives an agent a scenario and grades the answer, and most of them test a refusal: the agent should not run a migration from a worktree, should not call an unrun test lane green, should say that a mutation with no policy check is incomplete and name the rule that will catch it. The rule for adding a case is that the miss is the case: when an agent breaks a skill's rule in real work, that exact situation becomes the next eval.

The gate, and why a raw lefthook run proves nothing

The skill that matters most is verification, and its rule is short: run pnpm gate, and cite that.

Agents used to verify by running the pre-push hook directly and quoting the result. On a branch whose head was already pushed, that command exits 0 in about 0.09 seconds having run nothing. Lefthook scopes a push hook to the push range; an already-pushed head has an empty range, so every job is skipped with "no matching push files", and the summary still reads like a success. It failed most reliably in exactly the situation where an agent was double-checking work it had already pushed.

The hook isn't wrong. Skipping an empty push is the correct behaviour for a git hook. The mistake was using a hook as a gate: a hook is scoped to a push, a gate is scoped to a change, and on a pushed head those two scopes diverge.

You also can't fix it inside the hook config. Lefthook decides to skip before it reads any job, so a guard job would be skipped along with everything else. So pnpm gate is a small wrapper script:

  • it runs the same jobs with --force, so each job computes its own affected scope from the change, not from what git is about to push
  • the first job appends to an evidence file. The gate deletes that file before the run and requires it after, so if lefthook ran nothing, the gate fails no matter what exit code lefthook returned
  • it warns when the tree is identical to the base, because "everything passed" on an unchanged tree says the base is green, not that the change is

The same bug turned up one layer down. A freshly cloned cloud session has a shallow clone with no merge base, so turbo can't compute --affected. It prints "Unable to query SCM", falls back to zero packages, runs no tasks, and exits 0. The unit-test job then reported success in under four seconds having tested nothing. The gate now looks for that message and fails the run, because a broken base invalidates every affected-scoped job in it at once. The fix is to deepen the clone, not to ignore the warning.

The verification skill ends with the line I'd put on every agent's wall: never report a check you did not execute, and never quote a run whose output you did not read.

What I'd tell someone starting

Turn the review comments you keep repeating into rules, and make every exception carry a reason. Roll new rules out as warnings and ratchet them up, so they never block work that predates them. Plant a violation before you trust a rule's silence. And give agents one command whose green can't be produced by doing nothing, then tell them it's the only one they're allowed to quote.

⌘Kterminal⌘Lask