SimRig

when my coding agents moved to the cloud, they lost the simulator. simrig gives them a real one back, on a mac i own, without opening a single inbound port.

0mcp tools · one relay · both ends dial out · mit licensed
status
open source · mit · not on npm yet, build from a clone
stack
typescript · orpc · mcp · cloudflare workers + durable objects · simctl · maestro

The problem

Bagman is an Expo app, so most of what I need to check is on an iPhone screen. When my agents ran on my Mac, they could boot a simulator and look at it. Then I moved most of the work to cloud sessions, and the agents lost the device. A cloud sandbox is Linux, it has no Xcode, and it only allows outbound HTTPS and WebSocket traffic. Nothing can reach into it, and I didn't want to open my Mac to the internet so it could reach out.

So an agent could change a screen and run the unit tests, but it couldn't see the screen. It would report a UI change as done without ever having looked at it. I wanted the agent to lease a real simulator, drive it the way a person would, and bring back proof.

How it works

The rule that shapes everything is that both ends dial out. The Mac connects out to a relay. The agent connects out to the same relay. The relay is the only piece with a public address, and it is a small pipe with a room registry. That is the one layout that works from inside a sandbox that only allows outbound traffic.

rendering diagram…
both ends dial out

There are three parts:

  • the host, a daemon on the Mac that drives simulators through simctl, agent-device and Maestro, and runs the actual MCP server
  • the relay, one transport-agnostic core served either by a Cloudflare Worker with a Durable Object or by plain Node
  • the contract, the wire protocol, the schemas and the tool table that everything else is generated from, including the tool docs

The relay passes MCP calls straight through to the host without interpreting them. It also serves a typed control plane for the cli and the console, and a separate lane for screenshots and recordings so large files never go through the rpc path. It has no database on purpose: leases live in the Durable Object, device state is always read live from simctl, and run history is files in the artifact store.

The host runs on my laptop, or on any Mac with Xcode. Setting it up is a relay deploy and one command on the Mac, and an agent connects with a single claude mcp add line.

host: node 22, simctl, agent-device, maestro · relay: cloudflare worker + durable object, or a standalone node relay · contract: orpc, zod, mcp over streamable http · cli: rig (trpc-cli over orpc) · console: a dashboard spa served by the relay and the host

the tool table is the single source: the mcp server, the docs page and the contract tests all read it, so the docs can't describe a tool the server doesn't serve.

What an agent can do

An agent gets 34 tools. The ones it uses on every run:

  • lease a simulator, by model, device class or iOS version, with labels so retries land on the same warm device
  • take a screenshot with the accessibility element tree, and act on element refs instead of pixel coordinates
  • run a batch of taps, typing and swipes in one round trip
  • read app logs, network traffic, crash reports and a compact timeline of what happened
  • run a Maestro flow and get back the verdict, each step's result and screenshots of any failing step
  • save a screenshot as a baseline and diff the screen against it later, at the device's native pixel size

There are more for the edges a real app hits: deep links, push notifications, the clipboard, app files, the keychain, StoreKit, Reduce Motion, screen recording, and throwaway email inboxes for sign-in codes. The server also hands the agent its own operating procedure: lease, look, act, verify, export, record.

I use it on Bagman. A cloud session leases a simulator on my Mac, installs the build, points the app's Metro at the sandbox through a tunnel, and walks a screen, writing what it saw to a QA board. The same session that wrote the fix checks it on a device.

Chose

a relay both ends dial into

  • works from a sandbox that only allows outbound https and wss
  • the mac never accepts an inbound connection
  • the relay is the only public address, and it holds no device state
  • one core, two adapters: a cloudflare worker or plain node
Over

expose the mac directly

  • a cloud sandbox can't receive a connection, so the agent side still needs a middle
  • a public port on a laptop is a standing risk for a dev tool
  • every new agent host would need its own way in

the relay is deliberately dumb. it forwards mcp, routes control calls and serves artifacts behind signed, expiring urls. everything that touches a device happens on the mac.

a cold boot took longer than the agent would wait

create-ios-simulator now always answers within 20 seconds. A lease that isn't ready yet comes back as booting, with its id and an estimate, and the device is the agent's from that moment. The agent waits with get-instance-status, which can block for up to 50 seconds per call until the lease is ready or failed. Create, wait, drive is three calls.

The retry problem got its own fix: a create can carry a request id, and the same id within ten minutes returns the lease the first attempt made instead of leasing another device. A lease that never comes up moves itself to failed and gives its device back.

What I tried first

create-ios-simulator waited until the device was booted and ready, then returned.

Why it failed

a cold boot on a busy mac can take longer than the client's 60-second tool timeout. the call timed out, the agent retried, and the retry leased a second device while the first was still booting.

a reused lease could carry last week's build

Identity is the digest, not the name. When a lease installs a named artifact it records the digest it got. Before a warm lease is reused, each name is resolved again and compared to that digest; if any differ, or the registry can't answer, the lease isn't reused and a fresh one installs the current build.

The extra round trip is only paid where the bug can exist. A request that names no artifacts never touches the registry.

What I tried first

builds are published under names, like myapp-dev, and a lease can ask for names to be installed before it reports ready. with reuseIfExists, a warm lease that had installed the same names was handed back.

Why it failed

a name is a pointer. republishing myapp-dev changes the bytes behind it, but every lease that recorded the string still matches. the reused device had the old build and reported that it satisfied the request, and from the outside it looked exactly like a fresh install.

a dropped mac left agents polling dead leases

The host's heartbeat now carries the list of leases it holds. While a Mac is gone, its leases are marked lost and calls naming them are refused with when they were lost. The refusal also says how long to back off and the most it should wait, and to create no new leases meanwhile. When the Mac comes back, its next heartbeat settles it: a lease the host still reports survived, and one it doesn't is gone for good.

An empty list and a missing list mean different things. An old host that doesn't report leases says "can't say", and the relay doesn't retire its leases on that silence.

What I tried first

the relay kept no lease state. when a mac's connection dropped, calls were refused with one generic error and the health check showed no nodes.

Why it failed

an agent mid-run couldn't tell a two-second network flap from a rebooted mac. on one day three cloud sessions lost their runs that way, and the cause stayed unknown because the host's log was in /tmp and the mac had rebooted.

Sharing a Mac with agents

A simulator is cheap to drive and expensive to start. I measured it on a 10-core Mac: two booted simulators being driven held the load average between 6 and 20, while a boot or an app install pushed it to 150 to 240 for minutes. So the host samples its own load, memory, disk and thermal state, and a busy Mac queues a new lease with a position and an estimate instead of refusing it. A refused caller retries in a loop and adds to the load it was refused for; a queued one waits.

Every lease also reaps itself. It ends after 15 minutes idle by default, and there is a hard wall-clock cap that nothing resets, so a stuck agent loop that keeps calling tools can't hold a device forever.

Where it is now

SimRig is open source under the MIT license at github.com/simrig-dev/simrig. The packages aren't on npm yet, so today you build the cli from a clone, and the one-click relay templates install from the published packages, so they wait on that first release. The publish guard is already in the repo: it checks every package's manifest, license and files, refuses a published package that depends on a private one, and keeps every package on one version, because an npm tarball can't be amended after the fact.

⌘Kterminal⌘Lask