---
title: "SimRig"
description: "an open-source relay that lets ai agents drive real ios simulators on a mac. a mac host, a cloudflare worker relay and an mcp server: screenshots, the element tree, taps, logs, maestro tests and visual diffs."
canonical: "https://bokendell.com/projects/simrig"
last-updated: 2026-10-01
---

# SimRig

an open-source relay that lets ai agents drive real ios simulators on a mac. a mac host, a cloudflare worker relay and an mcp server: screenshots, the element tree, taps, logs, maestro tests and visual diffs.

## Metadata
- Canonical: https://bokendell.com/projects/simrig
- Markdown: https://bokendell.com/projects/simrig.md
- Lifecycle: building
- Live link: https://github.com/simrig-dev/simrig
- Tags: mcp, ios, typescript, cloudflare, developer-tools
- Updated: 2026-10-01

## Source
```mdx
<CaseStudyHero
	title={meta.title}
	thesis={meta.thesis}
	metric={meta.heroMetric}
	statusRows={[
		{ label: "status", value: "open source · mit · not on npm yet, build from a clone" },
		{ label: "stack", value: "typescript · orpc · mcp · cloudflare workers + durable objects · simctl · maestro" },
		{ label: "repo", value: "github.com/simrig-dev/simrig", href: "https://github.com/simrig-dev/simrig" },
	]}
/>

## The problem

Bagman is an Expo app, so most of what I need to check is on an iPhone screen. When my agents ran on my
Mac, they could boot a simulator and look at it. Then I moved most of the work to cloud sessions, and the
agents lost the device. A cloud sandbox is Linux, it has no Xcode, and it only allows outbound HTTPS and
WebSocket traffic. Nothing can reach into it, and I didn't want to open my Mac to the internet so it could
reach out.

So an agent could change a screen and run the unit tests, but it couldn't see the screen. It would report a
UI change as done without ever having looked at it. I wanted the agent to lease a real simulator, drive it
the way a person would, and bring back proof.

## How it works

The rule that shapes everything is that both ends dial out. The Mac connects out to a relay. The agent
connects out to the same relay. The relay is the only piece with a public address, and it is a small pipe
with a room registry. That is the one layout that works from inside a sandbox that only allows outbound
traffic.

<Mermaid
	caption="both ends dial out"
	chart={`flowchart LR
		agent["agent · mcp over https"] --> relay["relay · cloudflare worker + durable object"]
		cli["rig cli · orpc"] --> relay
		host["mac host · rig host start"] -->|"outbound wss"| relay
		host --> sims["ios simulators · simctl · maestro"]
		classDef hot fill:#ca653c26,stroke:#ca653c,color:#eae3e1;
		class relay,host hot
	`}
/>

There are three parts:

- **the host**, a daemon on the Mac that drives simulators through simctl, agent-device and Maestro, and
  runs the actual MCP server
- **the relay**, one transport-agnostic core served either by a Cloudflare Worker with a Durable Object or
  by plain Node
- **the contract**, the wire protocol, the schemas and the tool table that everything else is generated
  from, including the tool docs

The relay passes MCP calls straight through to the host without interpreting them. It also serves a typed
control plane for the cli and the console, and a separate lane for screenshots and recordings so large
files never go through the rpc path. It has no database on purpose: leases live in the Durable Object,
device state is always read live from simctl, and run history is files in the artifact store.

The host runs on my laptop, or on any Mac with Xcode. Setting it up is a relay deploy and one command on
the Mac, and an agent connects with a single `claude mcp add` line.

<TechStackLine
	groups={[
		{ label: "host", items: ["node 22", "simctl", "agent-device", "maestro"] },
		{ label: "relay", items: ["cloudflare worker + durable object", "or a standalone node relay"] },
		{ label: "contract", items: ["orpc", "zod", "mcp over streamable http"] },
		{ label: "cli", items: ["rig (trpc-cli over orpc)"] },
		{ label: "console", items: ["a dashboard spa served by the relay and the host"] },
	]}
	whyOneLiner="the tool table is the single source: the mcp server, the docs page and the contract tests all read it, so the docs can't describe a tool the server doesn't serve."
/>

## What an agent can do

An agent gets 34 tools. The ones it uses on every run:

- lease a simulator, by model, device class or iOS version, with labels so retries land on the same warm device
- take a screenshot with the accessibility element tree, and act on element refs instead of pixel coordinates
- run a batch of taps, typing and swipes in one round trip
- read app logs, network traffic, crash reports and a compact timeline of what happened
- run a Maestro flow and get back the verdict, each step's result and screenshots of any failing step
- save a screenshot as a baseline and diff the screen against it later, at the device's native pixel size

There are more for the edges a real app hits: deep links, push notifications, the clipboard, app files,
the keychain, StoreKit, Reduce Motion, screen recording, and throwaway email inboxes for sign-in codes.
The server also hands the agent its own operating procedure: lease, look, act, verify, export, record.

I use it on Bagman. A cloud session leases a simulator on my Mac, installs the build, points the app's
Metro at the sandbox through a tunnel, and walks a screen, writing what it saw to a QA board. The same
session that wrote the fix checks it on a device.

<DecisionPair
	chosen={{
		name: "a relay both ends dial into",
		bullets: [
			"works from a sandbox that only allows outbound https and wss",
			"the mac never accepts an inbound connection",
			"the relay is the only public address, and it holds no device state",
			"one core, two adapters: a cloudflare worker or plain node",
		],
	}}
	rejected={{
		name: "expose the mac directly",
		bullets: [
			"a cloud sandbox can't receive a connection, so the agent side still needs a middle",
			"a public port on a laptop is a standing risk for a dev tool",
			"every new agent host would need its own way in",
		],
	}}
	narrative="the relay is deliberately dumb. it forwards mcp, routes control calls and serves artifacts behind signed, expiring urls. everything that touches a device happens on the mac."
/>

<HardProblem
	headline="a cold boot took longer than the agent would wait"
	triedFirst="create-ios-simulator waited until the device was booted and ready, then returned."
	whyFailed="a cold boot on a busy mac can take longer than the client's 60-second tool timeout. the call timed out, the agent retried, and the retry leased a second device while the first was still booting."
>
	`create-ios-simulator` now always answers within 20 seconds. A lease that isn't ready yet comes back as
	`booting`, with its id and an estimate, and the device is the agent's from that moment. The agent waits
	with `get-instance-status`, which can block for up to 50 seconds per call until the lease is ready or
	failed. Create, wait, drive is three calls.

	The retry problem got its own fix: a create can carry a request id, and the same id within ten minutes
	returns the lease the first attempt made instead of leasing another device. A lease that never comes up
	moves itself to `failed` and gives its device back.
</HardProblem>

<HardProblem
	headline="a reused lease could carry last week's build"
	triedFirst="builds are published under names, like myapp-dev, and a lease can ask for names to be installed before it reports ready. with reuseIfExists, a warm lease that had installed the same names was handed back."
	whyFailed="a name is a pointer. republishing myapp-dev changes the bytes behind it, but every lease that recorded the string still matches. the reused device had the old build and reported that it satisfied the request, and from the outside it looked exactly like a fresh install."
>
	Identity is the digest, not the name. When a lease installs a named artifact it records the digest it got.
	Before a warm lease is reused, each name is resolved again and compared to that digest; if any differ, or
	the registry can't answer, the lease isn't reused and a fresh one installs the current build.

	The extra round trip is only paid where the bug can exist. A request that names no artifacts never touches
	the registry.
</HardProblem>

<HardProblem
	headline="a dropped mac left agents polling dead leases"
	triedFirst="the relay kept no lease state. when a mac's connection dropped, calls were refused with one generic error and the health check showed no nodes."
	whyFailed="an agent mid-run couldn't tell a two-second network flap from a rebooted mac. on one day three cloud sessions lost their runs that way, and the cause stayed unknown because the host's log was in /tmp and the mac had rebooted."
>
	The host's heartbeat now carries the list of leases it holds. While a Mac is gone, its leases are marked
	lost and calls naming them are refused with when they were lost. The refusal also says how long to back
	off and the most it should wait, and to create no new leases meanwhile. When the Mac comes back, its next
	heartbeat settles it: a lease the host still reports survived, and one it doesn't is gone for good.

	An empty list and a missing list mean different things. An old host that doesn't report leases says
	"can't say", and the relay doesn't retire its leases on that silence.
</HardProblem>

## Sharing a Mac with agents

A simulator is cheap to drive and expensive to start. I measured it on a 10-core Mac: two booted simulators
being driven held the load average between 6 and 20, while a boot or an app install pushed it to 150 to 240
for minutes. So the host samples its own load, memory, disk and thermal state, and a busy Mac queues a new
lease with a position and an estimate instead of refusing it. A refused caller retries in a loop and adds to
the load it was refused for; a queued one waits.

Every lease also reaps itself. It ends after 15 minutes idle by default, and there is a hard wall-clock cap
that nothing resets, so a stuck agent loop that keeps calling tools can't hold a device forever.

## Where it is now

SimRig is open source under the MIT license at
[github.com/simrig-dev/simrig](https://github.com/simrig-dev/simrig). The packages aren't on npm yet, so
today you build the cli from a clone, and the one-click relay templates install from the published
packages, so they wait on that first release. The publish guard is already in the repo: it checks every
package's manifest, license and files, refuses a published package that depends on a private one, and
keeps every package on one version, because an npm tarball can't be amended after the fact.
```

_Generated from live portfolio data and MDX source. Canonical site: https://bokendell.com_
