C_W

◄ Notes01 / 05

One person, many agents, and a phone

How I run many coding agents across a Mac and a rented Linux box from my phone, and the parts you need to build the same.

Topic
Systems
Reading
3 min

Most of my code is now written by agents, and most of my day is spent deciding which ones get to keep going. That only works if I can do it from anywhere, including a phone, without any of them reaching something they shouldn’t. If you want the same, this is the shape of it and the parts it is made of.

Two machines, one private network

The Mac does anything that needs a secret: the env files, the dev servers, the browser for UI work. A rented Linux server does the rest. It is a 16-core box on a one-week trial for a few euros, it runs headless builders that need no credentials, and it has its own logins, so nothing is copied over from the Mac.

Both sit on a private Tailscale network. The server’s firewall drops all public traffic; only the tailnet reaches it. Two small pages are served to the tailnet and nowhere else: a report page and the agent board.

Diagram
flowchart TB
  P["Phone<br/>Moshi · RustDesk · browser"] --> T(("Tailscale<br/>private"))
  T --> MAC
  T --> LIN
  subgraph MAC["Mac · holds the secrets"]
    I["Inbox steward"] --> S["Scrum master"]
    B1["Builders"]
    V["Dev servers<br/>board · reports"]
    S ~~~ B1 ~~~ V
  end
  subgraph LIN["Linux worker · no secrets"]
    B2["Headless builders"]
    L2["Its own logins<br/>Claude · GitHub"]
    B2 ~~~ L2
  end
  MAC --> G["GitHub<br/>cards · checks · reviews"]
  LIN --> G
  G --> Q{"One gate"}:::lit
  Q -->|proven| D["Dev branch"]
  Q -->|needs me| O["Me"]

One session orchestrates

Every agent is a Claude Code session in herdr, a terminal multiplexer that shows whether each one is working, idle, done or blocked, on either machine.

One long-running session never writes product code. It is the scrum master, and everything goes through it: it reads the board, starts builders, passes decisions along, merges what is proven and writes a hand-off note for when I come back. A second session, the inbox steward, reads mail, calendar and chat without touching them and reports to the scrum master, so I hear about what needs me from one place.

GitHub is the feedback loop

Triage is most of the job. Issues become cards on a GitHub Project with five columns, and a card is ready only when it says what done means and how to prove it. A builder claims one with a lease, works in its own git worktree, and opens a pull request.

From there the checks do the talking. CI, tests, linters and review comments, from people and from review bots, are what a builder reads and answers before anyone looks. The gate fits in one line: green checks, no open review threads, no protected paths touched, the right base branch. Anything that touches production, money, or an action an agent was refused comes to me, and a refused action is never retried by another agent.

The machine is code too

Nothing on the workstation is set up by hand. One skill installs it: Ghostty and herdr, Hunk for diffs, Lazygit, Neovim, Raycast, Nerd Fonts, Tailscale and Moshi. Another sets herdr to a baseline. Any change made live also lands in the skill, so running it again rebuilds the same machine. It looks like this site: Departure Mono on an amber theme, with amber kept for one thing, an agent’s state.

Build your own

  1. Put the machines on a tailnet and close everything else: Tailscale SSH and Tailscale Serve for pages only you can open. Relax SSH check mode for the agent user, or every agent waits for your approval at once.
  2. Rent a worker by the week (Hetzner Cloud) and give it its own logins rather than your secrets.
  3. Run agents in Claude Code inside herdr, one git worktree each.
  4. Make the board the source of truth: GitHub Projects, driven with the GitHub CLI, and merges held to required status checks.
  5. Steer from the phone with Moshi for terminals and RustDesk in direct IP mode for the desktop.

Open questions

  • How many agents can one person review before the review becomes the bottleneck?
  • Which decisions can safely move from “ask me” to the gate, and how would I know when one shouldn’t have?
  • Is a second remote worker worth it, or does it just move the waiting somewhere else?