I’m the maintainer of Hesper, an MIT-licensed native Mac workspace for Claude Code, Codex and shell agents. The UI is Swift/AppKit/SwiftUI with libghostty, but session ownership and multi-Mac connectivity live in Go.
The problem: several agents can be working on different Macs while their human needs one place to see terminals, notice approval/question states and return to the right session. The app talks only to its local hesperd over a Unix socket; that daemon owns PTYs and connects to the other machines.
Some implementation choices I’d like feedback on:
- One multiplexed link per controller/host pair carries agent events and attaches, with direct connectivity when available and a self-hosted relay as rendezvous/fallback.
- Terminal messages carry at most 32 KiB. A slow local attach past its 4 MiB queue is reattached with a fresh redraw rather than holding up the whole link.
- Input during reconnect is dropped, not replayed. Mutation delivery can have an unknown outcome; we don’t promise exactly-once keystrokes.
I ran the existing transport harness on main (aa76851), on an M1 Max with Go 1.25.5. With injected 20 ms RTT per relay leg (40 ms echo network floor), 300 keystrokes measured 48.4 ms p50 / 59.9 ms p95. Direct loopback measured 0.8 / 3.5 ms. This is a fake-shell transport test with software keys, not app rendering or physical WAN/LAN latency, and not a competitor benchmark.
Engineering note and raw output:
Project and Brew setup: GitHub - derzierau/hesper: Run Claude Code and Codex agents on every Mac you own, and steer them from one fast native window. · GitHub
v0.1.4 is released for macOS 14+, Intel and Apple silicon. Install Claude Code/Codex separately. One Mac needs no relay; multi-Mac currently requires your own relay and device setup. The measurements describe main, not the shipped Brew binary. Native review/diffs, checkpoint-based moves and live-agent forks on main are not yet in v0.1.4.
For people who’ve built PTY or remote-streaming systems: how would you test noisy-terminal interference and communicate dropped input during reconnect? I’d welcome criticism of the recovery boundary and benchmark methodology.