Every AI agent gets its own microVM: a real Linux machine behind a hardware
boundary. Set one up once, then fork it mid-task to try many approaches in
parallel. Forks share their parent's memory and disk, so each stores only what it changes.
shared with parentcopied on writeeach square is one 4 KiB page
— 01snapshot & fork / branch a live machine
Get to a good state once. Fork from it forever.
Cloning, installing and warming caches takes minutes. Do it once, then fork the
running machine as often as you like. Each fork wakes with the same memory, files and
processes, then goes its own way. Try several fixes at once and keep the one that passes.
try_fixes.pypip install upstm-py
import os
from upstream import Upstream
client = Upstream(api_key=os.environ["UPSTREAM_API_KEY"])
with client.sandbox(template="python") as base:
base.exec(["sh", "-c", f"git clone {REPO} app && cd app && uv sync"])
# one prepared machine, one fork per candidate fix
forks = [base.fork() for _ in patches]
for sb, patch inzip(forks, patches):
sb.write_file("/workspace/app/fix.patch", patch)
r = sb.exec(["sh", "-c", "cd app && git apply fix.patch && uv run pytest -q"])
print(sb.sandbox_id, "pass"if r.exit_code == 0else"fail")
— snapshot
Capture the whole machine
Memory, CPU state and disk are captured together, so a fork resumes mid-flight instead of booting cold.
— parent
The original keeps going
The parent pauses only while its state is captured, then carries on. Forking never costs you the machine you forked.
— forks
Each fork stands alone
Its own ID, token and network. Keep it, fork it again or throw it away without touching its siblings.
— 02sharing / one copy of what's identical
Store it once. Share it with every VM.
Most of a sandbox is identical to the one next to it: the same OS, runtimes and packages.
upstream keeps one copy of the common parts and gives each sandbox a private layer on
top. That is what makes a fork cheap, and it fits more sandboxes on every server.
Template snapshotsharedpre-built environment: Python, Node, Claude Code, tools
Base imagesharedminimal Linux + guest agent
↑ private on top · shared underneath
memory
Shared page by page. A fork maps its parent's memory instead of copying it. A userfaultfd handler tracks which VMs use each 4 KiB page and copies one only when a VM writes to it.
disk
Deduplicated by content. Images and snapshots are split into 512 KiB blocks named by their SHA-256 hash. A block used by a thousand sandboxes is stored once.
startup
Restored, not booted. New sandboxes are restored from a shared template snapshot, and a warm pool keeps some ready in advance, so claiming one skips the boot.
— 03isolation / a machine, not a container
Agents run untrusted code. Give them a wall.
Agents run whatever the model writes, so sharing must never weaken the boundary. Shared
pages and blocks are read-only, and every write lands in a sandbox's private layer.
Each sandbox is a Firecracker microVM, the technology behind AWS Lambda, with its own
kernel, network and disk. A container would share the host's kernel.
— 01 / compute
Its own kernel
Hardware virtualisation separates each sandbox from its neighbours and the host. An agent can be root inside and still reach nothing outside.
vmm
Firecracker on KVM
pid 1
Rust guest agent
fs api
path-confined with openat2
— 02 / network
Its own network
Sandboxes cannot reach each other, and outbound traffic goes only where policy allows it.
netns
dedicated namespace + TAP
egress
nftables rules per VM
dns
per-namespace forwarder
— 03 / access
Nothing listening
There is no open port on our servers to attack. Every request is checked three times before it reaches a sandbox, and its token opens that sandbox alone.
ingress
outbound tunnel only
auth
OIDC at three layers
scope
HMAC token per sandbox
— 04uses / what it's for
Anywhere an agent needs a real computer.
01
Coding agents
Run Claude Code, Codex or your own agent with a full shell and no risk to your laptop or production.
02
Code interpreter
Give an assistant its own Python machine to analyse uploaded files, run code and hand back results.
03
Evals & training
Start every rollout from an identical machine state, so differences come from the model, not the setup.
04
CI runners
Build and test every change, including agent-written ones, in a fresh VM restored from a warm template.
05
Live previews
Expose a port from inside a sandbox so a person or an agent can open the running app in a browser.
06
Debug from failure
Fork the machine the moment a test fails, then dig into the copy without re-running any setup.
07
Checkpoint & resume
Snapshot a long-running agent's machine and restore it later, exactly where it left off.
08
Untrusted code
Execute code from users, models or pull requests behind a VM boundary with controlled egress.
— 05how it's built / the whole stack, owned
From the page fault to the browser.
Built in-house, mostly in Rust, on open foundations like Firecracker and Kubernetes. It
runs on dedicated bare-metal servers rather than inside someone else's cloud, and the
same stack can run on your hardware.
compute
Firecracker microVMs on KVM · Rust guest agent as PID 1 · warm pool of restored VMs per template
memory
userfaultfd copy-on-write across fork chains · per-page reverse index · lz4 spill for diverged pages
storage
content-addressed block store · copy-on-write NBD overlays per sandbox · snapshot prefetch
network
namespace + TAP per VM · nftables egress policy · per-sandbox DNS
security
no inbound ports · OIDC at the edge, the proxy and the gateway · HMAC-scoped sandbox tokens · per-tenant rate limits
platform
Kubernetes on Talos Linux · bare metal · ingress only through an outbound Cloudflare tunnel
interfaces
ConnectRPC + REST API · WebSocket terminal · Python SDK · React console · metered per compute-second
Set up once. Fork everything.
One API, a Python SDK, and a browser console with a live terminal.