01Documentation
Pause And Resume
Sandbox0 can release runtime compute without deleting a sandbox. By default, pause publishes the writable block-COW RootFS head and retires the current runtime allocation. Resume claims a fresh resource-neutral carrier and starts a new runtime generation from the latest committed head.
The default filesystem-only pause preserves files, identity, and durable configuration. It does not preserve processes, memory, sockets, PIDs, open file descriptors, or live REPL connections.
Experimental memory pause/resume explicitly opts into preserving supported process
state with memory: true. It reuses the execution capture and restore machinery
of system maintenance migration,
but retains a checkpoint while no runtime is allocated. Existing calls and TTL
pause keep their filesystem-only behavior. Automatic resume restores memory
when the paused sandbox has a retained checkpoint, as described below.
Default Filesystem-Only Behavior#
| State | Across a successful pause/resume | After delete or hard_ttl |
|---|---|---|
| Sandbox ID and configuration | Preserved | Deleted |
| Writable RootFS | Latest committed generation is restored | Deleted with the sandbox |
| Named RootFS snapshots | Preserved independently until deleted/expired | Follow snapshot retention |
| Processes, memory, sockets, and live sessions | Not preserved | Deleted |
runtime_id | Replaced | Removed |
runtime_generation | Incremented on resume | Removed |
The writable RootFS is a block-COW filesystem. Ordinary RootFS paths such as
/var/tmp and workspace directories persist after a successful checkpoint.
/tmp is a runtime tmpfs and is recreated on pause/resume, as are
virtual mounts such as /proc, /dev, and /sys. Template-defined
ephemeralMounts are also excluded from pause, resume, snapshots, and forks.
Do not store durable data under any of these runtime-only mounts.
For named restore points, fork, and rebase, see Snapshot And Restore.
Explicit Memory Pause And Resume#
This experimental mode is still undergoing failure and capacity validation. Use it only with a deployment that supports memory checkpoints. Both actions accept an optional JSON object:
json{ "memory": true }
Omitting the body, using an empty object, or setting memory: false keeps the
filesystem-only behavior. Choose memory on both pause and resume. An explicit
memory resume requires a retained memory checkpoint matching the committed
RootFS generation; a filesystem-only pause cannot later recover its discarded
process memory. Unsupported or incompatible memory operations return an error
without falling back to an entrypoint restart.
A successful memory pause retains supported guest processes, memory, PIDs,
open files and runtime tmpfs state, together with the matching RootFS checkpoint.
Resume restores that execution image into a fresh compatible carrier. The
sandbox ID stays the same; runtime_id changes and runtime_generation advances.
Existing SSH, streaming and external TCP connections are not guaranteed to
survive. Reconnect clients after resume. A named RootFS snapshot still contains
filesystem state only.
pythonclient.sandboxes.pause_and_wait( sandbox.id, memory=True, timeout_sec=360, ) client.sandboxes.resume_and_wait( sandbox.id, memory=True, timeout_sec=360, )
typescriptawait client.sandboxes.pauseAndWait(sandbox.id, { memory: true, timeoutMs: 360_000, }); await client.sandboxes.resumeAndWait(sandbox.id, { memory: true, timeoutMs: 360_000, });
gooptions := &sandbox0.SandboxLifecycleWaitOptions{ Memory: true, Timeout: 6 * time.Minute, } if _, err := client.PauseSandboxAndWait(ctx, sandbox.ID, options); err != nil { return err } if _, err := client.ResumeSandboxAndWait(ctx, sandbox.ID, options); err != nil { return err }
Memory checkpoint cost depends on the live working set, image transfer,
storage bandwidth and restore compatibility. A warm carrier does not remove
these costs; the normal claim latency target is not a memory restore guarantee.
The example wait budget is a caller setting, not a completion-time guarantee.
Capture uses a bounded regional upload reservation shared with migration. If
that temporary budget is full, preparation retries until its deadline, then
cancels without capture. An image larger than the supported reservation limit
fails at admission. After pause commits and the source is physically terminal, the
temporary budget is released while the retained image stays available to its
owners. Neither case silently switches to filesystem-only pause.
Before capture is authorized, expired preparation is canceled through the exact
source protocol: API admission reopens and unused image staging is released.
The existing process, writer and resource lease stay owned by the running sandbox.
Once capture is authorized, this cancellation path is forbidden; an uncertain
capture requires physical failure resolution. It never retries capture or reports
successful memory retention. After exact source cleanup, the last committed
RootFS remains available, but memory resume fails without a retained image.
The public status is failed even though paused remains true for filesystem
recovery. Memory pause wait helpers raise SandboxLifecycleFailedError instead
of treating that recovery state as a successful pause.
Deletion records termination intent and waits for this cleanup or preparation
cancellation before reclaiming the source's resources.
After restore execution is authorized, a failed target is stopped and physically
reclaimed before its resource lease is released. The failed attempt consumes its
runtime generation while retaining the paused checkpoint and committed disk head.
A later explicit memory: true resume creates a new attempt using that image;
failure cleanup never silently starts a replacement. If restoring the parent of
a running memory fork fails, retrying the fork reports that failure instead of
recapturing the parent or creating another child.
After cleanup, the public status is failed; memory resume wait helpers raise
SandboxLifecycleFailedError for the failed generation.
An HTTP timeout or client disconnection does not cancel an accepted durable
operation. Inspect the sandbox and retry the same memory mode; never retry with
memory: false as an error-recovery shortcut.
Memory capture must match the destination's runsc/platform/CPU and immutable runtime configuration. Changing configuration after capture can prevent restore. If webhooks are configured, capture requires a durable webhook outbox; in-flight delivery and its acknowledgement drain before capture. Restored sessions retain their captured process identity rather than starting a new recovery attempt.
With auto_resume: true, supported inbound access restores a retained memory
checkpoint when the paused sandbox has one. If no memory checkpoint is retained,
automatic wake-up uses filesystem-only resume. A failed or incompatible memory
restore does not fall back to a cold start. An explicit POST /resume without
memory: true retains its filesystem-only default.
Filesystem Pause Transaction#
/api/v1/sandboxes/{id}/pause
bashs0 sandbox pause sb_abc123
Pause is asynchronous:
- manager records planned retirement for the exact sandbox generation, allocation, slot, writer grant, and resource lease;
- manager requests an idempotent stop for the exact Nomad allocation;
- the elected ctld primary fences guest writes, syncs and seals the local block branch, and publishes immutable generation descriptors;
- PostgreSQL atomically advances the RootFS head and commits the sandbox as paused; and
- the plugin-independent terminal reconciler proves runsc, writer, mount, network state, allocation directory, and lease cgroup absent before returning node capacity.
The initial response can report starting/paused: false while this work
continues. Poll sandbox details until status is paused and paused is
true before treating the checkpoint as resume-ready.
A repeated pause request also reports paused: false until the previous
runtime slot has completed physical cleanup. A committed RootFS head alone
does not make the sandbox ready to resume.
After the pause intent is durably recorded, a temporary node-channel outage
or completion deadline still returns 202 Accepted; the pause controller
retries the same operation. This response does not confirm a saved checkpoint.
Failure to record the initial intent still returns an error.
bashs0 sandbox get sb_abc123
pausing is an internal lifecycle transaction, not a public status. Repeating
the pause request is safe: retries recover the same durable operation instead
of creating a second RootFS publication.
If S3 or PostgreSQL is unavailable after local seal, Sandbox0 keeps the exact retirement retryable and keeps its capacity fenced. It does not silently mark the sandbox paused or release a writer before the regional transaction commits.
Filesystem Resume Transaction#
/api/v1/sandboxes/{id}/resume
bashs0 sandbox resume sb_abc123
Resume:
- locks the paused sandbox and reserves the next runtime generation;
- resolves the same attested image/platform and latest committed RootFS head;
- atomically leases a compatible ready carrier and exact node CPU/memory;
- asks ctld to attach RootFS, prepare the lease cgroup, and apply network policy;
- starts stock runsc and procd; and
- commits
runningonly after an authenticated first-command readiness proof.
Resume returns only after the new command-ready binding commits. The new
runtime_id names a different physical allocation and runtime_generation
increases by one.
A resume can fail as unavailable when there is no compatible carrier, node capacity, NBD device, node channel, or object-store access. The sandbox remains paused and its committed RootFS head is unchanged.
Runtime Failure Recovery#
A crash is not treated as a successful pause.
| Condition | Behavior |
|---|---|
| Driver process disappears | Regional reconciliation remains authoritative; cleanup does not depend on the plugin returning. |
| runsc/procd or allocation fails without a planned retirement | ctld fences the old writer and the regional transaction crash-abandons its uncommitted branch. Resume uses the last committed RootFS head. |
| ctld primary fails | The synchronized standby is promoted and resumes the durable node operation. |
| Manager restarts | Another replica recovers the PostgreSQL lifecycle/terminal operation. |
| Node reboots | Only the authenticated successor boot for the same durable node UID can reconcile the old incarnation, after proving old runtime state absent. |
| Nomad server record is purged early | The trusted exact client allocation directory and node absence proof remain required; server-record absence alone is insufficient. |
Writes after the last committed RootFS head can be lost after an unplanned runtime or node failure. Local dirty tail is never promoted without the exact writer and publication proof. Use explicit pause, snapshot, or application sync points when a workflow needs a stronger durability point.
Capacity stays reserved until terminal proof is complete. This avoids running a replacement sandbox on CPU, memory, writer, mount, or network state that may still belong to the failed generation.
Inspect State#
Sandbox details expose:
statusand the equivalentpausedconvenience boolean;- empty
runtime_idwhile paused; - the current
runtime_generation; - soft and hard expiration timestamps; and
- the durable sandbox configuration.
pythonsb = client.get_sandbox("sb_abc123") print("paused=", sb.paused) print("status=", sb.status) print("runtime_id=", sb.runtime_id) print("runtime_generation=", sb.runtime_generation)
During planned pause, the previous generation is not reusable once retirement begins. During resume, supported access may wait for the lifecycle transaction; callers never receive a mixed old/new runtime binding.
TTL And Auto Resume#
| Field | Behavior |
|---|---|
ttl | Soft runtime timeout. Expiry requests the same planned pause transaction. |
hard_ttl | Hard sandbox timeout. Expiry durably requests deletion of identity and RootFS state. |
auto_resume controls whether supported inbound access may wake a paused
sandbox. The wake-up mode follows the paused sandbox's committed state: retained
memory restores its processes, while filesystem-only pause starts a new runtime.
It does not disable platform-initiated crash fencing.
Common auto-resume entry points include sandbox runtime APIs, SSH, and
restartable public routes configured with resume: true. A public route also
needs sandbox auto_resume: true and a restartable cmd or function runtime.
Public URLs stay stable across runtime generations.
Caller-visible temporary results are:
503 unavailablewithsandbox is waking up: an accepted resume has not yet committed; and503 sandbox_resume_failed: that resume attempt ended unsuccessfully.
In filesystem-only mode, a supervised session preserves its logical identity and retained event journal,
then follows its runtime_recovery policy in the new generation. Its old PID
and sockets do not survive.
During graceful runtime shutdown, procd receives the stop signal before its
children so teardown does not turn recoverable sessions into normal exits.
Sessions that had already exited or were explicitly stopped remain terminal;
runtime_recovery: restart applies to sessions active at runtime replacement.
Sandbox0 Cloud Billing Pause#
Sandbox0 Cloud restricts new runtime use when prepaid credit is exhausted. Running sandboxes are automatically paused once outstanding usage and payment debt reach 10% of the team's most recent successful top-up amount. Without a successful top-up, any positive debt triggers the pause. Settlement and regional delivery are asynchronous, so this is a pause threshold rather than a precise limit on final charges.
Billing pause uses the default filesystem-only checkpoint. The sandbox ID and
committed RootFS remain available; processes and memory do not survive. Once
billing admission allows runtime use again, the sandbox can be resumed with a
new runtime generation. hard_ttl continues to apply while a sandbox is paused.
Typical Pattern#
- claim a sandbox with a soft
ttl; - perform active work and create explicit snapshots at important milestones;
- let the sandbox pause after idle time;
- resume explicitly or through an allowed access path; and
- set
hard_ttlwhen durable state must have a maximum lifetime.
Next Steps#
Snapshot And Restore
Create named restore points, fork running or paused state, and rebase a paused RootFS.
ContinueSupervised Sessions
Preserve logical process identity and replayable events across reconnects.
ContinueGateway timeout budget#
Runtime API requests for contexts, sessions, files, and preview creation or renewal may synchronously wake a sandbox. Gateway upstream proxies allow at least 45 seconds for these requests, as they do for explicit create, resume, and fork requests. Cluster-gateway bounds its manager resume call to 45 seconds, including when the access will become a long-lived stream. An earlier caller deadline or cancellation still applies. Metadata requests retain the ordinary proxy timeout, and runtime operations after startup retain their normal procd timeout. The default 60-second gateway write timeout accommodates this startup budget; custom ingress or client timeouts must also allow time for resume.