Skip to documentation
API + guides

01Documentation

Pause And Resume

Sandbox0 can release runtime compute without deleting a sandbox. By default, pause publishes the writable block-COW RootFS head and retires the current runtime allocation. Resume claims a fresh resource-neutral carrier and starts a new runtime generation from the latest committed head.

ℹ

The default filesystem-only pause preserves files, identity, and durable configuration. It does not preserve processes, memory, sockets, PIDs, open file descriptors, or live REPL connections.

Experimental memory pause/resume explicitly opts into preserving supported process state with memory: true. It reuses the execution capture and restore machinery of system maintenance migration, but retains a checkpoint while no runtime is allocated. Existing calls and TTL pause keep their filesystem-only behavior. Automatic resume restores memory when the paused sandbox has a retained checkpoint, as described below.

Default Filesystem-Only Behavior#

StateAcross a successful pause/resumeAfter delete or hard_ttl
Sandbox ID and configurationPreservedDeleted
Writable RootFSLatest committed generation is restoredDeleted with the sandbox
Named RootFS snapshotsPreserved independently until deleted/expiredFollow snapshot retention
Processes, memory, sockets, and live sessionsNot preservedDeleted
runtime_idReplacedRemoved
runtime_generationIncremented on resumeRemoved

The writable RootFS is a block-COW filesystem. Ordinary RootFS paths such as /var/tmp and workspace directories persist after a successful checkpoint. /tmp is a runtime tmpfs and is recreated on pause/resume, as are virtual mounts such as /proc, /dev, and /sys. Template-defined ephemeralMounts are also excluded from pause, resume, snapshots, and forks. Do not store durable data under any of these runtime-only mounts.

For named restore points, fork, and rebase, see Snapshot And Restore.

Explicit Memory Pause And Resume#

This experimental mode is still undergoing failure and capacity validation. Use it only with a deployment that supports memory checkpoints. Both actions accept an optional JSON object:

json
{ "memory": true }

Omitting the body, using an empty object, or setting memory: false keeps the filesystem-only behavior. Choose memory on both pause and resume. An explicit memory resume requires a retained memory checkpoint matching the committed RootFS generation; a filesystem-only pause cannot later recover its discarded process memory. Unsupported or incompatible memory operations return an error without falling back to an entrypoint restart.

A successful memory pause retains supported guest processes, memory, PIDs, open files and runtime tmpfs state, together with the matching RootFS checkpoint. Resume restores that execution image into a fresh compatible carrier. The sandbox ID stays the same; runtime_id changes and runtime_generation advances. Existing SSH, streaming and external TCP connections are not guaranteed to survive. Reconnect clients after resume. A named RootFS snapshot still contains filesystem state only.

python
client.sandboxes.pause_and_wait( sandbox.id, memory=True, timeout_sec=360, ) client.sandboxes.resume_and_wait( sandbox.id, memory=True, timeout_sec=360, )
typescript
await client.sandboxes.pauseAndWait(sandbox.id, { memory: true, timeoutMs: 360_000, }); await client.sandboxes.resumeAndWait(sandbox.id, { memory: true, timeoutMs: 360_000, });
go
options := &sandbox0.SandboxLifecycleWaitOptions{ Memory: true, Timeout: 6 * time.Minute, } if _, err := client.PauseSandboxAndWait(ctx, sandbox.ID, options); err != nil { return err } if _, err := client.ResumeSandboxAndWait(ctx, sandbox.ID, options); err != nil { return err }

Memory checkpoint cost depends on the live working set, image transfer, storage bandwidth and restore compatibility. A warm carrier does not remove these costs; the normal claim latency target is not a memory restore guarantee. The example wait budget is a caller setting, not a completion-time guarantee. Capture uses a bounded regional upload reservation shared with migration. If that temporary budget is full, preparation retries until its deadline, then cancels without capture. An image larger than the supported reservation limit fails at admission. After pause commits and the source is physically terminal, the temporary budget is released while the retained image stays available to its owners. Neither case silently switches to filesystem-only pause. Before capture is authorized, expired preparation is canceled through the exact source protocol: API admission reopens and unused image staging is released. The existing process, writer and resource lease stay owned by the running sandbox. Once capture is authorized, this cancellation path is forbidden; an uncertain capture requires physical failure resolution. It never retries capture or reports successful memory retention. After exact source cleanup, the last committed RootFS remains available, but memory resume fails without a retained image. The public status is failed even though paused remains true for filesystem recovery. Memory pause wait helpers raise SandboxLifecycleFailedError instead of treating that recovery state as a successful pause. Deletion records termination intent and waits for this cleanup or preparation cancellation before reclaiming the source's resources.

After restore execution is authorized, a failed target is stopped and physically reclaimed before its resource lease is released. The failed attempt consumes its runtime generation while retaining the paused checkpoint and committed disk head. A later explicit memory: true resume creates a new attempt using that image; failure cleanup never silently starts a replacement. If restoring the parent of a running memory fork fails, retrying the fork reports that failure instead of recapturing the parent or creating another child. After cleanup, the public status is failed; memory resume wait helpers raise SandboxLifecycleFailedError for the failed generation.

An HTTP timeout or client disconnection does not cancel an accepted durable operation. Inspect the sandbox and retry the same memory mode; never retry with memory: false as an error-recovery shortcut.

Memory capture must match the destination's runsc/platform/CPU and immutable runtime configuration. Changing configuration after capture can prevent restore. If webhooks are configured, capture requires a durable webhook outbox; in-flight delivery and its acknowledgement drain before capture. Restored sessions retain their captured process identity rather than starting a new recovery attempt.

With auto_resume: true, supported inbound access restores a retained memory checkpoint when the paused sandbox has one. If no memory checkpoint is retained, automatic wake-up uses filesystem-only resume. A failed or incompatible memory restore does not fall back to a cold start. An explicit POST /resume without memory: true retains its filesystem-only default.

Filesystem Pause Transaction#

POST

/api/v1/sandboxes/{id}/pause

bash
s0 sandbox pause sb_abc123

Pause is asynchronous:

  1. manager records planned retirement for the exact sandbox generation, allocation, slot, writer grant, and resource lease;
  2. manager requests an idempotent stop for the exact Nomad allocation;
  3. the elected ctld primary fences guest writes, syncs and seals the local block branch, and publishes immutable generation descriptors;
  4. PostgreSQL atomically advances the RootFS head and commits the sandbox as paused; and
  5. the plugin-independent terminal reconciler proves runsc, writer, mount, network state, allocation directory, and lease cgroup absent before returning node capacity.

The initial response can report starting/paused: false while this work continues. Poll sandbox details until status is paused and paused is true before treating the checkpoint as resume-ready.

A repeated pause request also reports paused: false until the previous runtime slot has completed physical cleanup. A committed RootFS head alone does not make the sandbox ready to resume.

After the pause intent is durably recorded, a temporary node-channel outage or completion deadline still returns 202 Accepted; the pause controller retries the same operation. This response does not confirm a saved checkpoint. Failure to record the initial intent still returns an error.

bash
s0 sandbox get sb_abc123

pausing is an internal lifecycle transaction, not a public status. Repeating the pause request is safe: retries recover the same durable operation instead of creating a second RootFS publication.

If S3 or PostgreSQL is unavailable after local seal, Sandbox0 keeps the exact retirement retryable and keeps its capacity fenced. It does not silently mark the sandbox paused or release a writer before the regional transaction commits.

Filesystem Resume Transaction#

POST

/api/v1/sandboxes/{id}/resume

bash
s0 sandbox resume sb_abc123

Resume:

  1. locks the paused sandbox and reserves the next runtime generation;
  2. resolves the same attested image/platform and latest committed RootFS head;
  3. atomically leases a compatible ready carrier and exact node CPU/memory;
  4. asks ctld to attach RootFS, prepare the lease cgroup, and apply network policy;
  5. starts stock runsc and procd; and
  6. commits running only after an authenticated first-command readiness proof.

Resume returns only after the new command-ready binding commits. The new runtime_id names a different physical allocation and runtime_generation increases by one.

A resume can fail as unavailable when there is no compatible carrier, node capacity, NBD device, node channel, or object-store access. The sandbox remains paused and its committed RootFS head is unchanged.

Runtime Failure Recovery#

A crash is not treated as a successful pause.

ConditionBehavior
Driver process disappearsRegional reconciliation remains authoritative; cleanup does not depend on the plugin returning.
runsc/procd or allocation fails without a planned retirementctld fences the old writer and the regional transaction crash-abandons its uncommitted branch. Resume uses the last committed RootFS head.
ctld primary failsThe synchronized standby is promoted and resumes the durable node operation.
Manager restartsAnother replica recovers the PostgreSQL lifecycle/terminal operation.
Node rebootsOnly the authenticated successor boot for the same durable node UID can reconcile the old incarnation, after proving old runtime state absent.
Nomad server record is purged earlyThe trusted exact client allocation directory and node absence proof remain required; server-record absence alone is insufficient.
ℹ

Writes after the last committed RootFS head can be lost after an unplanned runtime or node failure. Local dirty tail is never promoted without the exact writer and publication proof. Use explicit pause, snapshot, or application sync points when a workflow needs a stronger durability point.

Capacity stays reserved until terminal proof is complete. This avoids running a replacement sandbox on CPU, memory, writer, mount, or network state that may still belong to the failed generation.

Inspect State#

Sandbox details expose:

  • status and the equivalent paused convenience boolean;
  • empty runtime_id while paused;
  • the current runtime_generation;
  • soft and hard expiration timestamps; and
  • the durable sandbox configuration.
python
sb = client.get_sandbox("sb_abc123") print("paused=", sb.paused) print("status=", sb.status) print("runtime_id=", sb.runtime_id) print("runtime_generation=", sb.runtime_generation)

During planned pause, the previous generation is not reusable once retirement begins. During resume, supported access may wait for the lifecycle transaction; callers never receive a mixed old/new runtime binding.

TTL And Auto Resume#

FieldBehavior
ttlSoft runtime timeout. Expiry requests the same planned pause transaction.
hard_ttlHard sandbox timeout. Expiry durably requests deletion of identity and RootFS state.

auto_resume controls whether supported inbound access may wake a paused sandbox. The wake-up mode follows the paused sandbox's committed state: retained memory restores its processes, while filesystem-only pause starts a new runtime. It does not disable platform-initiated crash fencing.

Common auto-resume entry points include sandbox runtime APIs, SSH, and restartable public routes configured with resume: true. A public route also needs sandbox auto_resume: true and a restartable cmd or function runtime. Public URLs stay stable across runtime generations.

Caller-visible temporary results are:

  • 503 unavailable with sandbox is waking up: an accepted resume has not yet committed; and
  • 503 sandbox_resume_failed: that resume attempt ended unsuccessfully.

In filesystem-only mode, a supervised session preserves its logical identity and retained event journal, then follows its runtime_recovery policy in the new generation. Its old PID and sockets do not survive. During graceful runtime shutdown, procd receives the stop signal before its children so teardown does not turn recoverable sessions into normal exits. Sessions that had already exited or were explicitly stopped remain terminal; runtime_recovery: restart applies to sessions active at runtime replacement.

Sandbox0 Cloud Billing Pause#

Sandbox0 Cloud restricts new runtime use when prepaid credit is exhausted. Running sandboxes are automatically paused once outstanding usage and payment debt reach 10% of the team's most recent successful top-up amount. Without a successful top-up, any positive debt triggers the pause. Settlement and regional delivery are asynchronous, so this is a pause threshold rather than a precise limit on final charges.

Billing pause uses the default filesystem-only checkpoint. The sandbox ID and committed RootFS remain available; processes and memory do not survive. Once billing admission allows runtime use again, the sandbox can be resumed with a new runtime generation. hard_ttl continues to apply while a sandbox is paused.

Typical Pattern#

  1. claim a sandbox with a soft ttl;
  2. perform active work and create explicit snapshots at important milestones;
  3. let the sandbox pause after idle time;
  4. resume explicitly or through an allowed access path; and
  5. set hard_ttl when durable state must have a maximum lifetime.

Next Steps#

Gateway timeout budget#

Runtime API requests for contexts, sessions, files, and preview creation or renewal may synchronously wake a sandbox. Gateway upstream proxies allow at least 45 seconds for these requests, as they do for explicit create, resume, and fork requests. Cluster-gateway bounds its manager resume call to 45 seconds, including when the access will become a long-lived stream. An earlier caller deadline or cancellation still applies. Metadata requests retain the ordinary proxy timeout, and runtime operations after startup retain their normal procd timeout. The default 60-second gateway write timeout accommodates this startup budget; custom ingress or client timeouts must also allow time for resume.