Skip to documentation
API + guides

01Documentation

Sandbox

A Sandbox is an isolated execution environment for long-running agents where you can run code, manage files, and control network access. Each sandbox is created from a Template and has a durable writable RootFS that keeps its accumulated files across successful pause/resume checkpoints for the same sandbox identity. Incremental checkpoints and demand reads on resume avoid copying the entire growing filesystem.

For shell access and file copy over standard SSH clients, see SSH.

ℹ

Sandbox0 Cloud clients should use https://api.sandbox0.ai. Set SANDBOX0_BASE_URL only when you are connecting to a self-hosted or private deployment.

Sandbox Model#

Key Identifiers#

FieldDescription
idUnique sandbox identifier (e.g., sb_abc123)
template_idTemplate used to create this sandbox
team_idTeam that owns this sandbox
runtime_idOpaque identifier of the current physical runtime allocation. Empty while paused.
runtime_generationMonotonically increasing logical runtime generation. Resume and successful system migration commit a new generation.
cluster_idCluster where sandbox runs in multi-cluster deployments. Present on claim and list responses, not on sandbox detail responses.

Lifecycle States#

StatusDescription
startingSandbox is being initialized
runningSandbox is active and ready to use
pausedSandbox has no runtime; identity and latest rootfs checkpoint are preserved
terminatingSandbox identity and durable state are being deleted
failedSandbox encountered an error
ℹ

PostgreSQL stores durable lifecycle intent, the committed runtime generation, the exact Nomad allocation, and its resource/RootFS writer leases. running is published only after ctld and the task driver prove network, RootFS, runsc, and procd command readiness for that generation. During recovery, status returns to starting until a replacement generation commits.

ℹ

Pause and resume use internal lifecycle transactions, but pausing and resuming are not caller-visible statuses. Use the paused boolean as a convenience for status == "paused". Default filesystem-only pause does not preserve running processes, memory, sockets, PID state, or live REPL sessions. Experimental memory pause/resume requires an explicit memory: true request.

System Maintenance Migration#

Experimental system migration can move an eligible running sandbox to another compatible node during planned maintenance. The system selects the destination; there is no user migration API, SDK option, CLI command, or template setting. When enabled by the operator, the same mechanism can consolidate lightly loaded elastic workers onto a compatible fixed worker before a safe scale-in. A worker is removed only after its sandbox leases and runtime slots have been cleared.

A successful migration keeps the sandbox ID and preserves supported process state, memory, guest PIDs, open files, and runtime tmpfs data. The physical runtime_id changes and runtime_generation advances. Execution pauses during the move, and requests or existing connections may fail temporarily. Existing SSH, streaming, and external TCP connections are not guaranteed to survive. A failed restore is never reported as a successful migration by restarting the entrypoint. Recovery after an unrecoverable failure follows the existing filesystem checkpoint policy; it cannot reconstruct lost volatile memory.

This capability is still undergoing failure and capacity validation. It does not change user pause/resume semantics or recover live memory from a dead host.

Persistent Root Filesystem#

Sandbox0 persists the sandbox writable root filesystem as part of checkpointed pause/resume:

  • when ttl expires or pause is requested, Sandbox0 publishes a block-COW rootfs checkpoint before releasing the runtime allocation
  • when the sandbox resumes, Sandbox0 claims a fresh carrier and restores the latest rootfs generation before starting sandbox processes
  • files written to the writable rootfs survive pause/resume for the same sandbox identity after a checkpoint succeeds
  • rootfs checkpoints for the sandbox identity are deleted when the sandbox is deleted or hard_ttl expires

Use the root filesystem for transparent same-sandbox continuity. Use Snapshot And Restore for named rootfs snapshots, restore, and fork operations across sandbox identities.

As an agent keeps working, dependencies, repositories, and outputs accumulate in its RootFS. Pause publishes the writable branch's changes and reuses unchanged blocks. Resume attaches the committed generation and reads blocks on demand rather than downloading the complete filesystem before starting work. This avoids full-RootFS copies on the lifecycle path; latency still depends on unpublished writes, metadata, cache state, object-store access, and the workload's read working set. See Pause And Resume for filesystem and optional memory checkpoint behavior.


Claim a Sandbox#

Claim a sandbox from a template. Manager atomically selects a compatible resource-neutral carrier and leases exact CPU and memory from dedicated node capacity. A claim fails closed when carrier, physical capacity, RootFS device, or node authority is unavailable.

POST

/api/v1/sandboxes

Request Body#

FieldTypeDescription
templatestringTemplate ID to use
snapshot_idstringOptional rootfs snapshot ID used to initialize the writable root filesystem
configobjectOptional sandbox configuration

Sandbox Configuration#

FieldTypeDescription
env_varsobjectSandbox-level default environment variables for new procd-managed processes
ttlintegerTime to live in seconds (soft limit, triggers auto-pause; default: 0, disabled)
hard_ttlintegerHard sandbox lifetime in seconds (deletes identity and durable state; default: 0, disabled)
resourcesobjectOptional per-sandbox resource override. Only resources.memory is accepted; CPU is derived from the platform memory-per-CPU ratio with a 150m minimum limit.
networkobjectSandboxNetworkPolicy. Controls traffic rules, protocol controls, credential bindings, and destination-scoped egress auth
webhookobjectWebhook configuration
auto_resumebooleanAuto-resume when accessed (default: true)
servicesarraySandbox Services for public HTTP routes, including Sandbox Functions

Team quota failures return 429 with error.code set to quota_exceeded. The sandbox_claims policy controls sustained claim rate and immediate burst and includes a Retry-After header when exhausted; active_sandboxes controls the team's running sandbox capacity.

See Network and Protocol Controls for outbound control, Sandbox Services for public HTTP controls, Sandbox Functions for inline public handlers, and Credential for outbound auth and secret handling.

Sandbox env_vars override template/image environment variables for new contexts, supervised session attempts, command services, and function executions. Per-context, per-session, per-command, service runtime, and function env_vars override sandbox env_vars for that narrower scope. The resource-neutral carrier has no prestarted guest process; claim-time variables are available when procd starts.

Sandbox Resources#

Set config.resources.memory at claim time when one sandbox needs a different memory limit from its template. The minimum is 128Mi. The platform maximum defaults to 16Gi and is configured with manager's sandbox_max_memory. Sandbox0 enforces that maximum on template defaults and on sandbox claim, resume, and fork operations. Use the Update Sandbox API with config.resources.memory to change an existing sandbox.

The request only accepts memory. Manager derives CPU from team_template_memory_per_cpu and applies the platform minimum. PostgreSQL records the exact resource lease and ctld enforces it in the sandbox cgroup; Nomad carrier resources are overhead only.

Send a standalone resource update using PUT /api/v1/sandboxes/{id}:

json
{"config":{"resources":{"memory":"4Gi"}}}

Changing resources preserves the sandbox ID and durable files but restarts processes and terminal sessions. A running sandbox first completes filesystem pause and physical cleanup, then starts with a fresh resource lease. A paused sandbox stays paused and uses the new limits on its next start. Retained memory checkpoints are discarded; resize does not restore the old processes.

Submit resources separately from lifecycle and service updates. The operation survives request timeout and manager restart. A 503 can mean the change is still in progress; retry the same limit. Repeating an applied limit is a no-op. A conflicting resize or lifecycle operation returns 409. If fresh capacity is unavailable, files remain preserved in the paused sandbox while retries continue.

The CPU limit is the memory-derived amount rounded up to a whole millicore, with a minimum of 500m (0.5 CPU). The minimum also applies when a stored template predates this rule. Claim, resume and fork allocate a new resource lease; resource updates also replace the lease after completing pause. CPU admission and compute metering use the full amount recorded in that lease.

Claim-time helpers expose this as WithSandboxMemory in Go, memory in Python, memory in TypeScript, and --memory in the s0 CLI.

TTL vs Hard TTL#

Sandbox0 uses two-tier TTL to balance resource efficiency and flexibility:

FieldBehaviorUse Case
ttlRuntime soft pause: When expired, Sandbox0 checkpoints the writable root filesystem, releases the runtime allocation, and marks the sandbox paused.Keep sandboxes alive during active use while freeing compute resources during idle periods.
hard_ttlSandbox hard delete: When expired, Sandbox0 deletes the sandbox identity and durable state, including paused rootfs checkpoints.Enforce a maximum lifetime for compute and storage resources.
ℹ

The relationship: ttl <= hard_ttl. When ttl expires first, sandbox pauses but can be resumed. When hard_ttl expires, the sandbox is deleted and later access returns not found. Resume starts a new runtime generation from the latest rootfs checkpoint only while the sandbox is paused and still within its hard TTL.

When an expiration path is disabled or unset, Sandbox responses return null for the corresponding expires_at or hard_expires_at field. A disabled expiration is never represented as a sentinel date.

Example timeline (sandbox created with ttl=300 and hard_ttl=3600):

  • t=0: Sandbox created
  • t=300: TTL expires → sandbox auto-pauses
  • t=310: User calls refresh → TTL reset to 300, Hard TTL reset to 3600 (new hard deadline at t=3910)
  • t=610: TTL expires again → sandbox auto-pauses
  • t=3910: Hard TTL expires → sandbox identity and durable state are deleted
go
ctx := context.Background() client, err := sandbox0.NewClient( sandbox0.WithToken(os.Getenv("SANDBOX0_TOKEN")), sandbox0.WithBaseURL(os.Getenv("SANDBOX0_BASE_URL")), ) if err != nil { log.Fatal(err) } // Claim a sandbox from the "default" template sandbox, err := client.ClaimSandbox(ctx, "default", sandbox0.WithSandboxHardTTL(300), sandbox0.WithSandboxMemory("512Mi"), sandbox0.WithSandboxEnvVars(map[string]string{ "APP_ENV": "development", }), ) if err != nil { log.Fatal(err) } fmt.Printf("Sandbox ID: %s\n", sandbox.ID) defer client.DeleteSandbox(ctx, sandbox.ID)

Get Sandbox Details#

Retrieve full details about a sandbox.

GET

/api/v1/sandboxes/{id}

go
sb, err := client.GetSandbox(ctx, sandbox.ID) if err != nil { log.Fatal(err) } fmt.Printf("Status: %s\n", sb.Status) fmt.Printf("Template: %s\n", sb.TemplateID) fmt.Printf("Expires at: %s\n", sb.ExpiresAt)

Get Sandbox Status#

Get the current status of a sandbox (lighter weight than full details).

GET

/api/v1/sandboxes/{id}/status

go
status, err := client.StatusSandbox(ctx, sandbox.ID) if err != nil { log.Fatal(err) } fmt.Printf("Status: %s\n", status.Status.Value)

Observability#

Sandbox0 exposes runtime metrics, historical logs, and signed audit events without waking a paused sandbox.

DataUse it forEndpoint
Runtime metricsChart-ready CPU, memory, network, process, and rootfs seriesGET /api/v1/sandboxes/{id}/metrics
Historical logsRetained stdout, stderr, and PTY outputGET /api/v1/sandboxes/{id}/observability/logs
Audit eventsCanonical signed activity recordsGET /api/v1/sandboxes/{id}/observability/events

Audit ingestion and audit queries require the enterprise sandbox_audit feature. Logs, runtime metrics, and the metric catalog do not.

See Observability for metric names, query filters, watch streams, audit integrity, delivery modes, and coverage limits.


List Sandboxes#

List all sandboxes with optional filters.

GET

/api/v1/sandboxes

Query Parameters#

ParameterTypeDescription
statusstringFilter by status (starting, running, paused, terminating, failed)
template_idstringFilter by template ID
pausedbooleanFilter by paused state independently of status
limitintegerMax results per page (default: 50, max: 200)
offsetintegerPagination offset (default: 0)
go
limit := 10 sandboxes, err := client.ListSandboxes(ctx, &sandbox0.ListSandboxesOptions{ Status: "running", TemplateID: "default", Limit: &limit, }) if err != nil { log.Fatal(err) } for _, sb := range sandboxes.Sandboxes { fmt.Printf("- %s (status: %s)\n", sb.ID, sb.Status) }

Update Sandbox#

Update durable lifecycle or service configuration without replacing the runtime allocation.

PUT

/api/v1/sandboxes/{id}

Updatable Fields#

Only the following fields can be updated at runtime:

FieldTypeDescription
ttlintegerTime to live in seconds (soft limit)
hard_ttlintegerHard sandbox lifetime in seconds
auto_resumebooleanAuto-resume when accessed
servicesarraySandbox Services for public HTTP routes, including Sandbox Functions
ℹ

Use PUT /api/v1/sandboxes/{'{id}'}/network for network policy changes. Environment, resource, and webhook changes require a new runtime and are not accepted by this endpoint.

bash
curl -X PUT "$SANDBOX0_API_URL/api/v1/sandboxes/$SANDBOX_ID" \ -H "Authorization: Bearer $SANDBOX0_API_KEY" \ -H "Content-Type: application/json" \ -d '{"config":{"ttl":600,"hard_ttl":3600,"auto_resume":true}}'

Pause And Resume#

Pause and resume is covered in a dedicated page because it affects TTL, auto_resume, service routes, SSH, and webhook behavior.

See Pause And Resume for explicit pause, state inspection, resume, and auto-resume behavior.

See Snapshot And Restore for named rootfs snapshots, restore, and fork operations. Snapshot and fork accept a running or paused source sandbox; restore requires a paused target sandbox.


Refresh Sandbox TTL#

Extend the sandbox time-to-live. This resets both ttl and hard_ttl (if configured) from the current time while the sandbox has a runtime.

POST

/api/v1/sandboxes/{id}/refresh

Request Body#

FieldTypeDescription
durationintegerDuration to extend TTL in seconds (optional, defaults to original TTL)
ℹ

If duration is not specified, both ttl and hard_ttl are reset to their original configured values. Use this to keep a sandbox alive as long as the user is actively using it.

go
// Refresh with default duration (original TTL) resp, err := client.RefreshSandbox(ctx, sandbox.ID, nil) if err != nil { log.Fatal(err) } fmt.Printf("New expires at: %s\n", resp.ExpiresAt) // Refresh with custom duration (e.g., 1 minute) resp, err = client.RefreshSandbox(ctx, sandbox.ID, &apispec.SandboxRefreshRequest{ Duration: apispec.NewOptInt32(60), }) if err != nil { log.Fatal(err) } fmt.Printf("New expires at: %s\n", resp.ExpiresAt)

Delete Sandbox#

Terminate and delete a sandbox. This action is irreversible.

DELETE

/api/v1/sandboxes/{id}

A successful response means the deletion intent is durable. Runtime teardown and persistent RootFS cleanup finish asynchronously. For Nomad-backed Sandboxes, the manager fences the exact allocation and its writer authority before deleting the logical record; it does not depend on the Nomad task-driver process remaining available.

go
_, err = client.DeleteSandbox(ctx, sandbox.ID) if err != nil { log.Fatal(err) } fmt.Println("Sandbox deleted")

Next Steps#