01Documentation
Install
This is the shortest production-like path from zero to a running self-hosted Sandbox0.
Prerequisites#
- Kubernetes
1.35+ - Helm
3.8+ kubectlmatching your cluster minor version- A default
StorageClass
Kubernetes 1.35+ is required for the supported Sandbox0 runtime stack. Checkpointed pause/resume depends on ctld on sandbox nodes so Sandbox0 can save and restore writable rootfs checkpoints when releasing and recreating runtime pods.
Keep the Kubernetes ContainerRestartRules feature gate enabled. Sandbox0 uses its container-level restart policy to leave a terminated procd attempt available for crash rootfs recovery while keeping the ReplicaSet-compatible pod restart policy.
0) Create a Local Kind Cluster#
Create the cluster:
bashkind create cluster --config kind-config.yaml
kind-config.yaml:
yamlkind: Cluster apiVersion: kind.x-k8s.io/v1alpha4 name: sandbox0 nodes: - role: control-plane image: kindest/node:v1.35.0 kubeadmConfigPatches: - | kind: ClusterConfiguration apiServer: extraArgs: enable-aggregator-routing: "true" extraPortMappings: # cluster-gateway HTTP port - containerPort: 30080 hostPort: 30080 # registry port for template image push - containerPort: 30500 hostPort: 30500
1) Install infra-operator#
bashhelm repo add sandbox0 https://charts.sandbox0.ai helm repo update helm install infra-operator sandbox0/infra-operator \ --namespace sandbox0-system \ --create-namespace
Verify operator + CRD:
bashkubectl get pods -n sandbox0-system kubectl get crd sandbox0infras.infra.sandbox0.ai
2) Apply the Fullmode Sample#
Use the fullmode sample for a complete single-cluster data plane:
Fullmode requires Linux Kubernetes nodes because ctld uses host networking, FUSE, and node-local runtime access.
bashkubectl apply -f https://raw.githubusercontent.com/sandbox0-ai/sandbox0/main/infra-operator/chart/samples/single-cluster/fullmode.yaml
Watch status:
bashkubectl get sandbox0infra -n sandbox0-system -w
Expected status.phase lifecycle:
InstallingUpgrading(when reconciling changes)Ready
3) Initial Admin Credentials#
When local password bootstrap is enabled, if spec.initUser.passwordSecret is not provided, operator generates admin-password.
When built-in auth is enabled, the gateway bootstraps the initial admin with that password on first start. When built-in auth is disabled but OIDC is enabled, the operator skips generating the password secret and the gateway still bootstraps the admin user and initial team, but leaves the account passwordless so the first OIDC login for the same email can bind to that admin account.
bashADMIN_PASSWORD="$(kubectl get secret admin-password -n sandbox0-system -o jsonpath='{.data.password}' | base64 -d)" printf 'username: %s\npassword: %s\n' '[email protected]' "$ADMIN_PASSWORD"
4) Install s0#
macOS and Linux:
bashcurl -fsSL https://raw.githubusercontent.com/sandbox0-ai/s0/main/scripts/install.sh | bash
Windows PowerShell:
powershellirm https://raw.githubusercontent.com/sandbox0-ai/s0/main/scripts/install.ps1 | iex
Or with Go:
bashgo install github.com/sandbox0-ai/s0/cmd/s0@latest
Manual release archives are available from GitHub Releases. See the s0 README for platform-specific manual install steps.
5) Configure the API URL, Log In, and Optionally Create an API Key#
For the local kind setup above, the API endpoint is http://localhost:30080.
That is because the kind config maps host port 30080, and the sample Sandbox0Infra exposes cluster-gateway on the same NodePort.
Export the base URL, then log in with the initial admin account:
bashexport SANDBOX0_BASE_URL="http://localhost:30080" s0 auth login
At that point, interactive s0 CLI usage is ready.
Create an API key only if you want to use the SDKs or raw HTTP API:
bash# If the server did not auto-select a team yet: # s0 team list # s0 team use <team-id> unset SANDBOX0_TOKEN export SANDBOX0_TOKEN="$(s0 apikey create --name sdk-quickstart --role developer --expires-in 30d --raw)"
After that, continue with Get Started to make your first Sandbox0 API request.
Advanced: Deploy on GCP#
For Google Cloud, use GKE Standard rather than GKE Autopilot.
Recommended path:
- Create a GKE Standard regional cluster with Linux nodes.
- Pin Kubernetes to an explicit
1.35+version. Do not rely on the GKE default version. - Keep a regular node pool for host-level system components.
- Add a GKE Sandbox (
gVisor) node pool for sandbox workloads. - Install
infra-operator. - Apply the GCP
gVisorsample. - Expose
cluster-gatewaythrough aLoadBalancer. - Verify the deployment by logging in with
s0and claiming a sandbox.
1) Pick a Supported GKE Version#
Sandbox0 requires Kubernetes 1.35+.
Before creating the cluster, query the versions available in your region:
bashgcloud container get-server-config \ --zone us-east1-b \ --format='yaml(validMasterVersions)'
Choose any available 1.35.x version from validMasterVersions, then reuse it below as ${GKE_VERSION}.
2) Create a Regional Cluster#
This creates a three-zone regional cluster in us-east1, with one regular Linux node per zone:
bashexport GKE_VERSION="1.35.1-gke.1616000" gcloud container clusters create sandbox0-gke-gvisor \ --region us-east1 \ --node-locations us-east1-b,us-east1-c,us-east1-d \ --machine-type e2-standard-2 \ --disk-size 30 \ --num-nodes 1 \ --cluster-version "${GKE_VERSION}"
Get kubeconfig and confirm the nodes:
bashgcloud container clusters get-credentials sandbox0-gke-gvisor --region us-east1 kubectl get nodes -o wide
Keep a regular node pool for core system services. In the GKE gVisor sample, cluster-gateway, manager and its storage runtime, PostgreSQL, and the builtin registry run on the regular pool. The ctld HA pair and sandbox workloads run on the gVisor node pool, while ctld itself uses the node's host-compatible runtime.
3) Add a gVisor Node Pool#
Add one gVisor node per zone:
bashgcloud container node-pools create gvisor-lite \ --cluster sandbox0-gke-gvisor \ --region us-east1 \ --machine-type e2-standard-2 \ --disk-size 30 \ --num-nodes 1 \ --sandbox type=gvisor
Verify both pools:
bashkubectl get nodes -L cloud.google.com/gke-nodepool,sandbox.gke.io/runtime
Expected result:
default-poolnodes have nosandbox.gke.io/runtimelabelgvisor-litenodes showsandbox.gke.io/runtime=gvisor
4) Install infra-operator#
bashhelm repo add sandbox0 https://charts.sandbox0.ai helm repo update helm install infra-operator sandbox0/infra-operator \ --namespace sandbox0-system \ --create-namespace
Verify operator + CRD:
bashkubectl get pods -n sandbox0-system kubectl get crd sandbox0infras.infra.sandbox0.ai
5) Apply the GCP gVisor Sample#
Use the official gVisor sample, then switch cluster-gateway to LoadBalancer:
bashkubectl apply -f https://raw.githubusercontent.com/sandbox0-ai/sandbox0/main/infra-operator/chart/samples/single-cluster/fullmode-gke-gvisor.yaml
bashkubectl patch sandbox0infra fullmode -n sandbox0-system --type merge -p \ '{"spec":{"services":{"clusterGateway":{"service":{"type":"LoadBalancer","port":80}}}}}'
Wait for the deployment to finish reconciling:
bashkubectl get sandbox0infra -n sandbox0-system -w kubectl get svc -n sandbox0-system
Expected result:
Sandbox0InfrabecomesReadyfullmode-cluster-gatewayis exposed asLoadBalanceron port80fullmode-ctld-aandfullmode-ctld-brun ongvisor-lite; the elected primary owns rootfs, volume portal, and network runtime responsibilities- sandbox pods in
tpl-defaultrun ongvisor-lite
fullmode-manager serves lifecycle and storage APIs. On every sandbox node, the elected ctld primary owns rootfs persistence, volume portals, kubelet registration, and network enforcement; the synchronized standby takes over after promotion.
This sample uses the native GKE node label sandbox.gke.io/runtime=gvisor for sandbox placement. You do not need to add a separate custom node label just to target gVisor nodes.
RuntimeClass and Rootfs Clean/Restore#
services.manager.config.sandboxRuntimeClassName only tells Sandbox0 which Kubernetes RuntimeClass to put on sandbox Pods. It does not install runsc, edit the node's containerd config, restart containerd, or create a custom runtime handler.
On GKE Sandbox, Google manages the standard gVisor runtime setup and the sample uses sandboxRuntimeClassName: gvisor. On self-managed clusters, install the runtime handler and RuntimeClass before pointing Sandbox0 at it.
For gVisor rootfs checkpoint and restore support, use a dedicated handler with shared rootfs access. gvisor-rootfs is just an example name; the Kubernetes RuntimeClass.handler must match the handler configured in containerd on every sandbox node.
toml# /etc/containerd/config.toml version = 2 [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.gvisor-rootfs] runtime_type = "io.containerd.runsc.v1" [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.gvisor-rootfs.options] TypeUrl = "io.containerd.runsc.v1.options" ConfigPath = "/etc/containerd/runsc-rootfs.toml"
toml# /etc/containerd/runsc-rootfs.toml [runsc_config] overlay2 = "none" file-access = "shared"
Restart containerd after changing node runtime configuration, then create the matching RuntimeClass:
yamlapiVersion: node.k8s.io/v1 kind: RuntimeClass metadata: name: gvisor-rootfs handler: gvisor-rootfs
Then configure Sandbox0 sandbox Pods to use it:
yamlspec: services: manager: config: sandboxRuntimeClassName: gvisor-rootfs
The sandbox RuntimeClass does not apply to ctld. Keep ctld on the node's host-compatible runtime; it must not run under gvisor, gvisor-rootfs, or kata.
ctld assumes the node's containerd data root is /var/lib/containerd by default. If your node bootstrap stores containerd data somewhere else, set services.ctld.containerdHostDataRoot to the actual host path:
yamlspec: services: ctld: containerdHostDataRoot: /var/lib/containerd/sandbox0
For large writable rootfs checkpoints, ctld uses a disposable node-local object cache capped at 20Gi per node by default. Tune the cap and free-space floor for your node disk:
yamlspec: services: ctld: rootfsObjectCacheMaxBytes: 20Gi rootfsObjectCacheMinFreeBytes: 5Gi
6) Expose the API Publicly#
After the LoadBalancer patch above, GKE allocates a public IP for cluster-gateway.
Check the service:
bashkubectl get svc fullmode-cluster-gateway -n sandbox0-system -o wide
Wait until EXTERNAL-IP is no longer <pending>, then export the public base URL directly from the service:
bashexport SANDBOX0_BASE_URL="http://$(kubectl get svc fullmode-cluster-gateway -n sandbox0-system -o jsonpath='{.status.loadBalancer.ingress[0].ip}')" printf '%s\n' "$SANDBOX0_BASE_URL"
For production on GCP, this is a better default than exposing a node IP through NodePort.
7) Verify with s0#
Retrieve the generated admin password:
bashADMIN_PASSWORD="$(kubectl get secret admin-password -n sandbox0-system -o jsonpath='{.data.password}' | base64 -d)" printf 'username: %s\npassword: %s\n' '[email protected]' "$ADMIN_PASSWORD"
Log in from the public endpoint:
bashs0 auth login \ --api-url "${SANDBOX0_BASE_URL}" \ --email [email protected] \ --password "${ADMIN_PASSWORD}"
Claim a sandbox to verify the deployment is actually usable:
bashs0 sandbox create --api-url "${SANDBOX0_BASE_URL}" -t default
Optional deeper check:
bashs0 sandbox list --api-url "${SANDBOX0_BASE_URL}" s0 sandbox exec <sandbox-id> --api-url "${SANDBOX0_BASE_URL}" -- sh -lc 'echo ready && uname -s'
If s0 sandbox create returns a sandbox ID and the sandbox reaches running, the GCP deployment is working end to end.
Important Notes#
- Sandbox0 requires different runtime behavior for different components. On GKE, keep a regular node pool for system services and a
gVisornode pool for sandbox workloads. - Sandbox workload runtime selection should be handled by cluster defaults or administrator-managed service configuration, and the GCP
gVisorsample places sandbox workloads onto nodes labeledsandbox.gke.io/runtime=gvisor. - Use
sandboxNodePlacementto place sandbox workloads and the ctld HA pair onto the same sandbox node pool. infra-operatordoes not install containerd runtime handlers. Installrunsc/containerd-shim-runsc-v1, configure the handler, restart containerd, and create theRuntimeClassbefore using a customsandboxRuntimeClassName.- If containerd uses a non-default data root, set
services.ctld.containerdHostDataRootto that host path so rootfs pause/resume can use the overlayfs fast path. infra-operatorruns two independent ctld DaemonSets per sandbox node, with one ctld container in each Pod. The elected primary owns the CSI endpoint, kubelet registration, FUSE request channels, and network runtime. The synchronized standby preserves existing mounts across a single ctld process or Pod failure and assumes all node-local responsibilities during promotion. A filesystem, mount, or newly established network flow in progress during handoff can return an error and should be retried. Node reboot, kernel failure, and simultaneous loss of both ctld Pods remain outside this protection boundary.ctlddepends on host features such ashostNetworkandhostPath. This is why GKE Autopilot is not a good fit for full Sandbox0 deployments.- Even when ctld is scheduled onto
gVisornodes, its Pod uses the node's default host runtime. - If you hit GCP
SSD_TOTAL_GBquota limits while creating node pools, reduce--disk-sizeexplicitly instead of relying on the larger default boot disk size.
Why Not Autopilot#
Autopilot is not suitable for fullmode because Sandbox0 host-level components need capabilities that Autopilot restricts. Ctld needs node-level access, while template sandbox pods run on the gVisor pool.