The sandboxed transport
Run every work loop in a separate isolated process instead of inside the daemon. How to turn it on and off, what the kernel judges per tool call, the admission ceiling, what changes on your machine, and how to verify it on your instance.
Requires 0.197.1 or later; the per-tier default, the declared narrowing and the stop at a step boundary need 0.200.
What it is
A work loop, the thing that runs when an actor is given a goal and tools, can run in one of two places.
in_processruns it inside the daemon, sharing its memory with the provider credential, the operator's bearer and the store. It is the operator's default on their own machine.sandboxedruns it as a separate process with no network access, no view of the instance's home directory, a restricted set of system calls, limits on memory, processes, files and CPU time, one short-lived credential handed to it on standard input, and five reachable doors on a local socket. Everything else answers 404.
The daemon checks each isolation layer before the run is allowed to begin. A run it cannot isolate is refused rather than started, and it never receives a credential.
The transport works on hosted instances and on your own machine alike. On a hosted instance the daemon runs each worker as a separate unprivileged user.
Which runs are sandboxed
The transport is decided per run, at start, by the tier of the principal the run is for:
| principal | default | config key |
|---|---|---|
| the operator | in_process | work_loop_transport |
| every other tier (Trial, Paid, Team) | sandboxed | work_loop_transport_members |
Both keys are one config record on th_engine_config, live for the next run to start. There is no restart and no release in either direction, which is the point: a transport problem in production has to be one write away from off.
# The operator's own runs, sandboxed too
spl know --thread th_engine_config --actor did:sync:system:engine --act LEARN \
--body '{"kind":"core.engine.config.v1","topic":"work_loop_transport","transport":"sandboxed"}'
# The rollback for everyone else: members back in process (not recommended on a hosted instance)
spl know --thread th_engine_config --actor did:sync:system:engine --act LEARN \
--body '{"kind":"core.engine.config.v1","topic":"work_loop_transport_members","transport":"in_process"}'Check what the running service took, not only what was written:
curl -s -H "Authorization: Bearer $ADMIN" localhost:9100/v1/engine/health | jq .work_loop{ "transport": "in_process", "member_transport": "sandboxed", "admission_capacity": 16, "admission_available": 16 }An unrecognised value for the operator's key reads as in_process, and for the members' key as sandboxed: a typo moves nobody into a transport they were not already in. Before starting a run, POST /v1/work/loop/preview reports the transport a start would take for that principal.
Runs already in flight keep the transport they started under. Flipping a value stops new runs taking the old transport; it does not reach into one already running. To stop that one:
# At the next step boundary: the run's next tool call is answered with a stop verdict
# and the run ends with a cancelled outcome; its worker process is reaped.
curl -s -X POST -H "Authorization: Bearer $ADMIN" -H 'content-type: application/json' \
localhost:9100/v1/runs/$THREAD/events -d '{"type":"stop"}'What a tier's grant the sandbox cannot serve
The instance's built-in file tools (read_file, write_file, list_dir, search) are built from handles a separate process does not have. A sandboxed run whose tier grants them starts anyway: the four are withheld from its reach and the run says so. The run's anchor and its row at GET /v1/runs/{thread} carry narrowed, the list of granted tools the transport withheld, and the model is told the same in its opening. Nothing is dropped silently.
Every tool call is judged by the kernel
A sandboxed worker holds no permission rules. Before it runs any tool it proposes the call to the kernel, which evaluates the same guard the in-process loop would have used, over the live rules, and writes its judgment as a core.work.tool_verdict.v1 record before answering. Three layers answer in order:
| layer | meaning | what to change |
|---|---|---|
stop | the run was stopped; this call is refused and the run ends at this boundary | nothing, it is what you asked for |
grant | the tool is not in this run's reach | the caller's tool grant, or the tier |
clamp | the loop contract's hard forbidden_tools | the run's own request |
rule | a live permission rule denied it | spl config add-permission-rule |
allow | nothing objected | nothing |
Because the rules are read per call, a rule written while a run is going is honoured by that run's next tool call, with no restart:
spl config add-permission-rule --name stop-that --action deny --priority 100 \
--expression 'resource == "tool:bash"'The verdict is authored by the kernel, and so is the run's outcome. The run credential cannot write either kind, so a run cannot declare that it was allowed to act or that it succeeded. The outcome's success is also checked against what the tools did: a run whose every call failed is not a success whatever its final answer says, and the outcome carries tool_calls_ok and tool_calls_failed so a reader has the evidence beside the verdict.
After a call, not just before it
The permission rules above run before a tool call. After one returns, the kernel may add an observation: something the environment noticed about the run, put in front of the model on its next turn beside the tool's own result, never in place of it.
Three of them ship today. An actor calling the same tool with identical arguments three times in a row is told so, and told to change something or stop; the count is kept after execution, so a call the permission plane denied counts too, because an actor hammering a refused call is exactly the loop worth breaking. An actor whose second read of the same thread comes back empty is told that it has already looked, rather than being handed the same sentence twice. And a tool call arriving on a reply that hit the output limit is never executed at all: arguments cut off mid-write can parse cleanly and still be incomplete.
Each one is a core.work.observation.v1 record on the run's thread naming the
rule that fired and carrying the exact text the actor was shown. Hooks in
other agent harnesses are configuration files and leave nothing behind, so
nobody can ask afterwards why a run changed course. Here you can:
spl thread records $THREAD -o json | jq -r '
.data[] | select(.body.kind == "core.work.observation.v1") |
"\(.body.rule): \(.body.text)"'Observations compose forbid-wins: a rule may only add one, never suppress another, so no ordering of rules can quietly disarm the system. And they work on both transports, because the hooks are attached where the loop is assembled, which the worker shares.
To read what a run was allowed to do, walk its thread:
spl thread records $THREAD -o json | jq -r '
.data[] | select(.body.kind | test("tool_(call|verdict|result)")) |
"\(.sequence) \(.body.kind) \(.body.tool // .body.fulfills) \(.body.decision // .body.status // "")"'The admission ceiling
Concurrent sandboxed runs are bounded by host memory. The slot is taken before a worker is started, so a saturated instance never starts a process it then has to refuse. A start past the ceiling answers HTTP 503 with the code ADMISSION_CEILING, and /v1/engine/health reports the ceiling and how much of it is free.
What changes on your machine
- A second socket. The daemon binds
~/.syncro/run/spl-worker.sockbeside the operator socket. The operator socket still grants same-user trust; the worker socket requires a bearer on every request. - The daemon's process is protected. A same-user shell cannot read the daemon's memory or environment.
readlink /proc/<pid>/exeneedssudo, and the daemon produces no core dump. This is the hardening working. - Workers are found from the log, not from the process table. The daemon logs each worker's process id when it starts it. Do not use
pgrep -f 'spl loop-worker'; it matches the shell that issued it.
journalctl --user -u spl-serve -n 200 | grep 'running sandboxed'- Upgrades are stop, install, start. A worker is started from the daemon's own binary, so replacing the binary under a running daemon does not change already-running workers, and a live copy over the running binary is refused (
Text file busy). See In-place upgrades.
What a run is confined by
Precisely, so nobody reads more into it than is there: a user namespace (on a hosted instance, a separate unprivileged user instead), a network namespace with no interfaces, a mount namespace in which the instance's home is an empty read-only filesystem with only the worker socket rebound into it, a system-call filter that refuses namespace, mount, tracing and module calls, resource limits on address space, file size, processes and CPU time, an environment holding nothing but a path, a credential bounded to the run's own thread and to the kinds a run may author, and the narrowed door set. There is no process-id namespace: a worker can see other processes exist and cannot touch them.
Every release runs the escape suite: from inside a real sandboxed run of a non-operator actor, on the same-user shape and on the root-daemon shape, it tries each of those things and records the errno it got. A confinement layer that stops working turns its rows red before a release is tagged.
Verify it on your instance
Start one small run through the work loop door as a non-operator, or as the operator with work_loop_transport set to sandboxed, and watch it finish. The goal records a note and stops.
curl -s -X POST -H "Authorization: Bearer $ADMIN" -H 'content-type: application/json' \
localhost:9100/v1/work/loop \
-d '{"goal":"Record a note whose text is exactly: transport-check . Use the record_note tool with only the text argument. Then stop.",
"max_turns":3,"budget_usd":0.05}'If you leave budget_usd out, a sandboxed run is minted with its tier's per-run default ceiling, because the inference broker refuses a credential with no ceiling; the run row's budget_usd shows the ceiling the credential actually holds. An in-process run with no budget stays unbounded, as before.
The response carries the run's thread. Poll it until it is finished:
curl -s -H "Authorization: Bearer $ADMIN" localhost:9100/v1/runs/$THREAD | jq .state
# "done"Then read the outcome record on the thread and check three things: transport is sandboxed, tool_calls_ok is at least 1, and the note is there. On an instance with a data plane the run row also carries narrowed naming the file tools the transport withheld.
spl thread records $THREAD -o json | jq '.data[] | select(.body.kind == "core.work.loop_outcome.v1" or .body.kind == "core.work.note.v1") | .body'If the outcome says stopped_short and its summary names the worker socket or the inference relay, the instance is on a version before 0.197.1.
What is on by default
Every run for a principal other than the operator is sandboxed. The operator's own runs are in process unless work_loop_transport says otherwise. Both are per instance and both are one config record to move.
Host tools (bash and the filesystem set) run on the operator's own machine only, where the operator is already the trust boundary; no transport setting extends them to anyone else.
See also
- Security model for protection domains and the credential ordering
- Operator runbook for the upgrade sequence
Security model — secrets at rest, threat model, operator discipline
What spl serve protects against (default-secure auth + filesystem permissions) and what it does NOT (host compromise, cloud-sync of identity dirs). Read before deploying to a multi-tenant or shared-storage environment.
Account sign-in for your instance
Let people who hold a hosted account come back to your instance by signing in, by declaring which verifier your instance trusts. Hosted instances get this from provisioning; self-hosted instances opt in.