THE SIGNAL IN ONE SENTENCE

A model does not run code by thinking harder. It asks something else to run the code. That distinction sounds fussy until the model writes a command you did not expect, opens a file you forgot was attached or follows an instruction hidden inside a document. Then the location of the execution environment becomes the whole story. OpenRouter published a comparison of server-side code execution tools on October 5. The article places its own beta shell beside hosted execution from OpenAI, Anthropic and Google, and contrasts all four with a sandbox a developer operates independently. The basic trade is easy to understand. With a client-side tool, an application receives the model's tool call, runs a handler and sends the result back. The developer owns the runtime, the image, the network rules, the lifecycle and the security boundary. With a server-side tool, the model provider or gateway intercepts the tool call, runs it in a hosted container and returns the result inside the same request. The application may need only one entry in its tools array. That is a real reduction in plumbing. It is not a transfer of judgment. OpenRouter introduced its shell tool and Files API on September 8. Its documentation says the shell works with any tool-calling model on the Responses and Messages APIs. A related bash-shaped tool works on the Messages API. Setting the engine to OpenRouter sends commands into the provider's isolated container instead of asking the customer's application to execute them. The release is still beta. The provider says each container is scoped to the customer's account and workspace. Outbound network access is off by default. A developer can allow named domains, or allow all domains, before the container starts. The policy cannot be changed after startup. That default matters. An agent that can read a malicious spreadsheet but cannot call an arbitrary internet host has fewer ways to leak the spreadsheet. An agent that can run a package installer only against a named package registry has a narrower path than one with unrestricted outbound access. Narrower does not mean harmless. An allowed domain can still serve compromised software. A package can still contain dangerous setup behavior. A model can still delete files inside its own workspace, consume its command budget, produce a misleading analysis or return sensitive material in ordinary output. Isolation limits the blast radius. It does not make every command wise. OpenRouter documents hard command boundaries. A command defaults to a 120-second timeout and cannot exceed 300 seconds. Output defaults to 16,384 characters for each stream and can be raised to 65,536. A shell call with more than 100 commands is rejected. The wider server-tool loop defaults to and tops out at 30 tool calls per request. Those are useful circuit breakers. They are not a business policy. A 300-second command can still process the wrong customer file. Thirty permitted tool calls can still repeat the same bad assumption. A neatly bounded error is easier to contain than an unbounded one, but the user still receives the wrong result unless the application checks the work. The container lifecycle has similar tradeoffs. OpenRouter says a conversation receives a fresh container by default. A session identifier or explicit container reference can reuse one. The container sleeps after five idle minutes. Files inside the home directory are captured, and the Files API can move inputs and outputs between workspace storage and the container. Container files are kept for 30 days unless promoted into longer-lived workspace storage. Persistence is useful because an agent can inspect a file, write a script, run it, repair the script after an error and preserve the resulting artifact. Persistence is also state. If a team reuses a container across customers, tasks or trust levels, yesterday's files can become today's accidental context. The safe question is not only whether the provider isolates one customer from another. It is whether the application isolates one job from the next. The plain signal is that hosted execution moves the machine room into the API, while responsibility stays split across the provider, the application developer and the organization deploying the agent. The provider owns the container boundary it advertises. That includes isolation, patching, lifecycle enforcement, metering, documented limits and the behavior of the hosted tool. The application developer owns what the model may send into that boundary. That includes attached files, prompts, secrets, container reuse, network destinations, tool-call limits, output handling and the conditions under which a person must approve a step. The deploying organization owns whether the workflow belongs there at all. A convenient beta service is not automatically the right place for patient records, export-controlled designs, unreleased financial data, source code carrying production secrets or evidence governed by strict retention rules. OpenRouter's October 5 comparison makes that limit explicit in practical terms. It says a self-managed sandbox remains the better choice when a job needs a custom base image, GPU access or a session that runs for hours. Add several more reasons to that list. Use a system you control when you need a particular patch schedule, a private package mirror, a customer-specific encryption boundary, a formal evidence-retention policy, a guaranteed geographic processing location, specialized monitoring or a regulator-approved configuration. OpenRouter's own queue warning says its in-region endpoints do not support the shell and bash tools. That matters for teams whose data handling depends on a stated region rather than a general promise of isolation. The low-friction cases are smaller. A support agent can inspect a non-sensitive CSV and calculate totals. A research assistant can turn public data into a chart. A developer tool can run a short parser against a synthetic fixture. An operations assistant can analyze scrubbed logs. A publishing workflow can convert a harmless file and return the artifact. In each case, the workload is short, the inputs are bounded, the desired output can be checked and the tool does not need a standing credential to production. That last point deserves its own rule. Do not place a powerful reusable secret inside a model-controlled runtime merely because the runtime is called a sandbox. Give the job a narrow, temporary capability instead. A one-use upload URL is better than a storage master key. A read-only database view is better than production credentials. A package-registry allowlist is better than the entire internet. A fresh container is better than a shared workbench when jobs should not know about one another. The same principle applies to files. Attach only what the task needs. Strip hidden metadata when it is irrelevant. Separate customer workspaces. Name retention periods. Delete or expire artifacts on purpose. Log which files entered the run and which outputs left it. Then test the model and the tool together. OpenRouter reports running the same small Python calculation through six different models with one shell definition. All six returned the correct result. That demonstrates portability for that particular command, not universal reliability. A model can format a tool call incorrectly, choose the wrong command, misunderstand a file, ignore an error code or confidently summarize partial output. Model portability is valuable because teams can compare behavior without rebuilding the runtime. It also makes evaluation more important because the same tool surface can hide very different decision quality. Build a fixed test set before sending real work through the sandbox. Include malformed files, prompt injection inside documents, oversized outputs, timeouts, missing dependencies, denied network calls, confusing filenames, stale container state and commands that should require approval. Record the model snapshot, tool configuration, attached files, network policy, container identifier, commands, exit codes, output and final answer. That record turns a black box into a reviewable sequence. Cost needs the same clarity. OpenRouter bills sandbox time at $0.0001 per active second. A new or sleeping container carries a 30-second minimum. The request bill combines model tokens with sandbox time, and the provider shows the sandbox portion in its generation timeline. Files API usage has no separate charge, with total storage limited to 10 GiB. That looks tiny for one run. At scale, the operational question is not whether one second is cheap. It is whether the agent starts unnecessary containers, repeats failed commands, keeps choosing an expensive model for trivial execution or performs work a deterministic service could do faster. Measure cost per successful task, not cost per command. Also measure correction rate. If a person must repair one in five artifacts, a low sandbox bill can hide a costly workflow. Hosted execution will keep spreading because it turns a difficult infrastructure problem into an API feature. That is good news for prototypes and for many bounded production tasks. The responsible version comes with a map. Write down which boundary the provider guarantees. Write down which inputs the application permits. Write down which actions require a human. Write down where files persist, how network access is granted, which logs survive and who can investigate an incident. Then choose the smallest runtime that can finish the job. The sandbox may have moved server-side. The responsibility did not move into the cloud by itself.

01

WHAT ACTUALLY CHANGED

OpenRouter published an October 5 comparison of its hosted shell with server-side code execution from OpenAI, Anthropic and Google, plus self-managed sandbox infrastructure.

OpenRouter's beta shell can run commands for any tool-calling model on the Responses and Messages APIs, while its bash-shaped tool works on the Messages API.

The provider documents isolated containers, network access off by default, per-command time and output limits, short idle lifetimes, reusable sessions and file transfer through its Files API.

Hosted execution removes container provisioning and tool-loop plumbing from an application, but still leaves the application responsible for inputs, permissions, retention, evaluation and approval policy.

02

WHY THIS MATTERS

Moving execution into a provider sandbox can reduce infrastructure work and keep model-written commands away from production systems, but it creates a new dependency on the provider's isolation, retention and regional controls.

A sandbox limits where side effects occur. It does not determine whether a command is appropriate, whether an attached file should be processed or whether the model's conclusion is correct.

Model-agnostic execution makes it easier to compare models on one tool surface, while also requiring teams to test tool-call accuracy and recovery behavior for each model they may route to.

Regulated and sensitive workloads need an explicit responsibility map because a beta hosted runtime may not satisfy custom images, GPU needs, long sessions, data residency, evidence retention or approved monitoring requirements.

FIG. 316THE HOSTED EXECUTION RESPONSIBILITY MAP
1CLASSIFY THE INPUT→
2CHOOSE A FRESH OR REUSED CONTAINER→
3SET NETWORK AND FILE BOUNDARIES→
4CAP COMMANDS TIME AND OUTPUT→
5REQUIRE APPROVAL FOR HIGH IMPACT ACTIONS→
6CHECK THE ARTIFACT AND EXIT CODES→
7RECORD COST STATE AND RETENTION→
8ESCALATE OR MOVE TO A CONTROLLED RUNTIME
A hosted sandbox removes infrastructure work only after the team has defined the boundaries, approvals and evidence that make the run safe enough to trust.

03

WHERE IT COULD HELP

  • Analyze scrubbed CSV files, public datasets and non-sensitive logs with short scripts whose outputs can be checked independently.
  • Generate charts, reports and transformed artifacts without exposing the developer's laptop or production servers to model-written commands.
  • Run parser checks, reproducible examples and synthetic test fixtures in a fresh container before code reaches a normal build system.
  • Compare several models against the same tool definition, command budget and evaluation set before selecting one for an agent workflow.
  • Use narrow network allowlists, one-use credentials, per-job containers and durable execution logs for any production workflow that adopts hosted execution.

KEEP A HAND ON THE WHEEL

OpenRouter is the vendor of the shell being compared, so its October 5 article is a useful primary description and a commercial comparison, not an independent security audit. Shell, bash, Files and containers remain beta. Documented isolation and default network denial reduce risk but do not establish suitability for every regulated workload. Watch for third-party penetration testing, published incident handling, clearer data-region support, retention controls, service guarantees, independent cost comparisons and evidence that model portability holds on messy real tasks rather than one-line calculations.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 5, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US