OpenShell explained: where an AI agent is allowed to act
OpenShell performance and cost: measure the boundary you use
Build a benchmark that separates startup, mediation and provider charges
What you will learn
- Time the lifecycle
- Measure mediated work
- Account for external cost
Before you start
- One disposable workload and a supported compute runtime
- Authority to inspect policy and provider configuration
Turn the documented controls into a small reviewable operational report
Key takeaways
- No speed number is claimed without a run.
- Mediation and startup are different performance questions.
- Provider charges sit outside the runtime license.
Time the lifecycle
Measure gateway readiness, sandbox creation, first command and clean deletion on your chosen driver. A Docker socket, Kubernetes service and VM vsock have different setup costs, so combine their times only when the workload and environment match.
Record warm and cold starts and an explicit failure case. This series does not report latency or throughput figures because the upstream runtime was not executed.
Measure mediated work
Use a fixed set of allowed and denied DNS, TCP and file operations. Record network request latency, successful payload size and behavior when the supervisor connection drops.
The architecture describes separate streams with backpressure, which suggests what to measure; it does not establish a universal overhead percentage. Compare against a suitable baseline on the same host and policy.
Account for external cost
A local gateway still consumes host or cluster resources. Providers can add metered model inference, storage and outbound traffic. A Kubernetes deployment adds operational work to keep images, certificates and NetworkPolicy aligned.
Separate infrastructure cost from model tokens and from human policy-review time. The cheapest benchmark run can be misleading if it omits incident handling and provider revocation.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Measure cold and warm sandbox lifecycle separately.
- 2
Run the same allowed and denied operations with timestamps.
- 3
Record infrastructure, provider and review costs in different columns.
Copy-ready example
driver,gateway_release,policy_hash,cold_start_s,dns_allowed_ms,dns_denied_ms,provider_cost,notes
,,,,,,,not-runFrequently asked questions
Is there a published overhead percentage to rely on here?
This series does not validate one; measure your exact runtime, policy and workload.
Does a local gateway eliminate model costs?
No. Attached inference providers may still charge for requests.
Sources
- OpenShell / docs/about/architecture.mdxSource checked 2026-10-04
- OpenShell / docs/about/support-matrix.mdxSource checked 2026-10-04
- OpenShell / docs/how-it-works/providers/overview.mdxSource checked 2026-10-04