Install
Inside DeepSeek Harness, with dsh-market
dsh plugin --profile web add dshmarket
Or from the command line
dsh plugin --profile web add github:sandbaseai/sandbase-harness
Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.
README
AI-readable project metadata: llms.txt · installation guide
A local-first runtime for AI agents. Sessions, sandboxed tools, memory, credentials, audit trails, and a built-in Console — all running on your machine or in your own infrastructure.
Why
Agent SDKs handle the model loop. Production agents need more: persistent
sessions, tool governance, sandbox boundaries, credential handling, memory,
auditability, and a UI for humans to inspect what happened. managed-agents
is that runtime layer — not a visual workflow builder and not another model SDK.
| Need | What Harness provides |
|---|---|
| Run generated code safely | Local, Docker, Kubernetes, and self-hosted worker sandboxes |
| Inspect long-running agents | Persistent sessions, resumable event streams, audit, and replay |
| Control tool access | MCP toolsets, credential vaults, permission policies, and approvals |
| Operate any model | OpenAI, Anthropic, MiniMax, and OpenAI-compatible providers, including DeepSeek V4 |
| Keep infrastructure yours | Local-first SQLite and file storage with no required hosted control plane |
Features
- Claude Managed Agents-style
/v1API and local Console - SQLite metadata by default for agents, sessions, environments, credential vaults, memory stores, files, skills, and API keys — local file/skill bytes in the workspace state directory
- Resumable Server-Sent Events for session replay and debugging
- One active model provider boundary configured through Settings V2
- Sandbox backends: local process, Docker (per-session containers), Kubernetes (kubectl exec/cp), self-hosted worker queue
- MCP toolsets, permission policies, built-in tools, and skill packages
- TypeScript SDK at
managed-agents/sdk - Release gate:
npm run release:check
Quick Start
Requirements: Node.js 22+, npm 10+, and a model provider API key (OpenAI, Anthropic, MiniMax, or any OpenAI-compatible endpoint). Docker is optional and only needed for Docker-backed sandboxes.
git clone --branch v0.3.8 --depth 1 https://github.com/sandbaseai/sandbase-harness.git
cd sandbase-harness
npm ci
npm run build
mkdir ../my-agents && cd ../my-agents
node ../sandbase-harness/dist/index.js init
node ../sandbase-harness/dist/index.js start
init writes a workspace into the directory you run it from: an agent, a skills
folder, and config.yaml, whose provider reference is the ${OPENAI_API_KEY}
environment variable. start serves the API and the Console on
http://127.0.0.1:3000.
Two steps finish the setup, both on Settings > Setup at http://127.0.0.1:3000/dashboard:
- The provider. Paste your API key into the provider form and save. The page
then reports that the saved configuration is not active yet, so restart the
runtime — stop it with Ctrl+C and run the
startcommand again, or use the restart button. A saved setting only takes effect at startup. If your provider is not in the list, choose the OpenAI-compatible vendor and set its base URL. - The model. In the Agent models panel, set the model ID your provider
actually serves —
deepseek-chatfor DeepSeek, for example. An agent carries its own model ID, so thegpt-4othatinitwrites is not valid for every provider, and a wrong ID fails the turn withmodel_not_found.
Send the first message from the Console: open Sessions, create a session for the agent, and type into the composer. From a terminal it is one command:
node ../sandbase-harness/dist/index.js chat agent_assistant --message "hello" --tool-approval allow
chat sends that one message and exits once the turn settles; without
--message it keeps the session open and streams until you interrupt it.
--tool-approval allow
preauthorizes the tool calls the agent may make, which the init template
otherwise parks for approval and waits for a person to answer; see CLI.
The unscoped managed-agents name on npm is not this project. Until an
official scoped package is announced in this repository, install only from the
tagged GitHub source release shown above. Do not run npx managed-agents or
npm install managed-agents.
Try it in Codespaces
The included development container installs dependencies and builds the runtime. When the terminal is ready, start the server on the forwarded port:
node dist/index.js start --host 0.0.0.0
Open the forwarded SandBase Harness Console port, then configure a model in Settings > Setup. Codespaces usage may be billed by GitHub; the local quick start above remains free and keeps all runtime data on your machine.
Screenshots
| Console overview | Settings | API reference |
|---|---|---|
![]() |
![]() |
![]() |
Use the Official SDK
The runtime answers its own /v1 API on that same port, and an official
Anthropic TypeScript SDK client drives it unchanged: point the client's
baseURL at the runtime, give it the runtime API key, and the quickstart in
examples/official-sdk runs a whole turn —
message, tool call, tool result, final reply — against it. That example is
executed on every pull request by
tests/conformance/official-sdk-quickstart.test.ts, so the compatibility it
describes is compatibility that is tested rather than claimed.
The same surface is specified in docs/api.md, and this repository's own TypeScript SDK is documented under SDK below.
CMA compatibility
Coverage of the published Claude Managed Agents contract is declared entry by
entry in src/core/capabilities/matrix.ts:
of the official SDK's route surface, 76 routes are mounted, 29 refuse by name,
and 5 — the multi-agent thread surface — are deferred to a tracked issue.
Partial and Unsupported entries always name their reason.
The table below is generated by npm run docs:compat, and a contract-honesty
test fails when it drifts from the matrix.
| Area | Official capability | Status | Notes |
|---|---|---|---|
| headers | compatibility-header-admission |
Partial | Version, beta, and mutual-exclusion rules are enforced for any request that carries a compatibility header. A request with no compatibility header is accepted as a local caller, which the published contract does not define; this header-free path is a deliberate local-first extension for a self-hosted single-tenant runtime, recorded as such in the headers contract §4, and it is why this entry stays partial. |
| headers | extension-namespace-exclusion |
Supported | |
| pagination | opaque-cursors |
Partial | A collection's envelope follows the mount: the operations router serves /v1 with {data, prev_page, next_page} and its /v1/x mirror with the local {data, has_more, first_id, last_id}, chosen through one pager so no handler emits both spellings, and cursors that are readable base64url JSON rather than opaque binary. Every canonical /v1 collection serves the canonical envelope, with no exceptions: the complete-set listings carry {data, prev_page: null, next_page: null} and the windowed ones carry a followable cursor — /v1/sessions pages by number under {order, created_at bounds, page} and rejects a cursor replayed under another ordering or creation window, /v1/sessions/{id}/events carries {order, filter, after_id}, /v1/skills and the credential audit listings use {offset, filter} — so a cut page says so instead of looking complete. The one listing that is neither shape is /v1/environments/{id}/work-items, a windowed extension that adds a counts object and is named in the contract. |
| pagination | cursor-query-binding |
Supported | |
| errors | structured-error-envelope |
Supported | |
| agents | agent-crud |
Supported | |
| agents | model-object-profile |
Partial | String and object model forms parse field by field. effort and speed are stored, returned by the read projection (the agent read, the version read, and the session snapshot), and executed on the Anthropic provider under a model capability table — effort becomes output_config.effort, fast becomes speed: "fast" with the fast-mode beta, and adaptive-thinking models receive thinking: {type: "adaptive", display: "omitted"}; a listed model refused a level or speed it cannot take fails admission, an unknown model id or non-Anthropic provider sends nothing, and a deployment's own reasoning_effort model setting is operator-level and separate. inference_geo is refused by name with unsupported_model_field because this runtime has no inference-geography control; and a canonical multiagent roster is refused by name rather than executed. |
| agents | multiagent-roster |
Unsupported | A canonical multiagent roster is refused by name on both agent create and agent update, because no thread, coordinator, or advisor surface exists to honour it; accepting it would let a caller believe delegation by roster is in effect. Local delegation is registered separately as an extension. |
| agents | local-delegation-subagent |
Supported | |
| sessions | session-lifecycle |
Supported | |
| sessions | initial-events |
Supported | |
| sessions | prompt-caching |
Supported | |
| sessions | session-update |
Partial | POST /v1/sessions/{id} applies agent limited to tools/mcp_servers (merged onto the resolved definition, validated like creation, and materialized as agent_definition without touching the agent row), a metadata merge patch (null per key removes, null field is no change), a title replace (null clears), and a budget move under the budget contract's rules (budget_create_only, budget_not_raised, model_not_budgetable, budget_invalid_*). An agent change requires an externally idle session (session_not_idle while running); title, metadata, and budget move in any non-terminal state, and a terminated or archived session is session_terminated. One session.updated event carries only the changed fields — the full agent snapshot, the new ceiling or null, the whole post-update metadata bag, the new title — and a no-op emits none. vault_ids is refused with vault_ids_not_updatable, which is why the capability is partial rather than supported. |
| budget | session-budget |
Partial | Consumption is priced in integer microcents from the append-only log, and a session may declare a max_list_cost ceiling at creation. The builtin loop checks the ceiling inside a turn: the step that crossed the cap is the last one, the session idles with stop_reason budget_reached and a session.usage immediately before it, and an accepted budget update or removal resumes the session on its own — a tool call the ceiling stranded is settled so the resumed turn sees a paired transcript. At the ceiling the next work-starting event is refused with budget_reached while events that settle work already in flight are still accepted, so the next model request does not start; a declared outcome's revision loop reads the same spend and stops at the ceiling too, closing the outcome with result budget_reached rather than starting another grading pass or turn, because the loop's turns are internal to an event that was already admitted. Two deviations are deliberate: prices come from an operator-supplied cost profile rather than official list prices, so a session whose model the profile cannot price is refused a budget and usage.list_cost is withheld while any used model is unpriced; and the pause is reported on the session's own status_idle only, because the published thread-level budget_reached signal belongs to the thread surface, which this runtime does not implement. |
| events | append-only-event-log |
Supported | |
| events | processed-at-lifecycle |
Supported | |
| events | session-error-structure |
Supported | |
| events | error-enum-completeness |
Supported | |
| events | model-request-span-pair |
Supported | |
| streaming | resumable-sse |
Supported | |
| streaming | agent-message-stream-preview |
Supported | |
| tools | builtin-tool-execution |
Partial | File, shell, search, and web_fetch tools execute; web_search accepts configuration but has no search provider and fails admission before execution. |
| tools | web-fetch-execution |
Partial | WebFetch executes over HTTP/HTTPS with domain policy, per-redirect revalidation, private-address rejection, timeout and byte caps, HTML text extraction, and a max_content_tokens budget; it converts text-like content only (no image or PDF rendering), the token budget is a character estimate, and TLS hostnames are verified but content is not sandboxed beyond redaction. |
| tools | web-tool-domain-policy |
Supported | |
| tools | tool-output-overflow |
Partial | Overflow has one unified contract (spill path, preview, marker, retrieval), but the local threshold is 50,000 chars rather than the published 100,000. |
| tools | mcp-tool-approval-gate |
Supported | |
| custom-tools | custom-tool-declaration |
Supported | |
| system-message | system-message-events |
Supported | |
| memory-stores | memory-crud |
Supported | |
| memory-stores | memory-limits-and-preconditions |
Supported | |
| memory-stores | memory-version-audit |
Supported | |
| memory-stores | memory-multi-mount |
Supported | |
| github-repository | github-repository-materialization |
Supported | |
| github-repository | github-repository-identity-freeze |
Supported | |
| files | file-resources |
Supported | |
| files | file-mount-path |
Supported | |
| credentials | canonical-credential-wire-profile |
Supported | |
| credentials | credential-rotation |
Supported | |
| credentials | credential-injection-execution |
Partial | A turn on a session that attaches a vault injects its unrestricted environment variables as plaintext into the sandbox command environment and into any stdio MCP server the agent declares (a vault value wins over the value the agent configured), hands a url-transport server the credentials scoped to its own mcp_server_url (a static_bearer or mcp_oauth credential is attached only to the endpoint it names, on the SSE request and on every message POST), redacts every value a sandbox tool hands back and every value an MCP tool returns, and clears the retained values when the turn ends. The runtime composition supplies the resolver, so a CLI-started runtime resolves the session vault while an embedder that omits it runs sessions with no vault. Three deviations: the secret itself enters the child process environment, with no opaque placeholder and no substitution at the network egress, so any command the agent runs can read it and send it out — the published model keeps the credential out of the process and replaces a placeholder on the outbound request; the delegated child path builds its own sandbox tools and does not thread credentials, so a sub-agent receives no vault environment; and nothing is injected into model requests, so a credential authenticates an outbound call rather than a completion. |
| credentials | oauth-refresh |
Unsupported | No refresh loop or refresh-failure event exists, and none is scheduled. The official MCP OAuth validation endpoint explicitly refuses the capability with unsupported_capability. A supplied refresh block is parsed, stored, and answered with an explicit warning that it will not be executed, so a caller never assumes a token was renewed. |
| operations | webhook-subscriptions |
Partial | Locally implemented, but the delivery behaviour is not the published contract. Subscriptions are managed over REST under /v1/webhooks (with the /v1/x mirror) and delivery runs from a bridge the runtime composes at startup: each durable event is projected as it is broadcast and a 60-second tick retries due deliveries and runs due deployments, while POST /v1/webhooks/dispatch and POST /v1/webhooks/retry-due remain for on-demand passes. Every attempt carries the published header names and a Standard Webhooks v1 signature over id.timestamp.body — a retry keeps the event id and signs with its own timestamp, and each subscription holds its own whsec_ secret that is returned once at creation, and a rotation window keeps the previous secret valid in a second webhook-signature entry until it is retired — the payload is the published {type: "event", id, created_at, data: {type, id, organization_id, workspace_id}} reference envelope with webhook-id equal to the event id and local constant org/workspace values, subscriptions may only name events from the official catalog — , prefix., and unknown names are refused at write time — and the session stream reaches subscribers only through the published-name projection (status events mapped, budget_reached deduplicated per session and ceiling, internal events dropped) while resource routes publish the lifecycle events for sessions, agents, environments, vaults and credentials, memory stores, and deployments; the catalog names with no producing surface (session.pending/running/idled/requires_action, session.thread_*, agent.deleted, vault_credential.refresh_failed, deployment.deleted) stay subscribable-but-silent, nothing retires the previous secret automatically, of the three published auto-disable cases all three exist, two unconditionally and one opt-in (an attempt that observes a redirect disables the endpoint with the published disabled_reason and is never retried; an attempt whose host is an internal name or resolves to a private address is refused before any connection with its own published reason but only when the deployment sets MANAGED_AGENTS_WEBHOOK_SCREEN_PRIVATE_ADDRESSES, off by default because loopback is private and a self-hosted receiver normally shares the host; and an endpoint failing without interruption for at least a window is disabled with the published sustained-failure reason, where the contract publishes the trigger shape — duration, not attempt count, with a 2xx resetting it — but no length, so the window is a local parameter: it defaults to 10 minutes, a deployment sets its own with MANAGED_AGENTS_WEBHOOK_SUSTAINED_FAILURE_WINDOW_SECONDS, and the runtime records the window in force at startup because a deployment variable has no write path of its own). Retries follow the published jittered 5-120s exponential backoff: the ceiling doubles from 60s to 120s and each delay is drawn uniformly between 5s and that ceiling. |
| operations | scheduled-deployment-timers |
Partial | The published deployment surface end to end: the object answers with type deployment, a depl_ id, a pinned {type, id, version} agent, environment_id, a required non-empty initial_events list (the session admission plus the deployment-only system.message), resources, vault_ids, budget, metadata, and a schedule object with last_run_at and the next three upcoming_runs_at — null for a manual-only deployment, because cron is nullable since M052. Both mount spellings serve create/read/update (POST the published verb, PUT the local one)/archive/pause/unpause/run/run-due from one router, and the local flat aliases (agent_id, cron, timezone, payload) remain accepted. Each run creates its session through createWithInitialEvents and records a drun_ run readable at /v1/deployment_runs (deployment_id, has_error, trigger_type, created_at filters) with trigger_context carrying scheduled_at for a timed run. Failure is asymmetric per the published contract: a missing or archived bound agent archives the deployment with no run, a recoverable session_rate_limited_error records a failed run only, and other classified failures record the run and auto-pause the deployment with paused_reason.error mirroring the run's classified error.type. Manual pause and unpause exist and a paused deployment still accepts a manual run; timed runs publish deployment_run.started/.succeeded/.failed and manual runs publish none, while deployment.created/.updated/.paused/.unpaused/.archived publish on their transitions including the agent-gone cascade. The runtime's 60-second tick runs due deployments and startup re-arms their forward schedule without replaying a missed trigger. Remaining gaps: deployment.deleted has no producer because no delete route exists, and mcp_egress_blocked_error has no producing path because MCP egress is not gated. |
| operations | outcome-grading |
Supported | |
| operations | outcome-evaluation |
Supported | |
| capabilities | capability-inventory-endpoint |
Supported | |
| capabilities | capability-status-truthfulness |
Supported | |
| environments | environment-hosting-config |
Supported | |
| environments | environment-network-policy |
Partial | The published config.networking object is accepted in its own vocabulary (limited/unrestricted, allowed_hosts, allow_mcp_servers, allow_package_managers) and normalized into the recorded local config.network spelling by one normalizer, with fail-closed defaults for an unrecognized type and for an unset permission, and a request declaring both spellings inconsistently refused. It is partial because nothing enforces it: no sandbox provider shipped in this runtime reads an environment network policy, so a declared limited policy with an empty allowed_hosts grants the same egress as unrestricted, which the contract file, docs/api.md, and the Console API reference state plainly. The status becomes supported when a provider applies the policy it is given. |
| routes | documented-route-surface |
Supported | |
| unsupported | dreams |
Unsupported | Dreams are a memory-consolidation pipeline: they read memory stores and historical sessions and produce new, reorganized stores. This phase deliberately does not implement it; official SDK routes explicitly refuse it with unsupported_capability. Unavailable rather than not_applicable because the feature belongs in a local-first runtime — what it needs is a scheduled background worker and archived-session corpora, not a hosted service. |
| threads | threads-and-coordinator |
Unsupported | Not implemented: there is no thread resource, no thread lifecycle or per-thread event isolation, no coordinator or advisor role, no /threads route, and no thread-scoped budget event. A request carrying a multiagent roster is refused by name rather than silently stripped. Delegation exists only as the local single-level delegations / enable_general_subagent extension, which is not this surface. |
| unsupported | session-budget-alerts |
Unsupported | (not_applicable) Budget notification is a hosted billing feature: it needs an outbound channel to a party who pays for the account, and SandBase is single-tenant and local, so the operator is already the only party to notify. |
| unsupported | mcp-tunnel |
Unsupported | (not_applicable) MCP tunnel is a hosted connectivity feature outside the local-first scope. |
| unsupported | web-search-execution |
Unsupported | No search provider is bundled or configured, and search-engine HTML scraping is not an accepted substitute; enabling web_search fails admission before a session is persisted. WebFetch execution is a separate, implemented capability. |
CLI
managed-agents init
managed-agents start [--host 127.0.0.1] [--port 3000]
managed-agents list
managed-agents reload
managed-agents chat <agent-id> --message "hello" [--tool-approval ask|allow|deny]
managed-agents template list | install <name> | create <name>
A turn whose tool needs approval parks instead of failing, and chat asks before
running it, then lets the runtime continue the same turn. --tool-approval allow
decides every such call in advance, which is what a script or a CI job uses, and
deny refuses them. With no terminal to prompt, the default ask answers nothing
and exits non-zero with the calls that are waiting named, so a script states its
policy rather than inheriting one. A custom tool is the exception: only your own
client can produce its result, and chat says so and exits non-zero. See
usage.
SDK
import { ManagedAgentsClient } from 'managed-agents/sdk';
const client = new ManagedAgentsClient({
baseUrl: 'http://127.0.0.1:3000',
});
const session = await client.sessions.create({
agent: 'agent_...',
environment_id: 'env_...',
});
for await (const event of client.sessions.chat(session.id, 'Hello')) {
if (event.type === 'agent.message_chunk') {
process.stdout.write(event.delta ?? '');
}
}
The /v1 API follows Claude Managed Agents resource shapes, so you can also
point the Anthropic SDK at the local runtime:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic({
apiKey: process.env.MANAGED_AGENTS_API_KEY ?? 'local-dev-key',
baseURL: 'http://127.0.0.1:3000',
});
const session = await client.beta.sessions.create({
agent: 'agent_...',
environment_id: 'env_...',
});
Authentication
Open by default. Authentication activates when at least one API key exists:
# Static key via environment
export MANAGED_AGENTS_API_KEY=sk-local-example
# Or create a managed key
curl -X POST http://127.0.0.1:3000/v1/api-keys \
-H "Content-Type: application/json" \
-d '{ "name": "Local Console" }'
Clients send Authorization: Bearer <key>.
Integrations and examples
- DeepSeek Harness plugin — run this runtime as a DSH plugin over MCP
stdio: install, preflight, tool list, and troubleshooting live in
examples/deepseek-harness. A DSH project can also take a portable Skill from GitHub source —npx --yes github:sandbaseai/sandbase-skills add multi-source-searchinstalls into.dsh/skills/multi-source-search. - Agent Plugins 1.0 clients (Copilot CLI, VS Code) and the standalone
MCP bridge container — see
agent-plugin/PLUGIN.md. - Use cases — the Showcase walks through an auditable coding agent, DSH as an interactive front end, and controlled code execution across Local, Docker, Kubernetes, and self-hosted sandboxes.
- Agent configuration — the YAML agent definition,
config.yaml, and the workspace layout live in the usage guide; curl walkthroughs for every resource are in docs/api.md.
Documentation
- Machine-readable project metadata
- Agent / MCP installation guide
- Agent Plugin marketplace manifest
- Installation
- Usage Guide
- API Reference
- Skills
- DeepSeek V4
- MiniMax
- Deployment
- Architecture
- DeepSeek Harness integration
- Contributing
- Citation metadata
- Changelog
Development
npm ci
npm run typecheck # src + tests + Console
npm test # vitest
npm run build # runtime + console + SDK
npm run release:check # full local release gate
release:check runs typecheck, tests, both builds, npm pack --dry-run, CLI
init smoke, and examples/basic startup smoke.
Star and share
If this runtime solves a real agent-infrastructure problem for you, star the repository so other builders can find it.
Ecosystem directories, community guides, and related projects are in docs/ecosystem.md. Community use-case discussions: memory migration between Codex, Claude Code, and DSH, sandbox and filesystem protection for third-party plugins.
License
Comments
Comments live in GitHub Discussions. Sign in with GitHub to post or react.


