Install
Inside DeepSeek Harness, with dsh-market
dsh plugin --profile web add dshmarket
Or from the command line
dsh plugin --profile web add billion-context
Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.
README
Community
QQ Group: 1056132097 (full) 1108730198 (open)
📄 Paper / Preprint
- Model-Driven Incremental Hierarchical Compression: Training-Free Multi-Generational Context Management for Long-Lived Coding Agents (English, v0.2)
📝 The paper itself is open-sourced under the MIT License as part of the codebase (
paper/). It is a living document — anyone may edit it; improvements are welcome via pull request.
A production-scale longitudinal study: 4.5 months, three hosts, 174,327 model calls, 18.76B cumulative input tokens (~24.7B across all hosts), zero window violations on 204,800-token models, marathon sessions of 8,584–12,049 calls.
billion-context sits between any agent and its model API, rewriting Anthropic/OpenAI streams with acp-kernel compression. The model decides when and what to compress into high-fidelity summaries — not a hard truncation limit.
Why
Long coding sessions blow up context. Each provider charges per token, and once you pass the context window the session degrades or dies. billion-context compresses consumed conversation into layered summaries so you can run a single session for days — billions of tokens through one context window.
Unlike a host's built-in summarizer, compression here is incremental, reversible, and prefix-cache friendly: summaries are written in small ranges, can be decompressed on demand, and the cache prefix stays intact.
How it works
Agent (Claude Code / Codex / Cursor / Aider ...)
│ you point the agent's base URL at the proxy
▼
┌─────────────────┐
│ billion-context│ 1. parse the request (Anthropic or OpenAI shape)
│ proxy │ 2. run acp-kernel compression on the conversation
│ │ 3. inject a `compress` tool + compression philosophy
│ │ 4. forward to the real model API
│ │ 5. rewrite the streaming response
└─────────────────┘
│
▼
real model API (Anthropic / OpenAI / compatible)
The proxy injects four context-management tools (compress, decompress, search_context, acp_status) into the conversation. The model calls compress when the conversation grows, and the proxy executes it server-side — the compressed ranges are folded into the conversation history before the next turn.
An opt-in fifth tool, absorb (compress.absorb.enabled: true — see CONFIGURATION.md), compresses individual tool results the moment they arrive: large results (builds, logs, greps) get a forced absorb instruction, the model distills each into a compact summary, and the original pair is hidden from the wire from the next turn on — keeping mid-session pressure lower between fold rounds (#605).
The sixth tool, acp_rule (opt-in via compress.rules: true — see CONFIGURATION.md; once enabled the model has full rights over session rules and may call it unprompted, #1399), records persistent principle-level reminders: a short rule recorded by the model (user-emphasized lessons, behaviors to remember, major pitfalls hit) is hard-protected from compression — the call and its result stay in context across every fold — and omitting the argument lists the recorded rules; passing delete with a rule id (e.g. "rule3") removes one rule and clear: true removes all of them (ranxianglei/billion-context-pi#433).
The seventh tool, acp_retrieve (opt-in on every lane — set compress.ccr.enabled: true at any level, after local verification; plugin lanes require the explicit global true so the manifest advertises the tool, #1271/#1273 — see CONFIGURATION.md), backs the content-addressed message store (built-in CCR, #1097/#1179): oversized tool results are ID-referenced at arrival instead of force-distilled — the wire keeps a byte-stable placeholder and the original goes into a per-session content-store envelope (hash-deduped), retrievable on demand via one cheap tool call. V2 makes folds lossless too: covered originals are stored when a fold lands, decompress restores ranges (startId/endId refs) instead of whole blocks, and search_context hits carry the covered mNNNNN refs so you can fetch exactly what you need. Lossless by default: a retrieve not made costs nothing but the call; a detail distilled away by absorb is gone for good. Scope: proxy mode, plus plugin lanes on the anthropic + openai wires when explicitly enabled (acp_retrieve is advertised in the plugin manifest then, #1271); responses marker/text routes and google in plugin mode stay disarmed because no request-only round-trip channel exists there (silent loss, #1097).
An opt-in tool, image_full (compress.imageCompression.enabled: true — see CONFIGURATION.md), backs image pre-compression (#1095): screenshot-like images in tool results are downscaled once at arrival — the kernel decides routing and recipe, the host encodes via optional sharp — cutting billed pixels before they enter the wire (providers bill by pixel area; halving dimensions cuts billed tokens ~4×). Non-screenshot images pass through byte-identical. Lossy by nature: when the model can't read details it calls image_full with the message's ref to restore the original resolution for the rest of the session — no proxy-side storage needed, since the client's own history still carries the original bytes (it never saw the shrunk form). Default off.
A sibling protection knob, compress.protectedLatestTools (see CONFIGURATION.md), keeps the latest snapshot of a cumulative tool (a client's todo/task list, e.g. ["todo_list", "TodoWrite"]) un-compressible while older instances fold normally — so the agent never loses its live task list to a fold (#639). Its full-history counterpart compress.protectedTools hard-excludes every instance of a tool — for independent-content results no later result supersedes (e.g. opencode/pi skill loads); protecting all instances of a chatty or cumulative-snapshot tool grows context without bound (#639), so keep it to low-frequency, high-value tools. The inverse-direction knob compress.neverPreserveRecentTools (see CONFIGURATION.md, acp-kernel >= 0.0.92) removes tools from the soft-protected recent zone so their results fold immediately — by default only decompress/search_context/read/bash are exempt from recency; removing just read from that list is the recommended remedy for the batch-read fold→re-read death loop (#1198/#1277). Its positive-facing mirror compress.preserveRecentTools (see CONFIGURATION.md, acp-kernel >= 0.0.93) is the preferred one-entry form of that remedy — { "compress": { "preserveRecentTools": ["read"] } } subtracts read from the effective exclusion list without restating or freezing the built-in default.
Two compression modes — who executes compress
The proxy runs in one of two modes, and the mode decides who executes
compress, which in turn decides how the summary travels to the model (the
"carrier"). This distinction is the root of #377.
Launcher / plugin mode (bili pi, bili codex, …) |
Proxy mode (plain client → /bili/) |
|
|---|---|---|
| Client | ACP-native agent with the bili extension (pi/omp) | Any OpenAI/Anthropic client, no extension |
Who executes compress |
The agent (pi runs it locally) | The proxy (server-side compress loop) |
compress tool call in the re-sent history? |
Yes — part of the agent's own conversation | No — ephemeral proxy-loop traffic |
| Preflight blocks (no tool call)? | Last-resort backstop — the agent normally compresses on its own compress calls, but src/preflight.ts still fires (in both modes) when the input alone exceeds the window (#470) |
Yes — src/preflight.ts compresses behind the client's back |
| Summary carrier on the wire | the compress tool call |
an acp_summary user message |
| System messages on the wire | always exactly 1 (client + prompt) | always exactly 1 (client + prompt) — summaries ride on user messages |
| SGLang "single system" 400 (#377) | cannot happen | cannot happen (summaries are user messages, not system) |
Proxy-injected compress tools |
none — the agent registers the 4 ACP tools natively | the 4 context tools (when enabled) |
| Proxy-injected nudge | yes — the agent has no nudge channel of its own, so the proxy-side nudge is the proactive compression trigger (preflight alone only fires at the hard limit; #451) | yes (when enabled) |
Why the carriers differ. In plugin mode the agent owns compression: the
compress call + result live in the agent's own history and are re-sent every
turn, so the summary rides on the tool call and the agent's view never renders
the kernel's acp_summary fallback (billion-context-pi src/messages.ts
skips acp_summary_*). In proxy mode the client is not ACP-native, so the
proxy executes compress server-side; the tool call never enters the client's
history, and preflight blocks have no tool call at all — so the kernel's
acp_summary message is the only carrier. The kernel renders it as role
system, but strict OpenAI-compatible backends (SGLang) require exactly one
system message at index 0, so systemToUser (src/util.ts) re-voices it as a
user message, leaving it at its anchor position. This keeps the head system
message (the prefix-cache anchor) byte-stable across compress turns, so a new
block does not invalidate the whole-conversation prefix.
Why user, not system or a forged tool call. A mid-stream system
message is what SGLang rejects (#377). A forged compress tool call would be
the "pure" carrier, but in proxy mode it requires fabricating an
assistant tool_calls + user tool_result pair by id, declaring the tool in
the request, and handling preflight blocks that have no authentic call — far
more invasive than re-voicing a standalone note. A user message is allowed
anywhere in the conversation, so it is the minimal change that satisfies both
SGLang's one-system rule and prefix-cache stability. The accepted trade-off:
a summary is a stand-in for the folded history, and re-voicing it as a user
turn is a semantic mismatch the model tolerates (it is clearly marked
[Compressed conversation section]).
Do the two modes coexist?
- Same proxy instance: yes, by design. One proxy serves plugin and plain
clients at once;
pluginModeis decided per request (x-bili-pluginheader) and bound per session (session.metadata.pluginAgent). The launcher reuses a running proxy. - Same session: the mode is sticky. A session created in plugin mode stays plugin mode (metadata inheritance); a plain session can only be upgraded to plugin mode if a plugin request arrives with a matching conversation id (the header outranks) — and never downgraded. In practice a plain→plugin upgrade requires the plugin client's conversation id to match an existing plain session id, which doesn't happen (each client generates its own id).
- Cross-mode block hazard: theoretical only. It would require the same
conversation id to span a mode switch. plugin→proxy is safe (the tool call is
in the shared history); proxy→plugin could orphan proxy-created block
summaries (their tool call isn't in the agent's history and the agent's view
skips
acp_summary) — but that needs the id match above, which doesn't occur.
Verifying that a compression actually landed. After executing compress,
the proxy emits a confirmation marker (📦 [ACP] Compressed …) as plain
assistant text — but under sustained context pressure a model was observed
writing that marker format itself without ever calling the tool (#717): 17
fake "compressions" over ~2 hours while real usage climbed to 89%. A marker
line visible in the transcript is therefore not proof of persistence — verify
with acp_status (block count increased, compressible-range start advanced)
before trusting it. As a backstop, the proxy strips any marker-shaped line the
model emits on its own and logs a [marker-echo] warning, and both the nudge
and the injected prompt state explicitly that markers are proxy-emitted only.
Which do I need?
Pick by your client:
| Client | Use |
|---|---|
| pi | billion-context-pi (in-process extension) |
| opencode (1.x / 2.x) | billion-context — bili opencode (launcher) or bili plugin install opencode (native, no launcher); standalone opencode-acp remains usable on 1.x. Full guide: OpenCode |
| omp | billion-context via bili omp (built-in plugin) or bili plugin install omp (self-spawning native plugin, no launcher) |
| dsh | bili dsh (launcher — full native plugin via --patch: tools, session-bound /acp + /acp-cache, fetch intercept) or bili plugin install dsh ≡ dsh plugin --profile <name> add billion-context (one unified lane — pnpm-installs the package into each profile so dsh mounts the bundled patch layer; the bili form just drives dsh's own channel per profile and migrates legacy managed blocks) |
| kimi | bili plugin install kimi (self-spawning native plugin, no launcher — Kimi Code ≥ 2.0.0; per-session routing block in ~/.kimi-code/config.toml) or bili kimi (launcher, cert-MITM) or /bili/ prefix |
| hermes | bili plugin install hermes (self-spawning native plugin, no launcher — Python plugin, #958) or bili hermes (launcher, cert-MITM) |
| zcode (Z.ai / bigmodel coding plan) | bili plugin install zcode (self-spawning native plugin, no launcher — per-session routing block in the bigmodel provider store, #1145) or cert-MITM through the GUI's Settings → Network (HTTP proxy + CA path) or /bili/ prefix |
| claude | bili claude (launcher) or bili plugin install claude (native posture, #964 — managed settings block + session-owned proxy; see the notes below) |
| jcode | billion-context via bili jcode (launcher, cert-MITM) or /bili/ prefix — no native plugin possible: compiled Rust binary with no plugin seam, and its static per-provider config can't stamp per-request headers (#962) |
| gemini (Gemini CLI) | bili gemini (launcher, GOOGLE_GEMINI_BASE_URL /bili/ rewrite) or /bili/ prefix — launcher-only: gemini-cli's extension system reaches custom commands only, no in-loop tool seam (#1043) |
| iflow (iFlow CLI) | bili iflow (launcher, IFLOW_BASE_URL /bili/ rewrite) or /bili/ prefix |
| qwen (Qwen Code) | bili qwen (launcher, cert-MITM) or /bili/ prefix |
| mcode (MiniMax Code) | billion-context via bili mcode (launcher, cert-MITM) or /bili/ prefix — no native plugin possible: its plugin system is declarative event hooks only (no model-request/history seam), so compression rides the proxy (#1050) |
| aider | billion-context via bili aider (launcher, cert-MITM) or /bili/ prefix — no native plugin possible: Python script structure whose hook surface is shell commands around edits/notifications only, no tool-injection seam (#1048) |
| copilot (GitHub Copilot CLI) | bili copilot (launcher, cert-MITM) — closed Go binary, no plugin seam; model hosts (api.githubcopilot.com + per-plan subdomains) whitelisted (#1049) |
| amp (Amp CLI) | bili amp (launcher, cert-MITM) — closed Go binary, no plugin seam; ampcode.com whitelisted (#1049) |
| goose (Goose CLI) | bili goose (launcher) — rustls release builds trust no CA file, so no cert-MITM: built-in openai/anthropic legs redirected via OPENAI_HOST/ANTHROPIC_HOST, custom providers via a regenerated GOOSE_PATH_ROOT overlay (base_url → /bili/, real config untouched); fixed third-party providers unsupported (#1049) |
| everything else (no context hook) | billion-context — bili <client> (launcher, preferred) or /bili/ prefix |
Native mode vs standalone extensions. The host-native plugins (bili plugin install pi / opencode — they spawn the proxy inside the host process) and the standalone in-process extensions (billion-context-pi, opencode-acp) are mutually exclusive: both active means double compression. The installer makes the switch: bili plugin install pi replaces the legacy npm:billion-context-pi entry (with a reminder that a project-scope entry in <project>/.pi/settings.json from pi install -l lives outside the global settings), and bili plugin install opencode strips legacy opencode-acp entries from the global opencode.json — bare name, npm: alias, versioned (opencode-acp@stable), or path form, array or object shape; the original config is snapshotted to .bili-bak once. A project-local install (opencode plugin opencode-acp writes <project>/.opencode/opencode.json, not the global config) is not touched — remove it by hand; the installer note reminds you. As a runtime safety net for manual installs, the native entries set BILLION_CONTEXT_NATIVE=<host> synchronously at load so a standalone extension can stand down at action time — its own load-time BILLION_CONTEXT_PROXY check cannot see a proxy that native mode spawns asynchronously, and its /bili/ baseUrl check never sees the fetch-layer rewrite. On the pi side the marker needs billion-context-pi 0.1.72+ (the per-event re-check landed after 0.1.71); the pi-native entry additionally scans both pi settings files once its proxy is up and warns loudly when it spots a co-resident legacy entry the installer never saw — that warning is the only visible signal while an old billion-context-pi silently double-compresses.
Install
npm install -g billion-context
This installs the bili command (bili-proxy is kept as an alias).
Quickstart
Three ways to use it — pick one:
- Native plugin (no launcher):
bili plugin install <client>— bili becomes a plugin inside the client; start the client as usual. - Launcher (easiest): one
bili <client>command brings up the proxy and the client together — no real config file is ever touched. - URL change (persistent): prefix your client's baseURL with the proxy
origin +
/bili/.
Mechanism details behind these three options (plugin lifecycle, runtime-info protocol, injection priority) live in TECHNICAL-NOTES.md.
Option 1 — Native plugin (bili plugin install pi / omp / opencode / dsh / kimi / hermes / zcode)
The proxy lives inside the client: install once, then start the client exactly as you always do — no launcher command, no env vars, no fixed port, no URL edits. Supported today for pi, omp, opencode (1.x and 2.x), dsh, kimi, hermes and zcode:
bili plugin install pi # registers a "billion-context" entry in pi's settings (npm form when bili itself was npm-installed)
bili plugin install omp # registers an extensions entry in omp's config.yml (~/.omp/agent/config.yml)
bili plugin install opencode # registers the plugin in opencode's real config + disables native auto-compaction
bili plugin install dsh # runs 'dsh plugin --profile <name> add billion-context' for every existing profile
bili plugin install kimi # writes $KIMI_CODE_HOME/plugins/managed/billion-context/kimi.plugin.json (+ installed.json record); per-session routing block lands in config.toml on first start (Kimi Code >= 2.0.0)
bili plugin install hermes # copies the Python plugin into ~/.hermes/plugins/billion-context/ (+ machine-owned bili.json sidecar) and enables it via `hermes plugins enable billion-context`
bili plugin install zcode # writes hooks.enabled + a SessionStart hook + mcp.servers.bili into ~/.zcode/cli/config.json; per-session routing lands in the bigmodel provider store on first start
bili plugin remove <client> # undo (dsh removes through the same channel; config snapshots go to .bili-bak)
bili plugin update [client] # bring every lane's bili presence up to date, each through its own owner (see below)
Where a client has its own plugin channel you can also install natively, skipping bili commands entirely:
- dsh:
dsh plugin --profile <name> add billion-contextis the very commandbili plugin install dshdrives per profile — same end state either way (pnpm into the profile, bundled patch layer mounted by dsh itself); remove through the same channel. See the dsh section below. - opencode: add the bare npm name to your real config's plugin list —
"plugin": ["billion-context"](npm form only; a git checkout has no published entry). The package publishesexports["./server"]→dist/agent/opencode-native.js, so opencode loads it through its own Npm.add machinery and the plugin self-spawns exactly like the bili-installed form. Do the two things the bili installer would have done for you too: set"compaction": { "auto": false }in the same config (otherwise OpenCode's native auto-compaction double-compresses) and keep a manual backup of the file first.
For pi / omp / kimi / claude there is no client-side channel — bili plugin install <client> writes their config entries for you (kimi's declarative
kimi.plugin.json + registry record, claude's managed settings block, …).
Single-writer: who owns which copy (#991)
Every bili presence on a machine has exactly one writer — the thing that installed it is the thing that updates it, and nothing else ever overwrites that copy in place:
| Lane | Copy lives in | Updated by |
|---|---|---|
global bili |
npm global (npm i -g billion-context) |
bili update / background auto-update |
| pi | pi's package manager (npm form) | pi update — bili never overwrites it |
| opencode | opencode's plugin dir | opencode's plugin manager — bili never overwrites it |
| dsh | each profile's pnpm store | a periodic check re-runs dsh's plugin channel per profile — driven by the global bili self-update or by the profile copy's own proxy when the global isn't running (dsh-market installs, #1196); manual: dsh plugin add billion-context@latest. pnpm's hardlinked store must never be copied over in place |
| omp / claude / codex / kimi / zcode | no copy — entries point at the global bili install | they update together with the global copy |
| hermes | ~/.hermes/plugins/billion-context/ (copied files + bili.json sidecar pointing at the global dist) |
bili plugin update hermes re-copies the files; the sidecar tracks the global install |
This is enforced in code, not just convention: the self-updater
(src/update.ts → hostManagedInstall) detects install dirs under a pnpm
virtual store (.pnpm) or a host agent tree (pi / opencode / dsh / kimi /
omp homes) and skips them; installViaTarball refuses them structurally
so direct callers cannot corrupt a store either. Mixing commands is fine
(dsh plugin add ≡ bili plugin install dsh — same channel, same records);
mixing writers is what the guard forbids. bili plugin update [client]
is the one command that drives every lane through its own owner and prints
the per-lane update path (bili plugin list shows the same per-lane channel).
At load the plugin spawns its own proxy (attaches to a healthy running
one only when it passes the attach gate below; a parent-pid watchdog tears
it down when the client exits),
rewrites model traffic to <proxy>/bili/<upstream-url>, registers
compress / decompress / acp_status as native client tools (plugin
mode), and reports the client's own model config to the proxy so
compression budgets use the real window instead of a registry guess.
Opt-out envs: BILI_NATIVE_PI=0, BILI_NATIVE_OMP=0,
BILI_NATIVE_OPENCODE=0, BILI_NATIVE_DSH=0, BILI_NATIVE_KIMI=0,
BILI_NATIVE_HERMES=0, BILI_NATIVE_ZCODE=0. Full
mechanics: TECHNICAL-NOTES.md.
Reuse is identity-based (#1225) and lifecycle-gated (#1335): an existing
proxy is attached only when it runs the same code (sha256 of the entry
script, recorded in the instance file), its lane is compatible — each
launcher declares its client's lane, two different declared lanes never
share — and it owns a session lifecycle: its health endpoint reports an
armed parent-pid watchdog (watchdog.armed == true), i.e. it was spawned by
a launcher with a parent pid and dies when the last attached session dies.
An instance without a declared lane is wildcard-compatible on the lane axis,
but that alone no longer makes it attachable (see the gate below).
Instances written before #1225 carry no code fingerprint and are therefore
never attached: a rebuilt or updated install always starts a fresh proxy on
the next launch, so fixes take effect immediately instead of silently
serving stale code.
The attach gate (#1335). A native hook attaches to whatever answers on the port, so the three listener kinds get different treatment (TS lanes and the hermes Python plugin's discovery path alike, #1338):
| Listener | Lifecycle owner | Attach? |
|---|---|---|
| Its own session-spawned proxy | armed from birth | ✅ yes |
| Another session's armed proxy (shared, watcher set #1186) | watcher set | ✅ yes — sharing stays the design |
Manually started bili start daemon |
none — refuses watchers, never dies with sessions, often an older build | ❌ not by default |
The hook probes the candidate's /__bili/health for watchdog.armed before
attaching. Armed → attach + register a watcher (current behavior, README
lifecycle contract holds). Unarmed — or a pre-#1330 build that reports no
watchdog field at all (unverifiable, treated as unarmed) → do not
attach; the hook spawns its own session-owned proxy (ephemeral port, armed
from birth, dies with the last session). This also fixes version skew: every
session now runs the currently installed bili instead of whatever a
stale resident daemon happens to carry. The trade-off is one extra short-lived
proxy process per session when no armed proxy exists (session state is shared
on disk, so compression continuity is unaffected); the multi-instance warning
(#394) becomes correspondingly more common. Escape hatch: deliberately
run a resident daemon for your hooks to ride on → set
native.attachExternal: true in the config file or
BILI_NATIVE_ATTACH_EXTERNAL=1. That restores attaching to any compatible
listener regardless of watchdog state — you then own the daemon's lifetime
and version yourself. Explicit user-directed attaches (BILLION_CONTEXT_ATTACH
/ preset BILLION_CONTEXT_PROXY for kimi/dsh) bypass discovery entirely and
are exempt by construction.
Attach discovery is lane-aware across all live instances (#1232): the
launcher probes every live entry in the instance registry, not just the
single instance file (last-writer-wins — under concurrent multi-client use
it can point at another client's proxy), and applies the gate above to every
candidate. Among compatible candidates the newest instance with the launcher's
own declared lane wins; an instance without a lane is wildcard-compatible on
the lane axis (still subject to the gate). The another bili instance is running warning (#394) is lane-aware too: it fires for same-lane or lane-less
coexistence, but stays silent between two different declared lanes, whose
session files are disjoint.
Runtime-info protocol (#955). A native plugin reads the model config the client itself will use and pushes it to the proxy (per-request headers
- bootstrap report); the proxy prefers that truth over the models.dev registry / built-in table when resolving the context window. Protocol details, resolution order, and implementations: TECHNICAL-NOTES.md.
Notes:
- Native mode is mutually exclusive with the standalone in-process
extensions (
billion-context-pi,opencode-acp) — the installer swaps the entries and snapshots the original config (.bili-bak); migration details in the client table above (pi needsbillion-context-pi0.1.72+ to stand down cleanly). - OpenCode: legacy
opencode-acpsessions, the V1/V2 plugin shapes, and all caveats are consolidated in the OpenCode section. kimireports runtime-info at bootstrap only (staticcustom_headerscan't carry per-request window/model headers without going stale on model switch) and binds subagent conversations by per-callconversation_id— full mechanics in the "Kimi Code" section below.hermes's native plugin is Python (its CLI agent's plugin API is Python-only) — instead of patching fetch it points hermes' httpx stack at the proxy via env vars after a health check, and stamps per-request headers through anllm_requestmiddleware; full mechanics in the "Hermes" section below.codexhas a companion install too (an MCP shell), but it needs a running proxy — it is not native mode.claudealso has a native posture (#964):bili plugin install claudewrites a managed settings block (static/bili/URL +SessionStarthook) plus an MCP shell pinned to a stable port — the proxy lives and dies with the session. Opt out withBILI_NATIVE_CLAUDE=0(passthrough). Mechanics: TECHNICAL-NOTES.md.zcodealso has a native posture (#1145):bili plugin install zcodewrites~/.zcode/cli/config.json(hooks.enabled+SessionStarthook + stdio MCP server) and rewrites the bigmodel coding-plan provider'sbaseURLto<proxy>/bili/<upstream>per session (both store generations: legacyv2/config.jsonand v3.14+provider_config.json) — full mechanics in the "ZCode" section below.jcodehas no native mode at all: it is a compiled Rust binary with no plugin or extension seam, its only per-provider request surface is a static TOML header table applied verbatim to every request, and its MCP servers run in a global pool shared across all sessions — so there is neither a way to rewrite model traffic in-process nor one to stamp the per-request headers plugin mode requires (x-bili-plugin, conversation id, runtime-info). Full source-level analysis: #962 (closed wontfix). Usebili jcode.aiderhas no native mode either: it is a Python script structure whose hook surface is limited to shell commands around file edits and idle notifications (--git-commit-verify,--notifications-command) — there is no plugin or extension API and no MCP client, so there is no tool-injection seam for plugin mode. Usebili aider(#1048).copilot,ampandgooseare launcher-only (#1049): none exposes a tool-injection seam, so there is no native mode (a codex-style MCP-shell companion remains possible for amp/goose but is not shipped). Goose additionally cannot be cert-MITMed — its release builds run rustls/webpki and trust no CA file — so it rides plain-HTTP base-URL redirects instead of proxy envs.
Option 2 — Launcher (bili pi / bili codex / bili claude / bili omp / bili opencode / bili hermes / bili dsh / bili codebuddy / bili qoder / bili trae / bili jcode / bili kimi / bili gemini / bili iflow / bili qwen / bili mcode / bili aider / bili copilot / bili amp / bili goose)
The launcher wraps a client in one command: it starts a proxy on an
independent port (a fresh instance is always spawned — a port is never
reused), then points the client at it — certificate-based MITM where the
client honors proxy/CA env vars, or an isolated /bili/ config rewrite
where it doesn't. No real config file is ever edited; the client's own
config is READ to discover which HTTPS upstream hosts it talks to, and those
hosts are whitelisted for MITM so the proxy can TLS-terminate exactly them
and blind-tunnel everything else.
bili pi # launch pi through the proxy — file-free (#535): env + extension registerProvider, real ~/.pi untouched
bili codex # launch codex through the proxy
bili claude # launch claude through the proxy
bili omp # pi-style, file-free (#535): env + extension registerProvider + compaction cancel, real ~/.omp untouched
bili opencode # OpenCode (1.x & 2.x): full guide in the [OpenCode](#opencode) section below
bili hermes # file-free (#535): hermes proxy env (HTTPS_PROXY + combined CA bundle via SSL_CERT_FILE) — https via CONNECT MITM, http via absolute-form forward proxy; real ~/.hermes untouched
bili dsh # deepseek-harness: full native plugin injected via --patch (#941) — compress/decompress/acp_status registered as real dsh tools, requests stamped with the dsh session id (plugin mode), /acp + /acp-cache session-bound; non-loopback upstreams ride proxy envs (https MITM, http absolute-form), loopback keeps the overlay DSH_HOME (~/.dsh-bili) rewrite (#535), built-in deepseek route via DEEPSEEK_BASE_URL; dsh native auto-compaction disabled (compaction-basic auto:false)
bili codebuddy # Tencent CodeBuddy Code CLI: CODEBUDDY_BASE_URL /bili/ rewrite (OpenAI chat completions wire), budget aligned via CODEBUDDY_AUTO_COMPACT_WINDOW; real ~/.codebuddy untouched
bili qoder # qoder: model endpoint is hardcoded https (no /bili/ rewrite possible) — cert-MITM via HTTPS_PROXY + NODE_EXTRA_CA_CERTS, default model hosts whitelisted (#653)
bili trae # Trae CLI (ByteDance, closed Go binary, no base-URL override) — cert-MITM via HTTPS_PROXY + SSL_CERT_FILE, model host from TRAE_CLI_API_HOST or the default enterprise gateway (#655)
bili jcode # jcode (Rust agent harness) — env-only cert-MITM launch: HTTPS_PROXY + SSL_CERT_FILE, model host api.z.ai whitelisted, local loopback providers stay direct via NO_PROXY
bili kimi # Kimi Code CLI (Moonshot): honors standard proxy envs for all traffic EXCEPT an unconditional loopback bypass — non-loopback https via cert-MITM (HTTPS_PROXY + NODE_EXTRA_CA_CERTS/SSL_CERT_FILE), non-loopback http via absolute-form forward proxy; provider/model hosts from ~/.kimi-code/config.toml (KIMI_CODE_HOME respected) or the managed OAuth endpoints when none declared; loopback endpoints inventoried with a manual /bili/ prefix hint (#757)
bili gemini # Gemini CLI (Google): GOOGLE_GEMINI_BASE_URL /bili/ rewrite to generativelanguage.googleapis.com (Google native wire), real ~/.gemini untouched
bili iflow # iFlow CLI: IFLOW_BASE_URL /bili/ rewrite to apis.iflow.cn/v1 (OpenAI chat-completions wire), real ~/.iflow untouched
bili qwen # Qwen Code (multi-protocol gemini-cli fork, no base-URL hook): cert-MITM via HTTPS_PROXY + NODE_EXTRA_CA_CERTS, default DashScope/Qwen model hosts whitelisted, custom relays via --mitm-domain
bili mcode # MiniMax Code CLI: honors standard proxy envs for all traffic EXCEPT an unconditional loopback bypass — non-loopback https via cert-MITM (HTTPS_PROXY + NODE_EXTRA_CA_CERTS/SSL_CERT_FILE), non-loopback http via absolute-form forward proxy; provider hosts from ~/.minimax*/config.yaml (MINIMAX_DATA_DIR/MAVIS_DATA_DIR respected) or the official agent.minimax.* endpoints when none declared; loopback endpoints inventoried with a manual /bili/ prefix hint; session bound via the X-Mavis-Session-Id header (#1050)
bili aider # Aider (Python pair programmer): cert-MITM via HTTPS_PROXY + SSL_CERT_FILE/REQUESTS_CA_BUNDLE; endpoint from OPENAI_API_BASE / ANTHROPIC_BASE_URL etc., --openai-api-base, or .aider.conf.yml — api.openai.com + api.anthropic.com assumed by default; loopback endpoints stay direct via NO_PROXY (#1048)
bili copilot # Copilot CLI (GitHub, closed Go binary) — cert-MITM via HTTPS_PROXY + SSL_CERT_FILE, api.githubcopilot.com + per-plan subdomains whitelisted (#1049)
bili amp # Amp CLI (Sourcegraph, closed Go binary) — cert-MITM via HTTPS_PROXY + SSL_CERT_FILE, ampcode.com whitelisted (#1049)
bili goose # Goose (Block, Rust/reqwest): rustls release builds trust no CA file — no proxy envs at all; built-in openai/anthropic legs redirected via OPENAI_HOST/ANTHROPIC_HOST, custom declarative providers via a regenerated GOOSE_PATH_ROOT overlay with base_url /bili/ rewrites (real config untouched, user edits merged back); fixed third-party providers get a warning (#1049)
bili pi --mitm-domain api.foo.com # add a domain to the MITM whitelist
Option 3 — URL change (/bili/ prefix)
Start the proxy:
bili
Then just prefix your client's existing baseURL with http://localhost:8787/bili/.
The full upstream URL is embedded in the path, so the proxy knows where to
forward without any config:
client baseURL before: https://api.openai.com/v1
client baseURL after: http://localhost:8787/bili/https://api.openai.com/v1
That's it — put your real API key in the client config as usual (the proxy passes it through untouched). Context windows (gpt-5.1-codex=400K, glm-5.2=1M, claude-opus-4=200K, …) are looked up from models.dev automatically.
For per-client configuration examples (OpenCode, Codex, Pi, login-client MITM, …) see the web UI guide at http://localhost:8787.
Verify. With the proxy running and your config saved, check it answers and that your first real request shows compression activity in the log:
# Health check (proxy up + where it forwards)
curl -s http://localhost:8787/__bili/health
# → {"ok":true,"upstream":"https://api.anthropic.com"}
# Live session stats (after a real request)
curl -s http://localhost:8787/__bili/stats
Then send one message from your client and watch the log
(~/.local/state/billion-context/bili.log, also printed to stderr). You
should see a processTurn line per request, and once the conversation grows,
[acp-usage] round N input=X cached=Y (cache hit Z%) + a compress event.
dsh (deepseek-harness)
Two lanes, same plugin (#941):
- Launcher:
bili dshinjects the full native plugin through a--patchoverlay (~/.dsh-bili/.bili-acp.patch.yml) — every profile boots with the bili tools registered natively, model requests carryx-bili-plugin+ the dsh session id (plugin mode), and/acpis session-bound. dsh's native auto-compaction is disabled in the same patch (compaction-basic→auto: false); manual/compactstays available. - Profile install (no launcher) — one lane (#966):
bili plugin install dshrunsdsh plugin --profile <name> add billion-contextfor every existing profile — pnpm installs the package into each profile's ownnode_modules, and dsh mounts the bundled patch layer (dsh.bundle.patch.yml) automatically. The spec follows how bili itself was installed (#925): an npm-form install passes the registry name, a checkout/dev build passes its absolute path (alink:dependency, so local work stays live). Legacy managed blocks (# bili begin/# bili end, written by pre-#966 installs) are stripped on install and remove — user entries and comments survive, an emptied file gets its placeholder[]back. Run dsh once in each profile first so the profile dirs exist. The plugin spawns its own proxy at load (attaches to a healthy one instead of doubling; parent-pid watchdog), rewrites model-API traffic to<proxy>/bili/<upstream-url>via a global fetch patch, registers the manifest tools verbatim, and gates plugin-mode headers on tool readiness (round 1 rides wire mode). Opt-out:BILI_NATIVE_DSH=0. Remove withbili plugin remove dshordsh plugin --profile <name> remove billion-context— both go through the same channel. Registry installs require a published release that carriesdsh.bundle.patch.yml. If dsh fails to boot right after an add withERR_MODULE_NOT_FOUNDonbillion-context/dsh, the profile resolved a pre-bundle copy from a stale package-metadata cache (#953) — re-add pinned:dsh plugin --profile <name> add billion-context@latest. - Auto-update keeps profiles in lockstep: the refresh has two triggers —
after a global self-update, AND from the profile copy's own proxy when
its periodic check sees a newer registry version (so dsh plugin-market
users with no global bili running still refresh, #1196). Both scan
~/.dsh/profiles/*/package.jsonand bring any registry-pinnedbillion-contextdependency to the target version (the new global version for the global trigger, registry-latest for the self trigger), always through dsh's ownplugin addchannel — never an in-place copy — so the loaded plugin and the proxy never drift apart again (#953); profiles pinned to a local source are left alone. The refresh is best-effort, retries next cycle on failure, and never fails the update or the proxy. - Reported: zero proxy traffic for some transports under profile install
(#1158, under investigation): sessions served by some of dsh's
llm-pi-ai-layer transports show NO model request ever reaching the proxy (noprocessTurnlogged; bili tools 404 with "no model request has arrived") while other providers in the same host work normally. The root cause is still being pinned down with runtime evidence — candidates: the transport-level fetch shape (SDK-injected fetch / non-global dispatcher) or a host-side attribution gap leaving the traffic unclaimed by the takeover gate. Detection: the proxy logs a one-time[plugin] NO MODEL REQUESTS seen for conversation …warning, and the dsh plugin logs each distinct endpoint the attribution gate lets through unproxied (once per process). Reliable workaround meanwhile: launch throughbili dshinstead — the launcher's settings overlay rewrites those providers'baseURLs to/bili/URLs, so the traffic reaches the proxy regardless of which fetch the transport uses or what the attribution state is.
Under a bili dsh launch the plugin ATTACHES to the launcher's proxy (no
second spawn). Raw upstream URLs rewrite to <proxy>/bili/<url> like
spawn mode (a loopback proxy target is never proxied, so the MITM envs are
simply bypassed); already-routed /bili/-prefixed requests pass through
untouched except for header stamping. Known limitation: manual
/compact has no dsh-side event hook, so its boundary is left to the
kernel's natural ingest diff (auto-compaction is off, so this is rare).
Kimi Code (Moonshot)
Three aligned modes: bili kimi (launcher, cert-MITM — Option 2), /bili/
URL prefix, and native plugin mode (bili plugin install kimi, #963). Kimi
Code v2's plugin system is declarative only (kimi.plugin.json: MCP servers,
hooks, skills — no in-process JS execution), so bili cannot patch the client's
fetch stack like it does for p
…
Comments
Comments live in GitHub Discussions. Sign in with GitHub to post or react.