Skip to content
dsh-market Browse plugins GitHub 中文

drscrewdriver/dsh-prime-memory

Layered memory for DSH whose agent-facing surface is 10 tools: high-privilege writes (memory_add, memory_delete, memory_import), rumination controls (memory_ruminate, _cancel, _status), and a memory graph (memory_search_graph, memory_expand_graph_node), built over an L0-L3 distillation pipeline that recalls and injects relevant memories before each model step.

Stars ★ 1 Category Memory Listed 2026-09-18 npm dsh-prime-memory

Install

Inside DeepSeek Harness, with dsh-market

dsh plugin --profile web add dshmarket

Or from the command line

dsh plugin --profile web add dsh-prime-memory

Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.

Screenshots

README

dsh-prime-memory

A layered distillation memory plugin for DeepSeek Harness: conversations are processed in the background through L0 capture → L1 atomic memories → L2 scene consolidation → L3 persona distillation, and relevant memories are automatically injected into context before every model step — neither the user nor the model needs to do anything.

简体中文 · Latest release · Report issues

DSH Version Compatibility Matrix

DSH version settings registration API Status
0.1.1-rc.2 settings.register() (live scope) ✅ Verified
0.1.2-rc.1 settings.register() (fallback available) ⚠️ Inferred from framework docs, not field-tested
0.1.3-rc.1 settings.register() (fallback available) ⚠️ Not field-tested (0.1.3+ namespaces became plain strings; this plugin is compatible)
0.1.5-rc.2 settings.register() (fallback available) ⚠️ Not field-tested; Session V3 surface semantics and input-bar/settings slots pending regression
0.2.0-rc.1 settings.register() (live scope) ✅ Current adaptation line (event wiring agent/session-start → agent/created)

Compatibility mechanism: settings registration uses a three-way runtime branch (register → installSection bridge → always-on degradation); see the 0.11.0 entry in CHANGELOG.md. dsh.plugin.json declares engines.dsh: ">=0.2.0-rc.1 <0.2.1-0".

Getting Started

Requires Node ≥ 22.16. Two invocation styles — the npx prefix can replace dsh in any command below:

# Option 1: run the official CLI directly via npx (no pre-installed dsh; version can be pinned, e.g. dsh-prime-memory@0.8.4)
npx -y @deepseek-ai/dsh plugin --profile web add dsh-prime-memory

# Option 2: with the dsh CLI installed (dsh is a pnpm forwarder; npm i -g pnpm first if missing)
dsh plugin --profile web add dsh-prime-memory

# Alternative sources: GitHub repo / local path (dev & debugging, link: points at the repo; npm run build + restart dsh to apply)
dsh plugin --profile web add https://github.com/drscrewdriver/dsh-prime-memory
dsh plugin --profile web add /path/to/dsh-prime-memory

Install via an AI Agent (Recommended)

If your current agent can run terminal commands, send it this message as-is:

Please install the dsh-prime-memory plugin for the web profile of DeepSeek Harness.

Run only the two commands below and do not modify any other profile:
dsh plugin --profile web add dsh-prime-memory
dsh --profile web --dump-config

Confirm that dsh-prime-memory appears in the output, then report the result to me.
Do not close or restart my running DSH yourself; after installation, remind me to manually restart the DSH Web Host.

The agent should report the installation result and explicitly tell you whether dsh-prime-memory has appeared in the configuration.

This package declares a dsh.bundle composition layer (cordis.patch.yml); after installation the plugin entry is mounted automatically — no need to hand-edit $DSH_HOME/profiles/web/cordis.patch.yml. Then restart DeepSeek Harness and verify: the appearance of conversations/ records/ scenes/ and memory.db under ~/.dsh/memory/ means the plugin applied successfully; the "Memory" page in settings and the mode pill in the input bar mean the client half is ready.

⚠️ Security note: installing a plugin = running third-party code with your privileges. This plugin reads session content, writes files in its data directory, and calls the LLM/embedding services you configured; if that concerns you, review the source first (src/).

Uninstall: dsh plugin --profile web remove dsh-prime-memory + restart. Data stays in ~/.dsh/memory/; delete the whole directory manually if you don't need it.

Development from Source

git clone https://github.com/drscrewdriver/dsh-prime-memory
cd dsh-prime-memory
npm install && npm run build
dsh plugin --profile web add .        # link: install; after code changes, npm run build + restart dsh
npm run smoke                         # smoke test (rebuild first: see command below)
npx tsc src/smoke.ts --outDir dist-smoke --module nodenext --moduleResolution nodenext --target es2022 --strict --skipLibCheck --esModuleInterop

Runtime Data Flow

The plugin attaches to DSH-native event seams (session/event for capture, agent/pre-step for injection) and reuses the host's ctx.llm for distillation. Recall is presented as message-side injection: relevant memories enter the conversation as a synthetic message placed right before the user's new message, rendered as a "Context injection · memory" row in the chat flow (expand to see the hits) — so you can see "memory at work" directly. Injected content is bounded by length and time budgets — oversized lines are truncated (pointing the model at the memory tools for the full text) and a timed-out recall silently skips that turn, never slowing the chat. Per-session dedupe: a memory already injected in this session is not injected again (the model's context already holds it — follow-up questions on the same topic save tokens); the record resets when the context is compacted or cleared, so memories can flow back in, and an updated memory (new id after a content change) is never held back by the old suppression. Freshness weighting: recall ranking applies a soft weight relevance × max(0.5, 0.5^(days since last update / 30)) — among candidates of similar relevance the fresh one wins (slots rotate naturally), while a strongly relevant old memory still recalls fine (the floor caps its loss at half a ranking score, so long-lived facts never sink); tune via recall.decayHalfLifeDays, 0 disables. It also registers three model-callable memory tools: memory_search / conversation_search / memory_read_scene.

Cost dashboard: every distillation LLM call (extract / dedup / L2 / L3) writes its token cost to a SQLite detail table keyed by provider/model (configurable retention, default 365 days with rolling cleanup on write; accounting failures only log a warning and never block distillation). Visualize it under Settings → Memory → the Cost tab: per-model trend lines (day/week/month granularity + last-N-days window + L1/L2/L3 layer filter), a layer × time-window table (calls / output & reasoning tokens / mean / median), and per-model totals — distillation overhead at a glance. Input is counted in characters (DSH streaming usage carries no input tokens); output and reasoning in tokens.

In action: the "Context injection · memory" row surfaces relevant memories first, and the model then calls memory_read_scene directly to read scene blocks before answering from memory:

In restricted sessions where only the code-execution entry point is available, the model reaches the memory tools indirectly through run_code (nested as SUBTOOL calls in the trajectory view):

Layered Memory (L0–L3)

Per-Session Memory Modes

  • Control: the pill next to the mode selector in the input bar (Memory · Auto); clicking opens a macOS-style sliding picker above — release to snap to the nearest mode; adapts to light/dark themes;
  • The lower half of the popover is a per-session info area: recall hits (hit/searched turns plus cumulative items), batching progress (this session's slice x/effective threshold; the off mode shows parked slices instead), memories produced for this session, and session message count — plus status lines for anomalies (storage degraded / vector search unavailable) and a global summary (pending distill count, last distill time). Data comes from the dsh-memory/session-stats endpoint (in-memory registries + an indexed COUNT, zero file I/O), adaptively polled while open (2s busy / 5s idle) and stopped on close;
  • Each session's choice is persisted by sessionId to session-modes.json, surviving restarts/session restore; stacks with the global switches (global is the master gate); L2/L3 are fully family-isolated — content never leaks across families.
  • Write-only sessions (#38): a three-state "injection" switch inside the popover (follow global / on / off) — set to "off" for a write-only session: capture and distillation continue as usual (conversation still settles into L0→L1→L2/L3), but nothing is injected into this session (recall injection, the persona/navigation stable section and the tools guide all stop; memory_search and the other read tools return a write-only notice). The pill face changes to Memory · Write-only; the override persists per session, and switching back to "follow global" clears it to the settings-page recall toggle. Ideal for debug/eval/sensitive sessions that should absorb without interference. Orthogonal to the off mode: off remains full stealth (capture off too), while write-only keeps the "in" and gates the "out".

UI Preview

Measured Comparison (DSH-MemBench: Automated Benchmark)

Screenshots show what the plugin looks like — this section answers "what does enabling it actually buy you?" with measured numbers from an automated benchmark (bench/, one command to reproduce). Method: the same scenario bank with verbatim-identical inputs runs in Group A (memory on) with 3 merged repetitions and Group B (memory off) with 1 repetition (a memory-off long task burns multiples of the tokens per scenario — a cost guardrail); the dialog track now runs Group A only (memory-off probes in independent sessions cannot succeed, so the control carries no information — retired). Dialog-track environment: DeepSeek official deepseek-v4-flash, plugin 0.8.5 (judge same-source as tested; every answer archived for manual audit), Windows; taxonomy adapted from LongMemEval / LoCoMo / AMB, with the extended probe types and lifecycle track informed by MemoryAgentBench / GoodAI LTM / BEAM.

The dialog track below is the fresh 0.8.5 baseline (fixed plugin + corrected judging criteria); the workflow-track numbers remain the archived 0.8.3 run (the bank has since grown to 8 scenarios with a prospective-memory addition — re-run pending).

Dialog track (20 scenarios × 10 probe types × 3 reps = 420 questions): does it remember correctly

0.8.5 baseline (Group A data; the dialog-track B arm is retired — Group A only).

Dual-channel recall (Group A): passive injection hit rate 78.1% (the answer's key points appear in the recall injection, 281/360); most of the rest the model recovered by actively calling the memory tools — 106 questions with active queries, 75 rescued by tools. The end-to-end 95.2% is the composite of both channels plus model utilization. With the memory store accumulating across scenarios for the whole run, 295 probe injections carried other scenarios' memories (honestly counted) — yet accuracy actually rose from 92.8% (early, small store) to 97.7% (late, largest store), and offline flooding with 600 extra synthetic records moved retrieval recall@5 by only −2.8pp: interference resistance under a growing store, measured.

Layered weaknesses: offline retrieval metrics (recall@5, controlled replay) total 73.3%, with event ordering at 0% and scene recall at 50% — end-to-end still 93%+ thanks to model robustness over adjacent injected memories. Efficiency triangle (the cost of memory): injections add no latency (injected turns respond 210ms faster on average), recall text is ~10.3% of per-turn input, and the whole distillation pipeline costs ≈2727 input / 240 output tokens per captured message (1172 calls, 0 failures).

Workflow track (archived 0.8.3 · 7-scenario edition · Group A ×3 / Group B ×1, real tool sandbox): does it do it right, and cheaper

Probe-phase completion 85.5% vs 43.5% (+42pp): both groups have live context during teach/change phases — the probe phase (continuation task in a fresh session) is the pure memory window. Group A scored a perfect 12/12 on all three new probe archetypes (workflow knowledge update / twin-runbook disambiguation / style-convention continuity), consistent across all three reps; Group B scored 0/4 on style-convention probes (naming/structure/thousands-separator/footer conventions exist only in memory — they cannot be explored out of the sandbox), while on the workflow-update scenario it can reverse-engineer the procedure by reading the script (discrimination limited by sandbox affordances, honestly noted).

Long-task cost: Group B burns 6.8× Group A's input tokens per scenario (1.81M vs 266k) — without memory the agent advances by re-exploring, and under a high reasoning effort it even builds its own projects to probe what a one-line script convention would have done; output tokens 3× (46.2k vs 15.4k), steps +70%. This is memory's core value: what it saves is not task difficulty, but pointless round-trips and re-exploration.

Methodology & reproduction

node bench/harness/run.mjs --arm A --repeats 3 --provider deepseek-official --model deepseek-v4-flash   # dialog track (Group A only)
node bench/harness/run.mjs --track workflow --arm AB --repeats 3 ...                                  # workflow track (A/B arms in parallel)
node bench/harness/run.mjs --track lifecycle --arm A ...                                              # lifecycle track (gating/off/rebuild/forget)
node bench/harness/report.mjs --latest [dialog|workflow]                                               # aggregate report
node bench/harness/retrieval-metrics.mjs <runDir> --flood 200,600                                     # retrieval metrics + flooding curve
  • Scoring: programmatic contains-all plus an LLM judge against key points (every answer and verdict is preserved in result.json for human audit); for stale-bearing probes (updates/update-chains/forget) an old value only fails when stated as the current answer, and abstention probes allow citing real adjacent facts while denying the asked point; workflow completion is verified programmatically from produced files and their contents (four check kinds: positive / forbidden-word / must-not-exist / exists);
  • Metric surface: beyond the per-type accuracy table (6 core + 4 extended types), reports automatically include offline retrieval metrics (recall@5 / injection precision / stale leakage), the efficiency triangle (injection latency differential / injection share / distillation accounting per captured message), scale-position analysis (accuracy & contamination vs store growth), and the lifecycle-track section (family-gating matrix / off-mode dual assertions / rebuild fidelity / forget requests);
  • Live progress: running the benchmark auto-starts a local progress panel and opens the browser (--no-panel to disable) — per-arm scenario/phase/message-level progress, heartbeat & activity freshness (distinguishes "stuck" from "process died"), and cumulative cost as it accrues;
  • Metrics come from provider-reported usage (input with cache-hit split) and session-event folding; the steady-state cache rate excludes each session's first request (0.8.5 baseline: 89.1% — memory injection does not hurt caching);
  • Regression use: run before/after a plugin change and diff with compare.mjs (environment header check including git SHA + Group-B control-drift warning + retrieval-metric comparison);
  • Limitations (stated honestly): single machine; Group A ×3 merged, Group B ×1 (cost guardrail — noisier); judge vs tested model: same model in the 0.8.5 dialog baseline, heterogeneous in the archived workflow run (glm-5.3 judging v4-flash); the scenario bank is author-built (biased toward memory-advantage scenarios — reproduce it yourself); sandbox-file affordances partially leak procedures (Group B can reverse-engineer by reading scripts — discrimination limits honestly noted); dual-tier tool audit (strict violation voids the scenario / loose heuristic flags only), with 0 violations measured on both sides.

Full reports and per-question data: bench/baseline/.

Configuration

Override configs go into the profile's own cordis.patch.yml as a top-level bare patch entry (direct id:, not wrapped in insert: — an insert with the same id as the bundle layer appends and causes duplicate loader entry id startup failure):

- id: dsh-memory
  name: dsh-prime-memory
  config:                    # keys replace whole lines (no deep merge); write out all keys you want to keep
    family: auto             # default mode for new sessions: auto | chat | work
    llm:                     # static distillation route (both fields set = deployment pin,
      provider: ''           # which outranks the settings-page route chain; when empty the route
      model: ''              # follows the route-chain primary row in the settings page or the default model)
Field Default Description
family auto Default memory mode for new sessions: auto (both families) | chat (personal) | work (work); switchable per session via the input-bar control
dataDir $DSH_HOME/memory Data directory
capture.enabled true L0 capture
capture.stripCodeBlocks true Strip code blocks from assistant messages
capture.maxMessageChars 4000 Max characters per message
capture.redactSecrets true Payload redaction: 8 secret classes (PEM/Bearer/JWT/Cookie/vendor API keys/emails/long numbers/high-entropy IDs) become [REDACTED:<KIND>] placeholders at capture time — covers L0, distill input and manual writes; false restores plaintext (irreversible once enabled)
trace.enabled true Structured trace: recall/distill events as daily JSONL under <dataDir>/trace/; switch sources in the Log tab
trace.retentionDays 14 Trace retention in days (0 = forever)
trace.captureContent false true stores recall query text (default: length + sha256 only)
extract.enabled true L1 extraction
extract.minMessages 6 Steady-state trigger threshold: run L1 extraction once a session accumulates N new messages. The effective threshold ramps up 1→2→4→…→N (first turn yields memories immediately, then batches to save calls)
extract.idleSeconds 300 Idle flush: distill a session's pending slice after N seconds of silence (catches "user left before reaching the threshold"); 0 disables
extract.backgroundMessages 10 Background messages attached to extraction (fetched per session from L0 — no cross-session contamination)
extract.candidatePool 5 Dedup candidate pool size
l2.enabled true L2 scene consolidation
l2.minNewMemories 5 New-memory threshold since last L2 consolidation
l2.maxScenes 12 Scene block count cap
l2.sceneContextLimit 3 Max similar-scene full texts attached to the L2 prompt
l3.enabled true L3 persona distillation
l3.interval 20 L3 distillation interval (new-memory count)
recall.enabled true Auto recall
recall.maxResults 5 Max L1 records injected before each new user message
recall.maxCharsPerMemory 500 Per-memory character cap for injected recall (overlong lines truncated with a hint to use the memory tools for the full text); 0 disables
recall.maxTotalRecallChars 2000 Total character cap per injected recall batch (lowest-ranked tail dropped first); 0 disables
recall.timeoutMs 5000 Overall recall budget (ms): a timed-out recall skips that turn without blocking the chat; 0 disables
recall.includePersona true Inject persona context into the system prompt (<user-persona>, stable zone)
recall.includeSceneNav true Inject scene navigation into the system prompt (<scene-navigation>, stable zone)
recall.strategy hybrid Retrieval strategy: keyword / embedding / hybrid
recall.scoreThreshold 0.3 Recall score threshold (below is not injected; applies to keyword/embedding only, not pre-fusion hybrid; tool path unfiltered)
recall.decayHalfLifeDays 30 Freshness-decay half-life for recall ranking (days, 0=off): ranking applies relevance × max(0.5, 0.5^(days since last update / half-life)) — among similarly relevant candidates the fresh one wins (slots rotate), and an old memory loses at most half its ranking score (floor keeps long-lived facts afloat)
embedding.enabled false Vector retrieval switch; off = pure FTS
embedding.baseUrl empty OpenAI-compatible /embeddings endpoint (e.g. https://api.siliconflow.cn/v1)
embedding.apiKey empty API key
embedding.model empty embedding model name
embedding.dimensions 0 Vector dimensions (required when enabled; must match model output)
embedding.maxInputChars 5000 Max characters per text (overlong inputs truncated)
embedding.timeoutMs 10000 Per-call embedding timeout (ms)
embedding.allowLocalModels true Allow the local embedding tier (deployment ceiling; when off, no model downloads and no local tier in settings)
embedding.mirror https://hf-mirror.com Download mirror root for local models (can be changed back to https://huggingface.co)
embedding.proxy '' Three-state download proxy: '' (default) = auto-detect proxy env vars (HTTPS_PROXY/ALL_PROXY etc., honoring NO_PROXY); none = disable, always direct; any other value = proxy URL (e.g. http://127.0.0.1:7890). Direct connections to the mirror are intermittently unreachable on some networks (connect timeouts and poisoned bytes have both been observed) — keep the default auto-detection on machines with a proxy
llm.provider/model empty Static distillation route (deployment pin): when both fields are set the route is locked, outranking the settings-page runtime route chain and the default model (deployments can force distillation onto a specific route); when empty the route follows "settings-page route-chain primary → default model". At runtime, configure the primary route and fallback chain in the route-chain editor under Settings → Memory → Overview → distillation parameters (pick from configured providers, including custom ones added in dsh Settings → Models; the primary row may stay empty to follow the default model) — a non-empty chain takes over this static config wholesale, effective immediately with no restart
llm.fallbacks [] Distillation fallback chain: an ordered list of backup routes tried one by one when the primary route fails (error / cut-off / network error / empty output); each entry is {provider, model, reasoningEffort?} (a non-empty effort overrides the global llm.reasoningEffort, still clamped by model capability); entries identical to the primary route are skipped; each route gets the full timeoutMs; when all routes fail, the existing per-session backoff takes over. Empty list (default) = single-route behavior unchanged (see Distillation fallback chain & slow-TTFT models below); a non-empty settings-page runtime chain (distillChain) takes over both the primary route and the fallback chain (a single-row chain = explicitly no fallbacks), empty = follow this config
llm.layerRoutes {} Per-layer distillation routing (#34): keys l1/l2/l3, each holding a complete chain (entries like llm.fallbacks, head row must have both provider+model explicitly). A non-empty chain fully replaces that layer's resolution (its primary and fallbacks all come from the layer chain; the global chain no longer participates); empty/missing = the layer follows the global chain. l1 covers both extraction and dedup call sites. Layers can also be edited at runtime in the segmented panel under distillation parameters on the settings page (takes priority over this static config); a deployment pin does not disable static layer chains (same deployer-owned config as the fallback-chain precedent). Orthogonal to and composable with the fallback chain — one complete chain per layer (ADR-0005)
llm.maxTokens 65536 Fallback output cap for non-layered calls. Each distillation stage has its own budget (extraction 16k / dedup 8k / L2 32k / L3 16k; auto ×4 when the reasoning effort is high/xhigh/max, so thinking can't starve the text budget); the per-layer budgets are runtime-adjustable in Settings → Memory → Overview → distillation parameters (empty/0 = built-in defaults)
llm.reasoningEffort empty Distillation reasoning effort: empty = auto (resolved from model capability: the model's default tier, else high); an explicit value (off/none/minimal/low/medium/high/xhigh/max) is only sent when the model declares support — effort vocabularies differ across providers (deepseek accepts off, OpenAI-style APIs use none, models that declare no tiers get nothing), and unsupported tiers degrade to not-sending with a one-time warning; output budgets auto-×4 at high/xhigh/max. At runtime, override the effort per route in the settings-page route-chain editor (per-row dropdown; the tier list follows each model's declared capability live, defaulting to this value)
llm.temperature 0.3 Distillation temperature
llm.maxInputChars 700000 Input character budget per distillation call (over-budget L1 inputs are chunked automatically); runtime-adjustable in Settings → distillation parameters → input budget (empty/0 = follow this value)
llm.timeoutMs 120000 Per-call distillation timeout (ms)
tokenCost.retentionDays 365 Retention (days) for distillation cost details (the token_cost table); rows older than this are rolled away on write. 0 = keep forever. Also the upper bound of the cost dashboard's "last N days" window
tools true Whether to register model-callable memory tools
benchControl false Register the in-process bench control service (rebuild trigger / session-mode setting / distillation usage snapshot — used by the benchmark's lifecycle track). Off by default — zero surface in production deployments; do not enable casually
scope global Storage scope (visibility range): global (default) = visible across workspaces; workspace = isolate the work family per workspace. Orthogonal to family (content type) — the two axes answer different questions, and all four quadrants exist (chat×global / chat×workspace / work×global / work×workspace). The chat family stays global by default: personal memories are meant to cross projects. With the default global, behavior is byte-identical to before this setting existed — existing single-root data is all graded global, and migration only labels ownership, never moves or deletes (ADR-0008 / ADR-0009). When no workspace can be resolved, it falls back to global (no throw, no startup block)
conflictFreeze.enabled false Master switch for conflict freeze. When on, the dedup action vocabulary gains conflict: if the model judges that "both sides look right and it cannot tell", it no longer auto-updates or merges — the pair is parked in a pending queue instead. The new memory is still stored, and neither side is rewritten; a human adjudicates via the memory_resolve_conflict tool. While off, the dedup prompt is byte-identical to before the feature existed (zero drift). Off by default: freezing spends human attention, so it must not be on by default (ADR-0010)
conflictFreeze.maxPending 100 Pending-queue cap. Once the number of unresolved pairs reaches it, new conflicts are no longer parked and are settled on the spot using the LLM's winner/loser (the pair is still written to the queue for the audit trail, with resolution = auto to distinguish it from a human verdict). The semantics are "stop taking new ones", not "quietly delete old ones" — that is what makes the bound structural rather than a promise
conflictFreeze.timeoutDays 30 Timeout fallback (days): pairs left unresolved longer than this are settled automatically at the start of the next distillation run (again recorded as resolution=auto). 0 = no timeout fallback (explicitly off, not "everything expires immediately"). Without a safety valve, "two contradictory memories recalled side by side forever" stays in the database permanently

Distillation fallback chain & slow-TTFT models

Free/slow tiers of some inference providers have first-token latencies (TTFT) upwards of 20 seconds, while some upstream gateways cut a silent connection at ~20s — distillation calls then fail at a fixed ~20s (llm aborted) long before the plugin's 120s timeout could ever matter (the scenario measured in #31). Three mitigations, pick as needed:

  1. Switch route (most direct): change the primary route live in the route-chain editor under Settings → Memory → Overview → distillation parameters (or move a fast route to the head of the chain), or pin llm.provider/llm.model statically.

  2. Fallback chain (automatic demotion): when the primary route fails, backup routes are tried in order with no manual intervention:

    llm:
      provider: opencode-go          # primary route (may be left unpinned: settings-page route-chain primary / default model)
      model: ox-alpha-free
      fallbacks:                     # entry order = demotion priority; unset = single-route behavior unchanged
        - provider: opencode-go
          model: deepseek-v4-flash
          reasoningEffort: low       # optional: per-route effort override (defaults to the global value)
        - provider: deepseek-official
          model: deepseek-v4-flash
    
  3. Per-layer routing (each layer on its own channel): distillation layers want different things from a model (L1 is high-frequency and wants cheap/fast/stable; L3 tolerates slow first packets but needs strong capability), so diverging layers can get their own chain — one complete fallback chain per layer, while unconfigured layers keep using the global chain:

    llm:
      layerRoutes:                  # per-layer routing (#34); the head row must set provider+model explicitly
        l1:                         # l1 covers both extraction and dedup call sites: a cheap, fast, stable chain
          - provider: opencode-go
            model: deepseek-v4-flash
            reasoningEffort: low
          - provider: deepseek-official   # in-layer fallback: L1 failures demote only here, never onto the global chain
            model: deepseek-v4-flash
        l3:                         # L3 persona distillation: low frequency, large inputs — a strong-capability chain
          - provider: deepseek-official
            model: deepseek-v4-flash
            reasoningEffort: high
    

    Layers can also be edited at runtime in the segmented panel (global default / L1 / L2 / L3) under Settings → Memory → Overview → distillation parameters. In-layer priority: runtime layer chain > this static YAML layer chain > global default chain, falling back level by level.

    Failure = error / cut-off / network error / empty output (stream ends normally with 0 characters — worthless for distillation since parsing always fails, so it is treated as a route failure rather than an empty return); caller-initiated cancellation does not demote; each route gets the full llm.timeoutMs (a shared budget would give a slow-TTFT fallback route less time than its real first-packet needs, defeating the chain); token costs are recorded per attempt (failed attempts get a row too, with whatever tokens arrived before the stream broke), and successful calls are attributed to the route that actually served. The route chain can also be adjusted at runtime in the route-chain editor under Settings → Memory → Overview → distillation parameters (no config edit or restart needed); the YAML below suits deployments that want to pin the static chain.

  4. Raise the timeout: llm.timeoutMs only helps when the route is genuinely slow but the gateway doesn't cut; if the gateway kills at 20s, raising the plugin timeout is futile — use the first two layers.

Storage Layout

Vectors are off by default (pure FTS). DSH's ctx.llm has no embeddings endpoint; semantic retrieval is provided by a three-state embedding source (off / remote / local), switchable at runtime in the settings page — see the next section.

Semantic Retrieval (Embedding Source)

Pick the embedding source in Settings → Memory → Overview → Semantic Retrieval; it takes effect immediately, no config edit or restart:

Three sources: Off (default; no vector embedding at all, pure BM25 keyword retrieval), Remote (bring any OpenAI-compatible /embeddings service, selectable only when the embedding.* quartet is configured), Local (pick from a built-in model catalog, ONNX-quantized CPU inference — no API key, data never leaves the machine). The local catalog is a built-in allowlist (each model pinned to a revision with per-file sha256; arbitrary repos cannot be downloaded).

  • Download: one click on the model card (default mirror hf-mirror.com, resumable downloads + sha256 integrity checks; a proxy is used when direct access is unreachable — proxy env vars like HTTPS_PROXY/ALL_PROXY are auto-detected by default, see embedding.proxy). Per-file failures auto-retry with a rotated cache key (?dshmem-retry=N, sidestepping occasionally bad CDN cache objects); hash mismatches restart from zero, network errors resume from the checkpoint; stored under models/<id>/ in the data directory, deletable from the settings page at any time;
  • On-demand runtime: the inference runtime (transformers.js, ~100–200MB) is installed only on first switch to the local tier, into runtime/ in the data directory — never in the plugin's dependency tree or install directory; model loading and inference run on a dedicated worker thread, so the host event loop is never frozen (conversations and page interactions stay responsive while text is being embedded);
  • Live switching: one click to swap sources — everything is re-embedded in the background (visible progress, cancellable; retrieval silently degrades to keywords in the meantime, conversations unaffected; a dimension change rebuilds the vector table at the new size); a failed switch keeps the old source, which a restart still uses;
  • Effective = deployment ceiling AND runtime choice: embedding.allowLocalModels=false disables the local tier entirely; without the embedding.* quartet the remote tier is unavailable (enterprise deployments can lock this down). The choice persists in embedding-source.json.

Logging & Troubleshooting

The dsh host prints plugin logs to the console; the plugin mirrors info and above to memory.log in its data directory. The typical log path of one conversation turn: L0 capture → L0 flush → distillation pipeline start → LLM call (input/output chars, duration) → L1 extraction done → pipeline end; the next turn shows recall hit N L1 records. Empty LLM output carries full diagnostics (finish reason / token counts / reasoning excerpt); JSON parse failures include the first 400 characters of the raw model output; all failure warns carry the first stack frame. The JSONL fact source is appended per turn and relies on OS write-back (no per-line fsync); an extreme crash (power loss) loses at most a small tail, and the index DB can be fully re-derived from the fact source via "Rebuild memories".

Differences from MemoryCore

  • The full pipeline is embedded (no external Gateway); distillation reuses DSH's own LLM;
  • L2/L3 changed from "LLM manipulates file tools" to "LLM outputs operation JSON / full documents, engineering side executes";
  • Recall injection happens at agent/pre-step (message-side synthetic message, the official pre-step replacement semantics) plus agent-scoped systemPrompt.context (persona/navigation stable zone — DSH-native events/services);
  • Storage/retrieval is a single-machine slimmed version of the official sqlite backend (drops multi-tenant isolation columns, TCVDB cloud backend, audit tables; tokenization uses jieba like the official one — @node-rs/jieba prebuilt binaries union CJK character bigrams: word tokens give BM25 exact-word hits while bigrams keep sub-word recall; on load failure it falls back to pure bigrams, and FTS indexes are rebuilt automatically via a tokenizer version stamp).

Credits

The direct upstream of this repository is JunNanLYS/dsh-layered-memory — the layered distillation memory plugin for DSH. Thanks to JunNanLYS for open-sourcing it: this repository rewrites the implementation layer on top of it (the first commit 0b506b8 is "净室重写清场 — remove the old implementation and build artifacts"), while the documentation, images and module layout are carried over from upstream. Compared with upstream, this repository adds 12 agent-facing memory tools (high-privilege writes memory_add / memory_delete / memory_import, ruminate controls memory_ruminate, the memory graph memory_search_graph / memory_expand_graph_node, decision-receipt backtracking memory_receipts, and conflict adjudication memory_resolve_conflict), the skills/memport cross-tool memory transfer, and storefront screenshot declarations.

Two of those tools exist for traceability:

  • memory_receipts — trace where a memory came from. Every L1 dedup decision leaves a receipt (a digest of the candidate pool it saw + the verdict), so you can ask "which run did this record come from, and what candidates did it see?" per record, or "what did that batch decide?" per run. Receipts must exist before the event — input snapshots cannot be backfilled (ADR-0006).
  • memory_resolve_conflict — adjudicate conflict pairs parked by conflict freeze (see the conflictFreeze.* settings). The verdict is winner / loser / both: picking a side retires the other from retrieval, while both means the two records are really independent facts and both are kept.

The core memory capabilities (layered distillation pipeline, prompt design, and the dual-write storage architecture) are modeled after MemoryCore from TencentCloud/TencentDB-Agent-Memory. Thanks to the original project for open-sourcing its design and implementation.

Memory Retirement, Recovery, and Cleanup

Deletion has two tiers, and the tier is chosen by cost: the reversible one is the default, while the irreversible one must be requested explicitly and always carries an export.

Action Endpoint / tool Reversible Notes
Retire (soft delete) memory_delete · dsh-memory/records-delete ✅ Keeps the main-table row, closes valid_to, writes a supersede marker, and drops only the FTS/vector rows. Retires 1 record by default; for batches pass exact ids rather than relying on semantic matching
Recover dsh-memory/records-restore — Clears the retirement marker and rebuilds indexes, returning the record to recall
Physical cleanup dsh-memory/cleanup-retired ❌ The plugin's only irreversible action. Dry-run by default (omitting dryRun deletes nothing); even when executed explicitly it first writes a full-library snapshot and verifies it by content hash, aborting with zero deletions on any mismatch
List snapshots dsh-memory/snapshots-list — Lists the snapshots under snapshots/ that have a valid manifest, with reason and record counts
Restore from snapshot dsh-memory/snapshot-restore — Writes cleaned-up records back. Dry-run by default; accepts snapshot directory names only, never paths

All three retirement paths — conflict verdict, dedup supersede (update/merge), manual delete — share one primitive, so "delete" means the same thing in all of them: recoverable.

How the cleanup safety net is built

Physical deletion must be preceded by a snapshot that passes verification (see the table above). The way back is:

  1. dsh-memory/snapshots-list — get the snapshot directory name (shaped like l1-<timestamp>-<reason>);
  2. dsh-memory/snapshot-restore — dry-run first to read missing (how many records would genuinely come back, not the snapshot's total), then write with an explicit dryRun:false.

The restore entry point accepts directory names only, never paths: otherwise this RPC would incidentally gain the ability to read any directory and write its contents into the retrieval database. Restore itself is an idempotent upsert and can be re-run safely.

One semantic you need to know: cleanup-retired only removes already-retired records, and the snapshot is taken before deletion — so every record a cleanup snapshot can bring back carries a retirement marker. By default snapshot-restore only writes the row back to the main table (and reports it honestly in stillRetired); "back in the main table" ≠ "back in recall". For a genuine one-call rollback add unretire: true, or call records-restore on that batch of ids afterwards.

Roadmap

Features under planning — feedback and priorities welcome in the issue tracker:

  • Git branch awareness: associate memories with the current git branch; recall can filter/boost by branch (orthogonal to the existing memory modes)
  • Claude Code / Codex memory import: one-click migration of existing memory assets (CLAUDE.md, Claude Code memory files, Codex AGENTS.md, etc.), fed into the layered distillation pipeline

License

MIT

Content from the project README on GitHub ↗

Comments

Comments live in GitHub Discussions. Sign in with GitHub to post or react.