Skip to content
dsh-market Browse plugins GitHub 中文

Marquez807/dsh-experience-memory

Cross-session experience memory for DSH. Lessons are evidence-graded, injected when the turn is about the same thing, and shown just before an action only when the record declares which call it applies to. A rule that holds whatever the turn is about — "always answer in Chinese" — can be marked **standing** and is then carried every turn: measured on 24 isolated real-model turns, a standing rule is followed 6/6 against 0/18 for the controls (empty store, same record unmarked, unrelated standing record). A guess still cannot become always-on: a standing record needs a verified passage like any other. What nobody looks up is retired, reversibly. Five tools, seven slash commands, zero runtime dependencies.

Stars ★ 1 Category Memory Listed 2026-09-21

Install

Inside DeepSeek Harness, with dsh-market

dsh plugin --profile web add dshmarket

Or from the command line

dsh plugin --profile web add "https://github.com/Marquez807/dsh-experience-memory/releases/latest/download/dsh-experience-memory.tgz"

Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.

README

简体中文 · English

Domain-scoped long-term experience memory for DeepSeek Harness: it tells weight from noise, accumulates lessons, lets perishable ones expire, accepts corrections, and brings the relevant lesson back the next time the same kind of work happens.

  • 204 bytes per turn, unconditionally — one line of guidance; beyond that, content is injected only when there is relevant experience.
  • Zero third-party runtime dependencies; storage is a single SQLite file.
  • Five model tools and seven slash commands, usable with no configuration.

Contents

What you want Where to go
Install it, and confirm it is actually working Quick start
Understand the mechanism that makes it remember What it does
Find the tools, the slash commands, the config keys Tools · Slash commands · Configuration
See what it looks like to the model Model Experience
Know what it cannot do Known Limitations and Deferred Work
Bring an old memory store across Migration
Change the code, build it, run its tests docs/DEVELOPING.md

Quick start

Install it (point the path at the tarball you have):

dsh plugin --profile <name> add /path/to/dsh-experience-memory.tgz

The package name is @marquez807/dsh-experience-memory — scoped on purpose. npm carries a same-named dsh-experience-memory that belongs to someone else, and the desktop app resolves a dependency by package name: install by name and you get theirs (it ships no database binary, dies on load, and takes the whole back end — and every third-party plugin — down with it; that happened for real on 2026-09-23). With a scope, installing by name fails loudly with "no such package" instead of quietly loading someone else's code. The repository is also private: true: nothing is published to npm, and installs go through this repository's release tarball. To tell which copy you have, check whether package.json's repository is Marquez807/dsh-experience-memory.

That is the whole step. dsh plugin add does more than install a dependency — it reconciles dsh.profile.bundles with what is actually installed: any dependency declaring dsh.bundle is appended to the layer stack automatically (see reconcilePlugins in @deepseek-ai/dsh). No hand-editing of the profile's package.json.

Then restart the app. Zero configuration: it works with no config at all — the default store is created at $DSH_HOME/experience-memory/memory.db, and the five tools, the seven slash commands and the resident injection all take effect immediately.

After the restart, look at one line of the startup log (that line exists on purpose; the incident it comes from is in Known Limitations):

experience-memory: store <path> — <records> records, <confirmed> confirmed, <anchored> anchored

Check that store <path> is the store you expect. An empty store and a wrong store look identical from outside — both answer every query with nothing — so this line states the path and the count, and a mis-resolved store is visible instead of silently telling you nothing. The three numbers move with the store and none of them is a target; what matters is a path you did not expect, or 0 records.

To confirm it is working, use the slash commands:

/memory-status          # how many records, how many clear the resident bar
/memory-preview 部署     # what this turn would actually inject

There is also a self-check that runs only in a maintenance pass: when the file a verified-file record cites is no longer there (deleted, renamed), maintenance marks the record needs_review and names the missing file. It flags and never refuses — the file may simply not exist yet. To run it now: tools/provenance-audit.mjs.

The four surfaces it hangs on

Surface Content Triggered by
Automatic injection Resident digest: core layer (corroborated across projects) + query layer, sharing 1536 bytes; plus one unconditional line of record guidance every turn (204 bytes) nothing
Automatic maintenance bounded maintenance on agent/turn-stopping, batches of 32 with a cursor nothing
Model tools (5) memory_recall / memory_remember / memory_feedback / memory_forget / memory_stats the model
Slash commands (7) status, preview, maintain, audit, import, harvest review, repeat failures a person

The split between tools and commands is deliberate: auditing and importing reach outside the store (they scan arbitrary directories and write in bulk), so they stay behind a human trigger. memory_stats is the only operator-view tool the model gets — read-only, no arguments, for answering "what do you remember" or checking "did what I recorded ever arrive".

Why the guidance line must be separate from the digest, and unconditional: the digest renders an empty string when no record qualifies (so it is not injected), and "the store is empty" is exactly the moment the model most needs to be told that recording exists. Fold it into the digest and it disappears together with the memories — and an empty store sustains itself. This is not speculation, it is measured: across 5 real sessions and about 5,900 tool calls, the memory tools were offered on every request turn after installation, and memory_remember was never called once until someone explicitly asked for a record.

The same line now asks for a lookup too. See "Recording is not the same as being used": telling the model only to record, never to look, is asking it to write forever and never read. So the order is look first, then record — memory_recall before starting work, memory_remember when something is learned, memory_feedback once it helped.

What it does

Each of the three stages has one layer of mechanism, plus two lessons that tie them together: recording is not the same as being used, and recorded, but not there at the moment it acted.

1. When recording — grading the evidence

Recording a lesson requires the passage it rests on (quote) and the source (source_ref). The plugin checks them itself:

Grade Condition Base score (×3.0)
verified-tool source_ref is a tool call in this session that really ran and did not error 9.0
verified-user the passage appears verbatim in a message the user sent, and that sentence is not a question or a hypothesis 7.5
verified-file the passage appears in the cited workspace file 6.0
inferred none of the above 1.5 (always a candidate)

Only the first three grades can enter the injection layer; inferred is always a candidate. That is necessary, not sufficient: the resident bar is 5.5, and the base score of verified-file is 6.0 — 0.5 above it, about 60 days at a decay of 0.0083 per day. The three grades therefore behave like this:

  • verified-tool / verified-user are resident from day one and hold on base score for a long time (9.0 / 7.5 against 5.5, roughly 360 / 180 days);
  • verified-file holds about 60 days on its own; only if nobody looks at it and nobody confirms it useful in those 60 days does it sink below the line, after which it needs either a query that hits an identifier (a path, a class name, a file name — worth 1.0) or being looked up / successfully reused (a lookup caps at +1.0, a successful reuse adds +1.5 on a logarithmic scale) to come back. Below the line it is still retrievable on demand through memory_recall.

That 0.5 is deliberate. It did not exist at first: the bar was also 6.0, exactly equal to the base score of verified-file, so any age decay pushed a record below the line — that is not "earning its place through relevance", it is "must be used the instant it is written", which in practice means never. The gap is a line drawn on purpose: a new memory gets two months of exposure for free, and after that it lives on being used.

This is deliberate: a fact read out of a file is weaker than a tool measurement or a user assertion, so it earns prompt space through "relevant to this turn" rather than through "it exists". Measured on the audit, some verified-file records already qualify immediately (the recent ones do), and the pass rate rises once an identifier is hit — which is the rule working.

1.1 A failure has to say why

The judging logic did not change; the reason for a failure now travels outward. A caller's defect ticket forced this: to work out why three of its records only reached inferred, that caller ran 5 recording experiments and read the source, and found the real causes were "I passed an absolute path and the plugin never read it" and "my quote was missing one **" — two things one line in the write response can state.

Before the change, four different failures (absolute path / outside the workspace / file missing / unreadable) all collapsed into no session or workspace evidence matched the supplied passage — a sentence that points at the passage while the real cause was the path, the classic way to send someone in the wrong direction.

Now readWorkspaceFile returns the failure reason as data, and reason states what was tried, one case at a time:

Failure What reason says now
absolute path names it as absolute, asks for a workspace-relative path, and gives the workspace root
file missing gives the cited path and lists the contents of the nearest existing directory (when the repo lives under repos/x/ and the caller wrote lib/y.js, it is obvious at a glance)
path escapes the workspace / unreadable each gets its own sentence
passage not in the file if it matches once markdown decoration is ignored, says so and asks for the whole line verbatim; otherwise gives the closest line number and its content

Two boundaries are deliberate: decoration is used for diagnosis, never for admission — a match that only works after ignoring decoration still grades inferred, so the verbatim contract is not softened; and every verdict carries route (tool-call / file / user-message / none), because source_ref is a pun (a tool-call id or path:line) and callers previously had no way to know which one their value was read as.

1.2 grade is frozen at write time

The evidence grade is fixed the moment the record is written and is never recomputed; what every recall recomputes is importance (derived from stored facts: age, reuse, failure streak). So moving or deleting the cited file later does not change the record's grade — it keeps the verdict of that moment and can no longer be checked by anyone. Workspace membership works the same way: workspace_id is resolved from the session cwd at write time, so after switching workspaces the record is invisible (not "regraded"), unless it has been promoted to domain level.

1.3 A second gate before the injection layer: relevance

Grading decides "is this worth believing"; relevance decides "is this turn about that". The two gates are independent and both must pass:

Gate Criterion Passes when
Grading importance ≥ 6.0 (evidence + history) see above
Relevance whether the token overlap with this turn's query is specific an identifier is hit, or at least one content word is shared

The second gate was added after measurement. Grading alone once allowed this injection: a record about batchSize 上限 500 was injected into the turn "把这个仓库的 README 用一句话改写" — the two have nothing to do with each other. The only reason was that FTS5 matches bigrams with OR, and the record's body contained "不得动这个值", which collided with the 「这个」 in the question. Deterministic reproduction: identifierMatches=0, bm25 only −0.59, excluded empty — no filter objected.

The shape of the problem is "a common word is not evidence of relevance". The criterion is therefore not "how many words are shared" (a two-character Chinese word yields a single bigram, and demanding several would reject obviously correct matches — the first version did exactly that and the tests rejected it on the spot), but "is the shared word a content word": src/retrieve.ts keeps a CJK function-word list (这个/可以/一句/…), and the record is rejected only when every shared word is a function word. This fixes the root cause without punishing legitimate matches that share exactly one content word.

A perverse incentive disappears with it: importance rises with successful reuse, so the more useful a record is, the more easily it clears the grading line — and, when grading was the only gate, the more easily it could slip into an unrelated turn on the back of a 「这个」. The relevance gate does not depend on history.

2. When recalling — telling weight from noise

The old system required a human to register a script hash and replay 2–32 times before promotion was allowed. Rigorous, and fatally expensive: after 139 work cycles the store held 0 stable entries. Grading here is automatic, because strictness is only worth anything when it is cheap enough to actually happen.

The resident layer is recomputed every turn by ctx.systemPrompt.context (not a boot-time snapshot), at most two sections, hard ceiling 1536 bytes:

经验记忆(领域通用,已由多个项目独立印证):
- [id] 标题 — 教训          ← core layer: present whatever this turn is about
经验记忆(与本轮相关):
- [id] 标题 — 教训          ← query layer: matching the current topic

The two sections share one 1536-byte budget. That is what makes unconditional injection affordable: the core layer reallocates prompt space rather than adding any — it cannot conjure more tokens. A section with nothing in it does not appear at all (no empty heading), and with only the query layer the behaviour is exactly the single-section one.

Why the core layer exists: the query layer is query-gated, so when the user answers "继续" there are no tokens to match and the digest empties out precisely in the middle of a long task. The core layer has the narrowest admission rules in the framework:

Condition Why
scope = domain only content independently reported by two or more workspaces is promoted to domain level
status = confirmed candidates are never injected
evidence ≠ inferred content nothing has verified is not injected
distinctWorkspaces ≥ 2 one project's habit is not a domain rule
clears the same resident bar as the query layer a core record is always a subset of the resident layer, never a back door
bounded by coreMaxRecords so it stays bounded

A workspace-level record never becomes core, however important — nothing has corroborated it.

Inside the hit set, records are ordered by:

importance = 3.0 × evidence grade   (verified-tool 3.0 / user 2.5 / file 2.0 / inferred 0.5)
           + 1.5 × log2(1 + successful reuses)
           − 2.0 × consecutive failures
           − 1.5 × staleness
           + 0.5 × log2(distinct workspaces)
           + 0.3 × log2(1 + retrievals)
           + min(2.0, 1.0 × exact identifier hits)      ← capped

Sort key importance DESC, bm25 ASC, id ASC. The old system ordered its 8 resident slots by uuid4 string, which is random sampling frozen forever — with 100 records in the store, a new memory had an 8% chance of ever entering the overview.

2.1 Recording is not the same as being used: a loop that turns memory into write-only

This one the user named directly: "what gets recorded is never used". The cause was not an unwilling agent — four things formed a closed loop:

  1. to appear in the prompt automatically, importance had to be ≥ 6.0;
  2. a file-grade memory is exactly 6.0 the moment it is written (3.0 × 2.0) — a knife edge, and a few hours of staleness pushed it below the line;
  3. staying above the line required reuse credit, which requires someone to call memory_feedback and say "this helped" — an action that happened 3 times in the lifetime of all 76 records;
  4. so of 76 records only 2 sat in the automatic layer and the rest were reachable only if the model chose to call memory_recall — while the unconditional line asked it to record and never to look. Worse: the lookup itself was never recorded, so even when a later session dug a record out and used it, the record gained nothing — and stayed silent next time.

All four were fixed:

Change Effect
memory_recall now records "this was looked up" (retrieve_count / last_retrieved_at) retrieval leaves a trace for the first time, so "was it ever used" finally has an answer
a lookup counts as a touch: the staleness anchor is max(created, last used, last retrieved) a memory dug out and used later no longer decays as if nobody cared — it climbs back into the automatic layer, and the loop is broken
retrieval credit is capped at 1.0 (0.3·log2(1+retrievals), never above 1.0) one lookup is worth less than one confirmed reuse; otherwise calling memory_recall in a loop could keep anything resident forever
the guidance now says look first, then record, and names memory_feedback the guidance already fixed "offered but unused" once (see the measurement above); this applies it symmetrically to looking

Automatic injection does not count as a lookup, deliberately: if a record's own injection counted as use, it would keep itself in the automatic layer and the number would stop meaning "somebody went looking for it".

memory_stats therefore reports one extra line that answers the question directly: 被查过 N/M 条(已确认范围内) · 从没被查过也没被确认有用的 K 条. K is the write-only backlog, and it should fall as sessions go on.

2.2 Recorded, but not there at the moment it acted

This is the second thing the user named, and it is subtler than "recorded but unused": the memory existed, was correct, and had been injected — but not in the one turn where it mattered.

A real example: a record saying "before launching Bannerlord, confirm Steam is logged in, otherwise the game exits silently after 10 seconds" had file-level evidence and was injected in 9 of that session's 15 turns — just not in the turn where the user said "开始吧". The agent launched directly and the turn was wasted.

Two causes, neither of them "the memory is broken" — both of them "the delivery is wrong":

  1. The only text used to find memories was what the user said. When the user answers "开始吧", nothing in those characters matches "Steam" or "launch". So the digest layer emptied out mid-task, and what the agent was actually doing contributed not one character to the query.
  2. There was exactly one delivery moment — the start of a turn — and that moment is decided by the user's words, not by what the agent is doing.

Two changes, both measured against that session's real log (444 tool calls), not reasoned out:

  • The query now includes what the agent is doing: the arguments of the tool it is calling, what it has written itself, its to-do list. Messages the plugin injected itself are always skipped, otherwise a hint would feed itself into the next turn's query. With no activity, the assembled query is character-for-character what it was before — and an assertion pins that.
  • Delivery happens as a tool call is about to act (precall), and the record declares which calls it applies to — recall_for is filled in when the memory is written, with one of three facts: path:<file> (this call names that file), tool:<name> (this call is that tool), command:<word> (the command line contains that word). A call satisfying one of those gets the record attached — at most one per call, at most 300 bytes. A record with no declaration is not delivered just before an action — it still reaches the per-turn digest and memory_recall.

Why no longer "whatever is in the call, matched against the record": that rule fired on 57% of 15,383 real tool calls; a hand-audited sample of 47 found 5 genuinely about the call (10.6%), and 68.7% of the deliveries matched a word that appears only in the record's body and never in the rules the record itself states (trigger/failure_mode/lesson). Tightening the threshold to remove the noise dropped recall on 25 hand-labelled scenarios to single digits — and even at the loosest setting only 14 of the scenarios' correct records ever entered the candidate pool, so the right answer was never a candidate and re-ranking could not rescue it. Not a tuning problem: "two words collide" does not entail "this lesson applies to this call." Full measurements and the failure paths are in docs/DELIVERY-GAPS.md §12–§13.

Three things tried, measured and deleted (replay said no, not laziness. Kept to say why it is not done this way — all three belong to the word-matching rule that has since been replaced):

Attempt Replay result
treating argument keys as handles too (file_path, old_string) every edit carries those keys, so 232 of 444 calls could hit something; what got picked was not what needed reading. Restricted to values
one delivery per turn (1 through 6 all tried) the turn's slot was taken by "some other record encountered earlier in the turn", and the Steam record was never delivered once. So throttling rests only on a per-record cooldown and a per-session cap, and the code says why
using Bannerlord as the handle 13 of the 17 records this workspace can see mention it, so hitting it means hitting nothing; launch-a-runtime-clean.ps1 is mentioned by 2 and ERC403 by 1 — that is what the lesson is actually about

Those three are the history of the replaced word-matching rule. The new rule's own thresholds come from a different measurement: anchoring on the file name collided 508 times on a single record (this workspace has three tools.js), and anchoring on the relative path brought that down to 83.

The old rule's measured effect on that real session was 20 hints, landing in 4 of its 15 turns (the Steam record attached to the "write the launch script" call — the same turn as launching the game, before it acted). That is the number for the mechanism that was replaced, kept as the comparison point: the new rule's equivalents are 1.18% of calls and a worst case of 83 collisions on one record (see "Does it actually work" below). They are not the same quantity — the old one fired often and off-topic, the new one fires rarely and only where the record named it.

What it cannot do, stated plainly: it does not guarantee the hint lands on the call that most needs it. The first action in a turn that touches the topic takes the slot, so the "run" call may go without — the lesson is already in that turn's conversation, but it is not "attached to that line". That is a real trade-off, written here rather than glossed over.

3. Afterwards — forgetting and correcting

  • Retirement: explicitly forgotten by the user / two consecutive failures / expired / review overdue and never reused / 90 days unused and below the score floor
  • Never physically deleted: retirement is reversible, and only purge=true removes bytes
  • Cross-project promotion: a lesson stays in the workspace that learned it until two different workspaces independently report the same content — then it is promoted to domain level, confirmed, and becomes core memory injected unconditionally every turn
  • Identity is the claim itself, not the title: the title is only a label (often an automatic summary of the body), so two records with the same body and different titles are the same knowledge. Counting the title as identity would make cross-project corroboration impossible to reach, and domain promotion would never happen
  • Re-recording retires the candidate it replaces: the model has a stable habit — write a version with no quote first (→ candidate), notice it does not qualify, then rewrite it with a file quote. Because identity is the claim, the rewritten body is a different record, and the candidate stays in the store forever: not injectable, not visible, and nothing cleans it up. Measured in a real store, 3 such pairs had formed (43% of all records). Writing a graded record now retires same-workspace candidates with the same title, points supersededBy at the new record and writes a correction log entry. Title comparison folds punctuation — the store held a pair differing only by one pair of 「」, which exact comparison treated as two different claims
  • Maintenance runs on agent/turn-stopping, in batches of 32 with a cursor, and never enters the retrieval hot path
3.1 Perishable facts: a window on a record

Long-term memory that never expires is a liability — assertions like "the current test command is X" or "the current client version is 1.5.2" quietly become false once the world changes, and because they are verified facts they rank higher. So memory_remember accepts two optional windows:

Parameter Effect
expires_in_days stops being retrieved immediately on expiry; maintenance then sets the status to retired
review_after_days does not retire on expiry, it asks for a review; if another 30 days pass (REVIEW_GRACE_DAYS) with no reuse, it retires

The split is deliberate: an expired fact should not be answered, but "needs review" is not "is wrong". And a record that has been reused is not retired for an overdue review — the review window exists to find things nobody needs, not to punish age.

Reporting the same claim again is re-verification: the new window replaces the old one instead of being ignored.

Before these two parameters existed, expiresAt and reviewAfter were only ever filled by the legacy importer, so two of the three retirement paths were unreachable for records the plugin wrote itself — mechanism complete, nothing able to start it.

Operating it

Scope

Scope Who can see it
workspace only workspaces resolving to the same root path
domain any workspace resolving to the same domain

Domain resolution order (first hit wins): plugin config defaultDomain → domain: in the workspace's .dsh/memory.yml → name in package.json → the git remote repository name → empty (workspace level only).

The last step deliberately leaves it empty instead of falling back to the directory name: treating a name like dsh主工作区 as a domain would spread one project's quirks into every directory with that name.

Turning memory off for a mode (for example "model test mode")

A mode (agent preset) cannot switch this plugin off itself: the plugin is installed at the profile layer, and a preset's disabled flags affect only the rows that preset declares. So the switch lives in the plugin, keyed by preset id (disabledPresets, empty by default). A mode listed there gives its sessions:

What is switched off Why
digest injection it is the most direct expression of "memory"; with it the model is no longer bare
the "look first, then record" guidance it tells the model that memory is available, which a test mode should not be told
the pre-call hint same reason, and it would push past experience into a specific tool call
candidate harvesting + failure counting no recording: a test session's turns do not belong in the store
the behaviour of the five memory_* tools they refuse when called, and say why (the tool list itself is hidden by the mode, see below)

The decision reads the agentPreset in the session header and also watches agent-preset/selected events — so a session that started in a standard mode and was switched to the test mode also counts (reading only the header misses it, reading only events misses every session that never switched; both are read).

How the tools disappear from the list: the mode mounts a small local plugin that calls tools.restrict({deny}) — the tool registry only accepts restrictions inside a scope (a global restriction would mask every agent's tools, so the API refuses), and a preset is a scope. The two defences are independent: the config layer governs "no injection, no recording", the mode layer governs "not in the catalogue".

Maintenance still runs: it is store hygiene (expiry, eviction) that no session sees. Skipping it would let a memory-free mode quietly stop the whole store from ageing out.

Tools

Tool What it does
memory_recall search by query, cap 16384 bytes, over the cap it truncates in order and reports it. include_candidates reviews claims you recorded but never verified; include_retired audits retired ones. It records "was looked up" only for the records actually handed over (the truncated tail does not count) — the only trace that memory was used
memory_remember record a fact / experience / strategy; without a verifiable passage it is stored as a candidate. Optional expires_in_days / review_after_days put a window on a perishable fact; optional recall_for declares which calls this record applies to, and only a record that declares one is delivered just before an action (see below)
memory_feedback attach a real outcome; success clears the failure streak, two consecutive failures retire
memory_forget retire (default) or delete outright
memory_stats read-only census: how many records, how many clear the resident bar, reuse and correction counts, recent retirement reasons. No arguments. The first line is the build id, the last line this call's id; /memory-status is the human version

The two descriptions the model reads are instructions, not capability statements: memory_remember opens with the trigger ("call this the moment you learn something that will still hold next session"), memory_recall opens with the occasion ("before entering unfamiliar territory, or before repeating a decision already made"). That is measured, not stylistic — putting a tool in the schema is not enough to make the model use it (see the 5,900 calls above). The constraint sits at the end of the description: record reusable rules only, not one-off details, transient tool output, secrets or unverified guesses.

The source_ref parameter description also states which citation can be graded: for a file, write path/file:line; to claim "this command works", cite the id of a successful tool call; and a lesson learned from a failure cannot cite that failed call — a failed call is not evidence here (the existing gradeEvidence semantics, pinned in evidence.test as "a cited tool call that errored proves nothing") — cite instead the file that records the finding. That sentence came out of measurement: in an isolated turn the model cited a failed pytest call as its source, so the record could only land as a candidate and never reach the resident bar; in another turn it found its own way to "cite the test file committed to the repo", which is gradable, at the cost of one extra turn.

quote has a hard requirement of its own: the passage must itself say the thing (a rule, an order, a value, an error message), not "the paragraph the writer happened to be reading". The evidence is the model's own judgement: given a claim about a deploy vault and a quote that only says how to start the server locally, it answers "the cited evidence does not match the claim" and throws the record away. verified-file proves the passage is in the file; it cannot prove the passage is about the claim — that is a semantic judgement, and this framework deliberately makes no model calls. So the parameter also states the fallback: when no such passage exists, record the inferred grade and say what is missing.

recall_for decides whether it reaches the model's eyes just before it acts. One of three facts:

Value Meaning When to use it
path:<file name> this call names that file the lesson is about a file, or a kind of file
tool:<tool name> this call is that tool the lesson is about how to use a tool
command:<word> the command line contains that word the lesson is about a command

With no declaration it is not delivered just before an action — it only reaches the per-turn digest and memory_recall. That is the deliberate trade-off, and its cost and rationale are in "When recalling — telling weight from noise" above: it is not a choice of "fill it in or not", it is no declaration means no just-before-action layer at all.

Slash commands (for people; the model cannot see them)

Registered through ctx.commands.register, so they appear in the same slash menu as /compact and /goal. All are recordInput: false — operator commands and filesystem paths do not enter the session record.

Command Arguments What it does
/memory-status — store census: counts, status/evidence/scope distribution, how many clear the resident bar, reuse and correction counts, recent retirements and reasons
/memory-preview [<query>] prints the digest that would actually be injected for that query, plus what on-demand retrieval would add. With no query it uses the last two user messages — the same logic the plugin itself uses
/memory-maintain — runs bounded maintenance now and reports how many records retired and why (the same rules also run at the end of every turn)
/memory-harvest [--retire <id>] lists automatically harvested candidates, or retires one
/memory-audit <root> [--out <dir>] audits an archive of stores and writes four reports
/memory-import <root> [--selection <file>] [--apply] dry run by default; writes only with an explicit --apply
/memory-gaps [<count>] lists the repeated failure shapes in this workspace, the raw errors, whether the store has anything related, and which record was written but did not prevent it. Statistics only: no injection, no records written

Why auditing and importing are not given to the model: they scan arbitrary directories and write to the store in bulk, so the blast radius is large and this framework is fail-closed throughout. The model's tool list therefore holds 5 (4 knowledge operations plus one argument-free read-only census), with no per-turn token cost.

/memory-preview and real injection share one function (src/digest.ts), so it cannot disagree with what you actually receive — a preview that drifts would have no reason to exist.

Configuration

Deliberately not ## 配置: tests/docs.test.ts reads the config table out of ## 配置 in the Chinese README (the document the plugin is actually configured from) and then compares the keys listed here against the same code, so the two tables cannot drift apart.

Key Default Meaning
enabled true master switch
dbPath $DSH_HOME/experience-memory/memory.db store location
residentMaxRecords 5 per-section record ceiling
residentMaxBytes 1536 hard byte ceiling for the whole digest (all sections together)
coreMaxRecords 2 core-layer record ceiling; 0 turns the core layer off
standingMaxRecords 3 standing-rule ceiling; 0 turns the standing layer off
standingMaxBytes 768 byte ceiling for the standing section alone (the whole digest still respects residentMaxBytes); measured, that holds about three short rules or two long ones (lines cap at 240 bytes, the label costs 49)
recallMaxBytes 16384 per-recall byte ceiling
defaultDomain '' fixed domain; empty means infer it
maintenanceBatchSize 32 records processed per maintenance pass
failStreakLimit 2 consecutive failures before retirement
harvestEnabled true harvest candidates automatically at the end of each turn
harvestBroad false enable the broad criteria that measured unreliable (broad statements, recovered failures, changed goals)
harvestMaxPerTurn 1 candidates per turn (0 turns harvesting off)
harvestPoolLimit 200 candidate pool ceiling; the oldest is retired past it
harvestCandidateTtlDays 14 days a candidate survives unconfirmed and unretrieved
precallEnabled true deliver a lesson about the thing a tool call is about to do
precallMaxPerSession 20 hints per session (counted as deliveries actually made)
precallCooldownMinutes 30 minutes before the same record may be delivered again
failureTracking true count repeated tool failures in this workspace (counting only: no injection, no records)
failureShapeLimit 200 failure shapes kept per workspace; the rarest and oldest are evicted past it
disabledPresets [] which modes have no memory at all (by preset id). A listed preset gets no injection, no hints, no harvesting and no counting, and its tool calls are refused
anchorCostTable true check a declared anchor against what it would cost before storing it: an anchor that matches too many real calls is dropped and the caller is told (the record is still written; digest and recall are unchanged)
anchorCostMaxHits 300 hits above which an anchor counts as too common. Defaults to the pre-registered per-record gate (<300), so one number governs both
effectWeight 0 how much a measured deletion effect (remove this record and see whether the outcome changes) moves the ranking. 0 means not at all: the effect is recorded and shown but ranks nothing. It should only move once docs/GROWTH.md G5's experiment passes (same byte budget, decision-loss retention beats reuse-count retention by ≥5 points)
decisionLossRetirement false whether a record a deletion test measured as not changing the outcome may be retired for that reason. Off by default, and only a record that was actually measured (effect not null) can be retired this way — unmeasured is not the same as useless
guardHints true when a record describes an action that cannot be undone (wiping/clearing/overwriting), memory_remember answers with a computed line: which file the live store is, and where throwaway stores may live, so the rule can read "refuse by default, never touch this file" instead of a marker list the real store's own path may satisfy. Proposed, never written — the record's text is stored unchanged

An invalid value raises at load time and refuses to start the plugin instead of degrading silently. Exactly three limits may be 0: coreMaxRecords (0 = core layer off), standingMaxRecords (0 = standing layer off) and harvestMaxPerTurn (0 = stop harvesting). For every other limit, 0 is indistinguishable from "off", so the minimum is 1.

Model Experience

The per-turn experience digest

When a request is assembled, the plugin renders at most three sections: standing rules (records written with standing, at most standingMaxRecords, carried whatever the turn is about, with their own standingMaxBytes ceiling), domain-level experience corroborated across projects (the core layer, at most coreMaxRecords records), and relevant experience retrieved with the last two user messages as the query (the query layer, at most residentMaxRecords records). All three share the one 1536-byte hard ceiling, so the real line count is usually decided by the byte budget first — under the shipped configuration the record ceiling is 3+2+5=10 lines. One line per record: - [id] 标题 — 教训. When the standing section is cut by its own byte ceiling, the digest says how many rules were left out and names the setting to raise: this is the layer that promised to be present every turn, so it may not quietly lose one.

It is not added to the system prompt. The contribution of ctx.systemPrompt.context is composed by DSH into the "runtime context snapshot", and that snapshot is delivered to the model as a plugin-sourced message (source.kind === 'plugin', plugin dsh-system-prompt, form snapshot). That is not a detail: precisely because this text travels the same channel as the user's words, the plugin's query derivation and evidence grading must both skip plugin-sourced messages (one place each in src/digest.ts and src/evidence.ts), otherwise the digest would be read back as something the user said and the same few memories would reinforce themselves — the road Mem0's production store took to 97.8% noise. Both skips are verified against real session logs.

Token effect

The digest has a hard ceiling of 1536 bytes and costs 0 bytes when both sections are empty (no empty section is emitted); the core layer does not raise the ceiling, it only reallocates it. One further line appears unconditionally — the record guidance (204 bytes, RECORD_HINT) — outside that 1536 budget, because an empty store renders a 0-byte digest and that is exactly the case the guidance exists for. Its size is pinned by a test at a ceiling of 256 bytes, so it cannot grow unnoticed.

KV Cache effect

Content changes only when the hit set actually changes, so the effect on prefix caching is limited to the turns where it does. The core layer is stable, which makes it the cache-friendliest part.

Does it actually work: two controlled experiments

Everything above describes the mechanism. This section answers one question: did it actually prevent a mistake.

The method is two arms of N real model turns each — same repository, same task, one arm with the lesson in the store and one without. The conclusion is judged from the artefacts (file contents, file locations), never from what the model said.

Scenario one: what the config must say (docs/DELIVERY-GAPS.md §19, §21)

A convention that exists only in one thing the user said: a deploy config must declare the vault path first and the namespace second, and the order matters. No file in the repository says it.

fully correct 95% interval
no memory 0 / 18 0.0% – 17.6%
with the lesson 14 / 18 54.8% – 91.0%

Fisher's exact test, two-sided: p = 0.000002. Split out: single turn 0/12 versus 9/12 (p = 0.0003), cross-session 0/6 versus 5/6 (p = 0.0152 — one session hears the sentence and records it itself, a later session uses it).

Scenario two: where the file goes (§22 of the same document)

The convention becomes a placement rule: sample configs live in conf/samples/. The repository does not even have a conf/ directory.

in conf/samples/ 95% interval
no memory 0 / 6 0.0% – 39.0%
with the lesson 6 / 6 61.0% – 100.0%

Fisher's exact test, two-sided: p = 0.0022.

Both scenarios together: 0 / 24 without the lesson, 20 / 24 with it (p far below one in a million).

The detail that matters most: those 24 control runs are not failures to act — almost every one of them wrote a file, just wrong content or wrong place. Scenario one's control arm produced a plausible-looking deploy config of 695–1751 bytes and was wrong 18/18; scenario two's control arm wrote a file 6/6 times and put it in config/ 6/6 times — the model's own default guess for where sample configs live. So the difference is not whether it acts. It is whether what it writes is right and where it puts it.

The scope, stated plainly: this holds only for knowledge that lives in a conversation and not in a file. If the repository says it, the model reads it and memory is not needed; if nobody said it there is nothing to recall. What this framework is for is exactly the class of things with no second place to look — which is what separates it from reading documentation.

Automatic harvesting: catching the lesson the model never thought of

Whether something gets remembered depends on the model choosing to call memory_remember. That was measured here: across five real sessions and about 5,900 tool calls, memory_remember was never called once until someone explicitly asked. The unconditional guidance narrows that gap, but the structural problem remains — a lesson the model never thought of has nobody to catch it.

At the end of each turn the harvester reads that turn (not the whole session), and on five kinds of "moment worth keeping" it stores one candidate, one per turn, by priority:

Signal Criterion What is stored
failure-recovered within one turn a tool errored and then the same tool succeeded tool name + the raw error text
user-correction the user contradicted the previous turn (不对/错了/其实…, "no", "that's wrong", "actually"…) the user's verbatim sentence
user-statement the user stated an explicit durable rule (以后/一律/禁止/never…) the verbatim sentence
user-statement the user was not asking a question and named something specific (identifier/path/version/number/conclusion word) the verbatim sentence
goal-changed / action-refused goal/change; approval/decided and not allowed the new goal verbatim / that it was refused

It is not a judge; it only picks up verbatim speech. The criteria recognise a "moment", not a "lesson": what is stored is the verbatim sentence plus a mechanical title. Distilling a sentence into a claim is judgement, and the harvester has none — so it does not distil.

The broad criterion is the point: imperative-only matching would miss the shape lessons usually arrive in — "so the bug was because…", "this API does not fire in 1.5.2", "it turned out --preserve-symlinks was needed". None of those are commands.

Three properties keep it from becoming the thing this framework most wants to avoid:

  1. Always a candidate. Harvesting writes to the store directly and does not go through remember, so it can never be handed a grade out of thin air. It is a candidate by construction, the resident layer never looks at it, and the only way to promote it is for the model to restate the same sentence, at which point the normal evidence gate applies. That is what the test pins — with a harvested record whose quote is itself a user's verbatim sentence (which ordinary grading would call verified-user), it still has to stay a candidate.
  2. It infers nothing. All five criteria read markers the session has already written; origin and harvest_signal record which one fired, so it is auditable.
  3. It costs no LLM call. This plugin never spends one.

The bounds are invariants, not quotas: there is no daily allowance (the busiest days are the ones that teach the most, and a quota would quietly run out exactly when it is needed). Instead: at most 1 per turn, a candidate pool ceiling of 200 (the oldest is retired past it), and retirement after 14 days unconfirmed and unretrieved. That last one also closes an existing hole: maintenance used to scan confirmed records only, so candidates were immortal.

How a candidate is seen — otherwise harvesting just pours into a pool: the tail of a memory_recall response carries 另有 N 条自动采集的候选待确认 (only while the model is already looking at memory, so it costs nothing per turn); /memory-harvest lists them for a person and can retire one; memory_stats reports harvested / confirmed / pending.

The criteria were pinned against real logs, not against the event registry. The registry lists events this harness never emits: feedback/record is a known type, and in the busiest log in this workspace it appears 0 times in 11,735 events. That criterion was deleted before it was written — a criterion built on an event that never fires is a silent no-op.

And the criteria were calibrated on real logs, with the calibration deciding the defaults. Replaying the six largest logs in this workspace (235 turns):

Criterion Hits in 235 turns What sampling showed Verdict
user-correction 4 "not an implementation bug, my expectation was wrong…", "quant is quant, bigfat is value investing", "add a counter-example test: root=None must be rejected" acceptable precision (3 of 4), on by default
user-statement (broad) 105 skill directories, Objective: "…", Round: 5/256, questions, task requests about 5–10% precision, off by default
failure-recovered 5 (71 before the denylist) edit/write without reading the file first, old_string not found; the rest mostly

…

Content from the project README on GitHub ↗

Comments

Comments live in GitHub Discussions. Sign in with GitHub to post or react.