Install
Inside DeepSeek Harness, with dsh-market
dsh plugin --profile web add dshmarket
Or from the command line
dsh plugin --profile web add github:ZiYuan258/dsh-skill-router
Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.
README
English | 中文
Agent-driven skill discovery and on-demand loading. A Host plugin for DeepSeek Harness. It adds three tools — skill_search, skill_load, skill_ref — so an agent can find and load a skill from a library on demand.
dsh plugin --profile web add github:ZiYuan258/dsh-skill-router
Two things to check before installing.
1. The owner has to be
ZiYuan258. Eight repositories on GitHub share this name (this one is the newest, created 2026-09-23). Several of them auto-inject: they read the user's message before the model answers and put a skill's full body straight into the prompt.2. This plugin never auto-loads or injects a skill's body — but it does auto-suggest, deliberately. At the start of a task it may put a few candidate skill names (names only, never bodies) in front of the agent, and whether anything gets loaded is the agent's decision, including loading nothing at all. A comparison table of the same-named repositories is further down.
Only a dozen or two skills? This plugin is not for you. A single task usually uses just a few skills, so when the set is small, keeping them resident is cheaper — the catalog is DSH's native mechanism and it works. This plugin addresses the other situation: more skills than will fit in the catalog.
Do you need it
Two questions decide it: how big is the library, and how many does one task use.
| Your situation | What to do |
|---|---|
| Fewer than ~30 skills | Not needed. Keep them all resident; the catalog stays cheap and the built-in skill tool works |
| Tens to hundreds, a few per task | The typical case. Keep the frequent dozen resident; the rest live in the library |
| Several upstream repos, a thousand-plus skills | Where it matters most. All-resident is infeasible (1000 entries ≈ 150k characters per turn) but you do not want to lose access to any of them |
| A small library, but you want skills to trigger automatically | Not needed. That is pre-step routing, a different kind of plugin |
In one line: it is built for "a big library, a few skills per task". If you already keep your frequently used skills resident, its value is negative — you are paying for tool schemas and getting nothing back.
The technical problem it solves
DSH injects the session skill catalog into every model request (dsh-tool-skill emits a source.kind='skill-catalog' user message from its agent/pre-step listener). Therefore:
- Resident count costs continuously: every resident skill adds its name and description to every turn. A thousand skills ≈ 150k characters per turn, whether or not they are relevant.
- But the catalog is the model's only entry point: the built-in
skilltool resolves names throughctx.skills.list(), which only knows the roots the filesystem providers scanned. Anything outside those roots does not exist for the model — a library of 1000+ skills may as well be empty.
That leaves two options: put everything in the catalog (expensive every turn) or leave everything outside it (unreachable). This plugin offers a third: the library stays outside, and the agent searches and loads it itself when a task calls for it. The cost moves from "per turn, fixed" to "only when used".
Cost (measured, not estimated)
Tool-driven retrieval is not free. It costs in two places:
| Cost | Measured | Nature |
|---|---|---|
| Resident tool schemas | 3,603 B ≈ 1,001 tokens/turn | Fixed, independent of library size — the essential difference from catalog injection |
| One search round-trip | one tool call + about 1,533 B ≈ 426 tokens (limit=12) |
On demand; the extra step compared with picking from an injected catalog |
For scale: a 28-skill resident catalog measured 1,541 tokens/turn. Swapping 1,541 + 3,603 B for "six resident plus the tools" is where the arithmetic works — which is why a small library should not use this (see the section above).
Two ways to cut the round-trip: lower limit (or names_only: true, roughly 60% smaller), and call skill_load directly when the name is known, skipping search.
How it works
Data flow
user message
│
├─ resident skills (native DSH) ← catalog injected every turn, auto-triggered
│
└─ library skills (this plugin)
the model decides it needs specific knowledge
│
├─ skill_search ← keyword lookup over the index (reads no skill body)
│ returns: name / repo / copies / absolute path
│
├─ skill_load ← loads only the chosen ones as <skill_content> + base directory
│
└─ skill_ref ← reads one bundled references/ or scripts/ file, on its own
The key property: search and load are separate. skill_search only queries the index and never reads a skill body; skill_load reads only what you named. Searching 12 hits therefore costs metadata, and only the loaded skills cost their text.
The index is a contract, not a cache
The plugin reads <workspace>/.skill-src/skill-index.tsv — tab-separated, one row per skill:
| Column | Required | Meaning |
|---|---|---|
repo |
yes | upstream directory name, used for disambiguation and filtering |
relpath |
yes | path under that repo; with repo it locates the SKILL.md |
name |
yes | the skill name, and the key skill_load takes |
description |
yes | the retrieval corpus, and the text shown to the model |
files |
no | how many files the skill directory holds |
KB |
no | size |
whenToUse |
no | trigger phrasing, scored above the description (0 of 1025 SKILL.md files in the reference library carry it — absent is normal) |
Why TSV rather than YAML or JSON: the index is machine-generated, and TSV has the least structural ambiguity — whereas YAML is exactly what this project was bitten by in practice (a frontmatter block scalar >-/|- leaks its marker into the description). The parser is hand-written, because the plugin has zero imports and cannot use node:path or a CSV library; it follows CSV quoting rules, and it finds the header by shape rather than by literal text, so adding a column never breaks an existing file: a 6-column index keeps working, and a missing column reads as an empty string.
The index is located by walking up to 8 levels from the session working directory, so no drive or path is hard-coded. When it is absent, the tools say so and list where they looked.
Retrieval: weighted AND with an honest fallback
Scoring accumulates per keyword. The weight order reflects information density:
| Field matched | Weight | Reason |
|---|---|---|
name |
+100 | the skill name is the strongest signal |
whenToUse |
+40 | trigger phrasing by definition |
description |
+24 | descriptive prose |
path (repo/relpath) |
+6 | weak, but it rescues "find by repository" queries |
| exact name equality | +400 | an exact hit dominates |
Sort order: all-keywords-match → matched-keyword count → score → name hits → name length → lexicographic (deterministic: the same query always yields the same order).
The default is strict AND. Only when AND returns nothing and there is more than one keyword does it try a partial match, and a candidate must then hit every keyword but one — otherwise it answers "nothing found" and sets fallback: "weak". Partial hits carry matchCount, the result carries strict: 0, and even the header the model reads says "0 exact match(es); N partial match(es)": a near-miss is never dressed up as a real hit. A single-keyword query never falls back, because there is nothing to degrade to.
That threshold was forced by measurement. On the 1026-row reference library,
test setup config helperused to return 1026 entries, nearly all of them sharing one common word; tightened, it returns 7. In the same experimentmake a moviematched the whole library onmakeanda, so words with no discriminative power —a,the,make,use— are dropped during tokenization (STOP_WORDSintokenize).
With explain: true the score breakdown shows why something did or did not match:
- beta-gadgets [beta-skills]
why: widgets: -; beta: +130 (name+description+path)
Duplicates: determinism over guessing
Libraries assembled from several upstreams carry the same name more than once — one repo publishes every skill under skills/, plugins/<name>/skills/ and antigravity/skills/ (the reference library holds 5 copies of test-driven-development).
The handling: skill_search reports copies, making the ambiguity visible; skill_load accepts repo to disambiguate; with no hint it picks deterministically (shallowest path wins, so skills/<name> beats plugins/<x>/skills/<name>); and a repo filter that matches nothing lists the repos that do have the skill instead of silently falling back to another.
Lookup order and limits
skill_load searches the library (.skill-src) first, then the resident catalog (ctx.skills); the source field says which won. Limits: 8 names per call, 120,000 characters per skill body, 8 bundled entries listed. An oversized body is truncated and reported (truncated: true, plus the original length and the path to read in full) rather than silently shortened.
Path containment
skill_ref runs a containment check on the normalized path before any I/O, so ../ never reaches the filesystem. Known limitation: the check is lexical and not symlink-aware (the reference library has zero symlinks, so this is theoretical today). resolvePath is exported and unit tested directly — a security rule verified only through the tool is a rule that can quietly stop holding.
Install
Requires a DSH installation and a profile. dsh.engines.dsh declares >=0.1.5-rc.1 — that is the verified floor, not a claim that newer is required: the plugin uses only ctx.tools.register and ctx.fs.*, with no event hooks and no imports. Older versions have not been verified, and over-claiming compatibility would be worse than being conservative. ctx.get, ctx.effect and ctx.skills are all optional and degrade rather than crash when absent, which test/minimal-host.mjs pins.
# from git (recommended)
dsh plugin --profile <profile> add github:ZiYuan258/dsh-skill-router
# or from a downloaded release tarball / checkout
dsh plugin --profile <profile> add /absolute/path/to/dsh-skill-router
Restart DSH once, then confirm all three tools appear in the tool list. Uninstall:
dsh plugin --profile <profile> remove dsh-skill-router
No registry publication. DSH composes a plugin from any package it can resolve, and a git URL or local path is enough — so this package is
private: true. Release tarballs live on GitHub releases for offline installs.
⚠️ Check the name first: several repositories share it
Eight repositories on GitHub are called dsh-skill-router (this one is the newest, created 2026-09-23), so installing the wrong one gets you a different plugin:
| This repository | The other family | |
|---|---|---|
| Install command | github:ZiYuan258/dsh-skill-router |
github:lau4tin1/dsh-skill-router (holds the bare npm name dsh-skill-router) |
| Mechanism | tool-driven: the model calls skill_search / skill_load / skill_ref itself |
routing plus auto-injection: embeds tasks and skills with a local model, keeps the clearly-relevant ones by a gap rule, and puts their bodies in the prompt |
| Problem it solves | library skills are invisible to the model | the model does not use a skill it should have |
| Who chooses | the agent — it reads the candidates and may pick none | the plugin — a routing hit is injected |
| Dependencies | zero dependencies, zero imports | varies; some need a local model or embeddings |
The two take opposite positions, and that is this plugin's deliberate design rather than a gap. See the task-aware discovery section below: that layer automatically suggests a few candidate skill names (at tier === HIGH, one line of names and matched fields) but never auto-loads or injects a skill's body — a body only enters the context through skill_load, and that call is always the agent's.
Others in the family include MJorgin/dsh-skill-router (rule-first pre-step routing) and Phantomcyber-ai/dsh-skill-router (intent-level auto-routing); some are placeholders or unfinished. Check that the owner is ZiYuan258 before installing.
If a tool reports this repository as unreachable, check which kind of failure it is: GitHub rate-limits unauthenticated API calls to 60/hour and answers
403 API rate limit exceeded, while the web page and raw files keep working — anddsh plugin adduses those.
Usage
0. It works once installed (nothing to set up first)
The plugin ships with a starter library (5 skills covering search before guessing, evidence before claims, debug with evidence, scope before building, and report results clearly). When you have no library of your own, that is the library — restart DSH and skill_search has something to find immediately:
skill_search "debug"
→ debug-with-evidence ← from the starter library bundled with the plugin
The result carries starterLibrary: true and a library: path inside the plugin package. Seeing that flag means you are reading the starter library, not your own — the two situations call for completely different diagnosis, so they have to be distinguishable.
The starter skills never enter the resident catalog. They live under the plugin's
resources/starter-skills/and go throughskill_search/skill_loadlike any other library skill; not one byte reaches the per-turn model catalog. Putting them in.dsh/skills/would inject them into every turn — which would destroy the entire point of this plugin.test/starter-library.mjsasserts this.
Shipping 5 rather than 1,000 is deliberate: the problem this plugin solves is a large library that must not stay resident. Bundling a large one would bring package size, update lag, license mixing and version coupling all at once. For real coverage, attach your own library (section 1).
1. Where the library lives
The plugin walks up to 8 levels from the session working directory looking for .skill-src/skill-index.tsv, so the convention is to put the library at the workspace root:
<your-workspace>\ ← you start DSH sessions here
├─ .skill-src\ ← the library root; the name is what the plugin looks for
│ ├─ skill-index.tsv ← the index (generated in step 3)
│ ├─ remotion-skills\ ← one upstream repo = one top-level directory
│ │ └─ skills\remotion-create\
│ │ └─ SKILL.md
│ └─ trailofbits-skills\
│ └─ plugins\semgrep\skills\semgrep\
│ └─ SKILL.md
├─ .dsh\skills\ ← the DSH resident set (the plugin never touches it)
└─ AGENTS.md
Three rules:
- The directory must be named
.skill-src. The leading dot keeps it invisible to DSH’s skill scanner — that is precisely the mechanism that makes the library free. Name itskills/, or put it under.dsh/skills/, and DSH will inject every skill into every turn, defeating the point of installing this. - It must sit at or above the session cwd. If your session starts in
<your-workspace>\projects\foo, the plugin walks up and finds<your-workspace>\.skill-src— that works. - The layout inside does not matter. The plugin only needs "some directory named like the
repocolumn, then therelpathcolumn, thenSKILL.md". An upstream repo mixingskills/,plugins/<name>/skills/andantigravity/skills/can be dropped in as-is.
2. The library lives elsewhere (another drive, another directory)
The plugin looks for cwd/.skill-src, so when the library lives elsewhere, put a directory link in the workspace. Measured to work (Windows junction / POSIX symlink; verify with node tools/check-link-support.mjs):
# Windows: a junction needs no administrator rights
New-Item -ItemType Junction -Path "D:\work\.skill-src" -Target "E:\skills-archive"
# Linux / macOS
ln -s /mnt/skills-archive "/home/me/work/.skill-src"
skill_search and skill_load both work through the link. Known limitation: the returned path is the link path, not the real one — when debugging, use dir or ls -l to see where it points.
Do not use DSH's
customSkillDirsfor this. That setting registers skills as resident, which would put every skill into the per-turn catalog — the exact opposite of what this plugin is for.
3. Generate the index
Any writer will do, as long as it emits those columns. A reference implementation (PowerShell, for a library of upstream repos checked out under one directory):
node tools/build-index.mjs <library-root> # writes <library-root>/skill-index.tsv
node tools/build-index.mjs <library-root> --check # reports drift only (exit 1 when stale)
node tools/doctor.mjs --root <library-root> # check-up: path / count / stale / duplicates
Zero dependencies, cross-platform. It does four things: finds every SKILL.md, parses the YAML frontmatter, computes relpath relative to the repo root (not the library root — one extra level and nothing resolves, which is the bug its first version shipped), and writes real TSV escaping so a description containing a tab, quote or newline cannot break the columns. The whenToUse column is written only when at least one row has it.
After this you never hand-edit the index again. Re-run it when the library changes; if you forget, doctor says so — it compares "on disk" against "in the index" in both directions.
That block was a script you had to copy and edit a $root in. It worked, with two problems:
- the index format is the plugin's internal data model, and it should not be part of the install flow;
- it scanned one level only (
$root/<repo>/**), never emittedwhenToUse, and missed skills without frontmatter — the reference library had 3 of those (in a Microsoft monorepo), so they could never be found. Measured:build-index.mjsproduces 1,028 rows for that library while the old reference implementation's index held 1,025.
The index is still a public contract: anything that emits those 7 columns can replace build-index.mjs. The contract is the table above.
4. Verify it works
The install command is in the section above. After restarting DSH (a plugin row is composed only when a new host process starts), check in this order:
- Are the tools there? Ask the agent which skill-related tools it has;
skill_search,skill_loadandskill_refshould all be listed. - Did it find the index? Have the agent run
skill_searchfor a name you know is in the library. The result carries alibraryfield — the root it actually used. Check that path; it is the fastest way to catch "it looked in the wrong place". - Does loading work? Have it
skill_loadone hit.sourceshould readlibrary(aresidentvalue means it came from the resident set, not the library). - Did the library leak into the catalog? Confirm a new session's skill catalog did not grow because of this install. Anything under
.skill-srcshould be absent from it.
Two fields separate the three failure modes: no index (the error lists where it looked), an index that is not this library (the library path is wrong), and a malformed index (a non-empty error). explain: true additionally shows why a search did not match.
Checking the index itself by hand:
# header (6 or 7 columns) and row count
Get-Content "D:\work\.skill-src\skill-index.tsv" -TotalCount 1
(Import-Csv "D:\work\.skill-src\skill-index.tsv" -Delimiter "`t").Count
5. Day to day
You do not need to remember skill names. That is the point — the agent does the choosing:
- Just describe the task ("build me a video with Remotion"). The tool descriptions say to search before any non-trivial task, so the agent looks on its own.
- To see what is available, ask: "does your skill library have anything about X?" The agent will run
skill_searchand show you. - To make a skill auto-trigger (no reminder needed each time), that skill has to become resident:
The price is its entry in every turn's catalog — that is what you are buying.& "D:\work\.skill-src\install-more.ps1" -Name remotion-create
6. When the library changes
| What you did | What to do |
|---|---|
| Added, removed or renamed a skill directory | Regenerate the index, or the search works from stale metadata |
Edited a SKILL.md description |
Same — the description is the retrieval corpus |
| Edited only a skill body | Nothing; skill_load reads files live every time |
| Moved the whole library | Update the link target; the old index's repo/relpath values stop resolving |
Keep the generator as a script (for example .skill-src\scan-skills.ps1) and run it after changing the library. A stale entry shows up as "searchable but unloadable": the path is still in the index, the file is gone. skill_load names the path it could not read rather than failing silently.
7. Backup and quarantine
- Back the library up. It is your capability set, usually assembled from several upstream repos. If those repos are re-cloneable, the minimum is
skill-index.tsvplus your admission notes. - Quarantine first, investigate second. Move a suspicious directory out of the library root (say
.skill-src\_quarantine) and regenerate the index; it disappears without uninstalling anything. For an audit,node tools/audit-library-risk.mjs <library-root>counts risky patterns separately for code blocks and prose — only the former is what a model may copy and run.
Tool reference
skill_search
Keywords are lowercased and matched with weights across name / whenToUse / description / path.
| Parameter | Type | Notes |
|---|---|---|
query |
string, required | e.g. "kubernetes helm", "remotion video" |
limit |
integer | 1–40, default 12 |
repo |
string | case-insensitive filter on the upstream directory name |
names_only |
boolean | names and repos only, no descriptions (about 60% smaller) |
explain |
boolean | also return the score breakdown per hit, for diagnosis |
Returns total, strict, shown, more, fallback, and per hit: name, repo, description, copies, matchCount, stale, whenToUse, files, path, libraryRelative; with explain also score and why.
stale: true means the index has the row but its SKILL.md is gone — the library changed and the index was not regenerated. Such rows do not make the search fail (an earlier version threw ENOENT, so one stale row took down the whole retrieval); they are flagged instead, and the flag is visible to the model.
skill_load
Loads the full text of one or more skills into context.
| Parameter | Type | Notes |
|---|---|---|
name |
string | one skill name; an absolute path to a SKILL.md also works |
names |
string | several names in one call, separated by commas or newlines |
repo |
string | upstream repo filter, applied to every name in the call |
Returns requested, loaded, failed, and a skills[] array of name, source, repo, copies, path, resourceDir, content, referenceFiles, truncated, error. The tool card renders each skill as a <skill_content> block with its base directory, so relative paths (scripts/, references/, assets/) resolve correctly.
Why
name+namesrather than an array. An earlier version declarednameasoneOf: [string, array]. It reads well in a schema and fails in practice: array arguments can arrive at the tool stringified, so["gh-cli"]becomes the literal'["gh-cli"]', matches the string branch, and is treated as one nonexistent skill name. Three shapes are accepted now (a real array, a JSON string, comma/newline-separated) because robustness here should not depend on how the transport serializes arguments.
skill_ref
Reads one file bundled with a skill, or lists what is bundled.
| Parameter | Type | Notes |
|---|---|---|
name |
string, required | a skill name skill_search returned |
path |
string | file relative to that skill's base directory, e.g. references/rulesets.md |
list |
boolean | list every bundled file instead of reading one |
repo |
string | upstream repo filter, for a name that exists several times |
The referenceFiles list from skill_load is usually enough to decide whether to read a file — this reads only that one instead of pulling the whole directory into context.
Skill usage tab (conversation.view)
The plugin ships a Client half that adds a 技能 / Skills tab to the conversation view ring — the row holding Chat, Trajectory, Approval and Context — listing what this session actually loaded:
技能调用清单
本会话共 17 次技能调用,涉及 10 个技能。其中 1 次调用一次点名了多个技能,故按技能名分行列出;
1 cordis·插件·开发 (cordis-plugin-development) skill 第 1 轮
2 editing-cordis-compositions skill 第 1 轮
3 remotion·创建 (remotion-create) skill_load 第 1 轮
…
16 验证·前置·完成 (verification-before-completion) skill_load 第 63 轮
17 系统化·调试 (systematic-debugging) skill_load 第 67 轮
18 系统化·调试 (systematic-debugging) skill_ref 第 67 轮
已读到本会话最早一条记录,上面的数字是完整的。
dsh-skill-router v1.15.4 · 第 43 页 · 已读完 · 可翻页 是
Zero model tokens. The data comes entirely from the session ledger, handed to the component by the session-scoped slot:
// The registration. `conversation.view` is declared scope: "session", so the renderer calls inject
// with the scope binding's key and spreads the result over the component's props. (That is also the
// mechanism behind following a session switch: injected props are cached per scope.)
inject: (sessionId, binding) => {
const b = ctx.get('sessions').binding(sessionId ?? binding?.key)
return { source: b.eventSource, session: b.session }
}
| You might assume | The actual contract |
|---|---|
eventSource can page |
❌ SessionEventSource = ObservableSnapshot<SessionEventWindow> — only getSnapshot() and subscribe() |
| Then how is older history read | ✅ loadOlder(): Promise<void> is on the session (SessionFace extends ISession), which is also how the shipped trajectory tab calls it |
Subscriptions are cancelled with unsubscribe() |
❌ No such method; cancelling is the return value of subscribe(fn) |
Where turn / callId come from |
✅ { type: 'tool/call', seq, time, data: { turn, step, callId, name, arguments } } (measured, not inferred). Identity is the envelope's seq, not callId |
It reads the whole history, and says only what it knows. The ledger window is bounded (observed ~1,664–1,900, seen resetting 3,336 → 1,664) and hasMore is true on the newest page, so page 0 is the recent end — a skill loaded five pages back is invisible until paging reaches it. The tab backfills to the oldest record in the session (43 pages in practice), and:
本会话共 N 次技能调用…is printed only once the oldest record has been reached; until then it says how many names it has read and that older records remain;- with no ledger to read it says only that and prints no count at all (missing service / no binding / no eventSource are three different sentences);
- paging can stop for three reasons and they are stated separately — no answer (a 4-second deadline), an answer that moved nothing, or the 200-page cap. "I gave up on the rest of the history" and "the history ended here" are different claims.
Why the row count and the call count can differ. This is correct, not double counting:
- the row key is
event identity + normalized nameand the call count groups by event identity, so one call naming several skills becomes several rows (in practice a singleskill_loadloaded bothcode-review-and-qualityandgh-cli); - the event identity is the
seqon the event envelope (SessionEventdeclaresseq: SessionSeqon every event, so it is contractually unique, and onetool/callevent is one call).callIdis the pairing id between a tool call and its result, and nothing in the contract says two different calls cannot share one — deduping on it would silently merge two real calls into one row and report a quietly low count. SocallIdis only part of the fallback token for an event that carries no numericseq; - the call count is derived from the final rows, never accumulated alongside them — two sources for one fact drift, and that is exactly what a live report's "18 rows / 17 calls" forced into the open;
- when the two differ, the header explains why, so nobody has to stare at the numbers.
The order is the session's, not the arrival's. Events arrive newest-first and older pages are prepended, so ordering by arrival gives you the reverse (a live report had turn 31 above turn 5). Records are sorted ascending by the event's own seq, whatever order the pages happen to arrive in.
Paging is judged by whether the window reached further back — not by the promise, and not by its length. The real session.loadOlder() silently does nothing in several conditions (a session still opening, events not yet arrived, a concurrent read) and returns an already-resolved promise, so "resolved" does not mean "a page arrived". The judgement is:
the OLDEST seq in the window decreased <- primary: only acquiring older history can do this
or
the window grew <- secondary: a live append can grow it too
Length alone is wrong: the window is bounded, so it can slide — constant length while the whole content moves older. Such a page was judged "no progress", and after two of those the tab gave up with "无进展停止", reporting giving up as nothing there. test/usage-tab.mjs pins this with a true sliding window (constant capacity 40, advancing 20 per page) that hides a skill in the older history.
Every page is folded into the accumulator the moment it is read. This matters more than the judgement: the accumulator used to be written only when the ledger notified, and a notification can be a long time coming. That produced a successful read that was never kept — loadOlder() brought a page into the window, the UI rendered it, it slid out before the next notification, and the accumulator never saw it. Pages are now folded in while they are still on screen. Stopping the paging does not stop the live tail — calls arriving later still show up immediately.
The last line says which build you are running. The Client half is served with cache-control: immutable, cannot be imported by a Node test, and its served bytes sit behind the Desktop capability check — so "which build is the browser running" used to be unanswerable. It is now printed in the tab: dsh-skill-router v1.15.4 · 第 43 页 · 已读完 · 可翻页 是.
The Chinese name is display only. A skill name is the match key for skill_load, for index search and for the /skill command, so:
- what actually gets called is always the English name; the Chinese form never leaves the render layer;
- search still runs against the English text, and neither
SKILL.mdnor the index is changed by a single byte; - proper nouns (
azure,vercel,semgrep,figma…) are left alone — the most frequent tokens in this library's names areazure(148) andgoogle(44), and translating those only makes a name harder to recognise.
The translation is a glossary plus a proper-noun allow list, not 872 hand-written pairs: phrases first (best-practices → 最佳实践), then single words (troubleshooting → 故障排查), with filler words (and, from, the) dropped.
This tab used to be empty, and the reason is worth keeping. It read the wrong source four times: a guessed node shape; per-turn tool declarations from request headers (counting skills merely offered to the model as loaded);
useChat().legacy.nodes(measured at one instant: 210 nodes with zero tool calls, against a ledger holding 2,778+ events); and paging written againstsource.loadOlder, which lives on thesession— so the guard returned on the first line every time and four "fixes" changed a code path that never executed. None of the four crashed and all four rendered a plausible list, which is exactly why the data contract had to be measured rather than inferred. The retrospective is indocs/release-notes-v1.8.0.md.
Task-aware skill discovery (HIGH injection — an experiment with one variable)
The three tools above answer "there are many skills — how does the agent find one". They do not answer the other half:
Will the agent think to look at all?
A library skill is invisible to the model, so using one requires the model to first remember that searching is possible. If the user says "run a Semgrep security audit" and the model decides to just answer, skill_search never happens. This layer exists for that gap.
It hooks agent/pre-step — DSH's waterfall that runs before a request is assembled — and at step 1 of a turn ranks the library for the incoming task, locally, with zero model calls:
the user task (step 1 only)
↓
tokenize (the same tokenizer skill_search uses)
↓
STOP_WORDS (language-level noise)
↓
**corpus-frequency filter**: a word carried by > 80% of rows is dropped (under 20% discrimination)
↓
scoreRow (the same scorer skill_search uses, weights in one place)
↓
**dedupe by skill name**: several copies take one slot, best score kept
↓
top 5, or an explicit "nothing"
↓
tier computed over the **deduped effective tokens**
↓
tier === HIGH -> one line put in front of the model; otherwise nothing happens
Why the scorer is shared and the query semantics are not. skill_search's strict AND is built for a short query written by a model; a task is prose, and "分析这个 React 项目的性能问题" has no interpretation under strict AND — which would make this layer fail silently. So the difference stays in the caller and the weights stay in one function.
Why injection is on, and why HIGH only
What was measured before this shipped: across 273 turns the library was almost never searched, and every record said injected: false. The hypothesis under test is therefore narrow and causal:
The agent does not fail to use skills — it is never told which ones are worth considering.
Testing that needs exactly one changed variable, so this version does one thing: when tier === HIGH, the candidates are put in front of the model. The retrieval policy — strict AND, the tokenizer, the tier thresholds — is unchanged, so a change in behaviour can be attributed to the hint rather than to two edits at once.
The cost is why it is worth trying:
| already paid every turn | tokens |
|---|---|
| the resident skill catalog | ~3,238 |
| the three tool schemas | ~1,001 |
| injecting 5 candidates (new in this version) | 532–549 bytes ≈ 148–152 (measured on a real library) |
What gets injected
Maybe relevant skills for this task: semgrep — Runs a Semgrep security scan over a codebase: detects langu…; code-review-and-quality — Conducts multi-axis code review. Use before merging any cha….
Load any that fit with skill_load, or ignore this and continue without one.
The name plus a real description, capped at 60 characters. And it says explicitly that ignoring it is fine: this layer discovers, the agent still chooses.
This slot used to read
(name, description, path)— which is not a description, it is the list of fields that matched.In a measurement over 467 candidates, 304 matched all three fields, so nearly every line rendered the same parenthetical: a literal. The model got five unfamiliar names with no statement of purpose, which left it nothing to judge relevance by; the observed behaviour was to continue with
read/edit/pwsh, and telemetry showed 14 injections with zero library loads following them.The code meant to show what the matched fields say and literally rendered their names instead. So this is not a feature, it is a rendering bug fixed: the hint goes from zero information back to information. The cost is ~57 more bytes per injected turn (329–341 → 532–549), and what it buys is the end of an injection that was paid for every turn and returned nothing.
Two implementation details, both verified in the harness rather than assumed:
- It goes into
decision.messages, not throughagent.inject().preStepcallsinbox.claim()before dispatching the waterfall, so an injected message is only claimed at the next step — and this layer has to work on the first one.decision.messagesis the authoritative batch for the current step. - The injected message carries a unique
id. The shape is copied from the framework's owncreateUserMessage:{ role, content: [{type:'text',text}], source: {kind}, id }. The id is not decoration — framework messages always carry one, and two injections sharing an id would be indistinguishable downstream.
Debug switch: write the task's tokens into telemetry
By default only counts are recorded, never the words — a keyword can carry a project name, a customer name or a vulnerability id. Turn it on only when investigating candidate quality:
# on the row in cordis.patch.yml
- insert:
- id: skill-router
name: dsh-skill-router
config:
discovery:
debugTokens: true
Or the environment variable DSH_SKILL_ROUTER_DEBUG_TOKENS=1 (both paths kept, because whether cordis hands a mounted row its config is not something this repo has verified). With it on, records gain a tokensUsed field.
Telemetry now records two things, because they answer two questions
{"at":"…","turn":3,"step":1,"tier":"HIGH","reason":"ok","tokenCount":7,"indexRows":1025,
"candidateCount":5,"candidates":[…],"arm":"treatment","injected":true,"hintBytes":330}
{"at":"…","kind":"turn-calls","turn":3,"tier":"HIGH","arm":"treatment","injected":true,
"skillSearchCalls":1,"skillLoadCalls":1,"skillRefCalls":0,"residentSkillCalls":0,"otherToolCalls":7}
The first is written at step === 1, when nobody can yet know whether the turn will search — so "did the hint make the agent search" has to be answered by the second, and the two are joined by (sessionKey, turn) — not turn, since 14 turn numbers were measured duplicated across sessions. The second records counts only: no tool arguments, no skill bodies, no user text.
Per-turn counting settles at the next turn's first step, because this harness has no hookable end-of-turn waterfall (verified: agent/turn-stopping does not exist in 0.1.7-rc.2, so hooking it would have been instrumentation that never runs). A turn abandoned mid-flight loses its numbers rather than misattributing them — the right failure direction for a measurement.
Randomised arms: why "not injected" is not a control, and why the unit is the SESSION
An earlier version treated non-injected turns as the control. They are not one. Injection is decided by the tier, so the two groups are different task populations by construction:
injected: run a semgrep security audit -> clear candidates -> injected=true
"control": what is the weather -> no candidates -> injected=false
A difference in search rate between them is attributable to the tasks being different, not to the hint.
The question is far narrower:
Among tasks the system already considers high quality, does merely showing the candidates change what the agent does?
And the intervention is durable, which decides the unit. dsh-agent-loop lands the step's messages on the session like this:
this.session.append('user/message', message, { surfaceOp: 'append' })
surfaceOp: 'append' means the hint joins the session surface, and every later step derives its history from that (observed directly: an injected hint appears in a session projection as a skill-router surface node of about 90 tokens).
So per-turn arms are wrong, and wrong in one direction only:
turn 10 treatment -> hint enters the session surface
turn 11 control -> no new hint, but turn 10's hint is still in context
Treatment contaminates every later control opportunity while control never contaminates treatment — a bias in a known direction, with numbers that look entirely normal.
So the unit is the session:
session
↓
sha256(sessionKey) -> roughly 50/50
↙ ↘
control treatment
no HIGH in this every HIGH in this
session gets a hint session gets one
A control session never receives a hint at all, so nothing of the intervention can leak into any of its opportunities.
The cost: the two arms are now different conversations, so they may differ in capability or context. That is the trade once the intervention persists — matching within one conversation is only worth having if the arms are still independent, and per-turn arms are not.
The unit of measurement: each session's FIRST HIGH
Even with a per-session arm, later opportunities inside one session are not independent samples: the agent searched once, learned a skill, and its later turns are shaped by that. So the primary metric counts only each session's first eligible opportunity — the one observation that provably precedes any hint this experiment could have shown, because at that point nothing has been injected.
The plugin marks firstEligible on every record; the readout reports the primary metric over those, and lists the exploratory (same-session later) and pooled numbers separately, labelled as unusable for causal conclusions.
The experiment protocol
① restart DSH, record the start time T0
② freeze skill-index (usable skills should stay 1025; if it changes, this batch is void)
③ **do not resume an existing session** — but what is excluded is **residue**, not age (see below)
④ collect >= 50 SESSIONS, each with exactly one pairable first observation (roughly half per arm)
⑤ node tools/discovery-report.mjs --since <T0>
Note that step ④ counts sessions, not turns: the gate is 50 sessions each holding one pairable first HIGH opportunity.
--since is required: legacy records have no sessionKey and are reported as "unreliable pairing", but only the analysis window actually keeps them out of the denominator.
What step ③ actually excludes: residue, not age
firstEligibleSeen is plugin process memory and is empty after a DSH restart, while a DSH session is durable and resumable. Together:
before the experiment session A already received a skill-router hint
↓
restart DSH firstEligibleSeen is cleared
↓
resume session A
↓
its next HIGH -> firstEligible = true <- but A already has intervention history
The contamination did not disappear — it moved from across turns to across processes.
What makes an observation invalid is that it was exposed, and "the session was created before T0" is only a proxy for that. The proxy errs in both directions:
| Reality | What the birth-time rule does | |
|---|---|---|
| false positive | created before T0 but never injected: not one hint in its history, so the observation is clean | wrongly excludes it |
| false negative | created after T0 but already injected inside this same process | the rule never sees it |
So admission judges the contamination itself: the session must have no injected: true record before T0.
primary = unique sessionKey
∩ firstEligible === true
∩ paired
∩ **no hint delivered before T0** <- the actual criterion
∩ sessionCreatedAt is known
∩ armViolations === 0
∩ duplicateFirsts === 0
∩ a single indexRows
That set has to be computed by the readout before it trims the log at T0 — once trimmed, "was this session injected before T0?" can no longer be asked.
The report prints admitted observations in two classes (created after T0 / created before T0 but never injected), so you can see how much a clean old session contributed instead of having that hidden behind a blanket "created after T0".
A session whose creation time cannot be established is still excluded: a session with an unreadable birth time could be new or could be a resumed old one, and admitting it on a guessed time defeats the gate. Exclusions are reported in two classes (pre-T0 residue / time unknown) and are not hidden from the exploratory and pooled numbers.
With no --since, the gate is off and the readout says so.
Three success metrics, in causal order: ① does the agent start calling skill_search after a HIGH hint (the trigger hypothesis) → ② does it then actually skill_load (the candidates produced behaviour, not just a glance) → ③ is the loaded skill relevant to the task (manual sampling).
50 sessions is the gate for deciding what to do next, not proof the project works
A proportion read alone hides how small n is. 8/25 and 20/25 both yield a Δ, but the strength of evidence differs by an order of magnitude. The readout therefore prints, per arm:
treatment n=25 searched 20 / not 5 -> 80.0% 95% interval [60.9%, 91.1%]
control n=25 searched 4 / not 21 -> 16.0% 95% interval [6.4%, 34.7%]
delta +64.0 percentage points
intervals overlap no -> the direction is credible (still not an effect size)
Three deliberate choices:
- A Wilson interval, not the naive normal approximation
p ± 1.96·sqrt(p(1-p)/n). The latter breaks on values this experiment will actually hit: with 25 sessions and zero searches it returns[0, 0], reporting "not observed" as "cannot happen", and one search in 25 gives a negative lower bound. Wilson stays interpretable there, for two extra lines. - "not searched" and "not loaded" are explicit fields, so no reader has to subtract.
- Metric ② reports two denominators, because they answer different questions: the conditional rate (denominator = sessions in that arm that searched) answers "having searched, will it actually use something", and the full funnel (denominator = all first observations) answers "what share of opportunities end in a load at all". Reporting only the latter mixes in "never searched" and reads like "searched but refused to use".
This is not a statistical model — it turns "how uncertain is this proportion at this n" into one number.
Two intervals not overlapping is not a significance test. That is a description, not a conclusion. The readout therefore prints the raw 2x2 table and computes no p-value by default — testing belongs to the analysis stage, a product should not bind one particular test into its output, and no reader should meet a p-value on every readout that gets mistaken for a verdict:
④ raw 2x2 table
search no search total
treatment 20 5 25
control 4 21 25
(copy-paste line: a=20 b=5 c=4 d=21)
Ask for it explicitly with --fisher (two-tailed Fisher exact, zero-dependency, validated against the textbook case 3/3 vs 0/3 -> p=0.1). At small n it is safer than a normal approximation. No correction for multiple comparisons within one readout, and the p-value answers only "do these data look like one distribution" — it is not an effect size.
What each outcome would mean next
| Outcome | Reading |
|---|---|
| LOW HIGH relevance | go back to discovery |
| High relevance, large Δ, intervals disjoint | the original trigger hypothesis holds: the agent does not fail to use skills, it never entered the decision space |
| High relevance, treatment ≈ control ≈ 0 | a valuable negative result: showing the candidates is not by itself enough to trigger a search → study the hint's wording and placement rather than the retriever |
| Both arms searched, low load | the skill_search → skill_load step is the problem |
| treatment searches and loads | the discovery layer works; Chinese/AND/cost come next |
**One conservative behaviour already observed (not changed here, logged for the data)
…
Comments
Comments live in GitHub Discussions. Sign in with GitHub to post or react.