安装
在 DeepSeek Harness 里通过 dsh-market 安装
dsh plugin --profile web add dshmarket
或使用命令行
dsh plugin --profile web add dsh-answer-reviewer
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络。请先审阅源码,并尽量锁定 commit(github:owner/repo#sha)。
README
该插件的 README 只有英文版本。
dsh-answer-reviewer
A dsh host plugin: every time the agent is about to close a turn, the
plugin extracts the assistant's final user-facing text and asks a dedicated
review model to grade it on a 1-100 scale. If the score is below the
configured threshold (default 80), the plugin steers the agent back
into the same turn with the reviewer's single most important finding as
concrete feedback. Once a turn has been steered maxChallenges times
(default 5), the agent's reply is allowed to close the turn as-is — the
cap is a leak guard, not the real gate; the score threshold is.
Tune the gate, switch the reviewer on/off, change the max-challenges cap,
or pick a dedicated review model — all live, no host restart — from the
Reviewer 配置 tab in the conversation view ring above the composer,
from the optional sidebar tab, from the self-hosted config UI on
http://127.0.0.1:3987, or via the same JSON API from curl. Every reviewed answer then shows its own
score chip in the message action row (see
Score chip). The on-disk config lives at
~/.dsh/answer-reviewer.json (override with REVIEWER_CONFIG_PATH);
turns of the host go through with the defaults if the file is absent.
The review model sees every real user prompt the user has issued in
the session (any user/message event whose source.kind is not
'plugin'), tagged <all_user_prompts>, plus the assistant's final
reply under <reply>. Plugin-injected context (file-change notices,
AGENTS.md, skill content, system reminders) is filtered out so the
reviewer scores against what the user actually asked, not against the
context the harness happened to be carrying.
The review model is independent from the agent's working model when
reviewProvider / reviewModel are configured; otherwise the plugin
falls back to the agent's current route so it works out of the box.
The score is not just an internal gate — it is shown on the answer it graded:
⤷ ⧉ ↻ 👍 👎 ● 评分 92 ← in the message's action row
Score chip (since 0.7.0)
Every finalized assistant message carries the score the reviewer gave it, as a small tinted chip in the message's action row, next to the shipped copy/retry and Like/Dislike buttons:
| Verdict | Chip | Meaning |
|---|---|---|
| Cleared the gate | ● 评分 92 on a green tint |
score >= threshold; the turn closed on the first look |
| Below the gate | ● 评分 61 · 第2次 on an amber tint |
Steered back for a retry; the number in the chip is the attempt that finally counted |
| Out of retries | ● 评分 55 · 第4次 on an amber tint |
Still below the gate and maxChallenges was spent, so nothing was pushed back (since 0.7.1) |
Hover the chip for the gate, the verdict and the reviewer's own reason:
阈值 80 · 通过 · 第 4 轮
The reply answered the question but never named the failing file.
Design notes:
- It binds to the answer, not to the turn. The score is recorded under
the same durable message id the shell hands to the action row
(
assistant/message→data.message.id), so the chip and the gate can never disagree about which answer a score belongs to. - It renders nothing when there is no score. Disabled plugin, no review route, an unparseable reply, or an exhausted retry cap all leave the row byte-identical to stock dsh — no placeholder, no spinner, no reserved space.
- Retries are visible. A turn that was steered shows
第2次on the answer that finally passed, so a passing score on a third attempt does not read like a first-try pass. A failed intermediate answer keeps its own chip if it has one. - The last answer always carries a chip. The retry cap silences steering,
never the review — a turn whose
maxChallengesran out still publishes its final score, markedcapped. Anything else would make the answer the user actually reads the one answer with no score (fixed in 0.7.1). - Colours are the host's. The chip uses dsh's own
--dsw-alias-state-success-*/--dsw-alias-state-warn-*tokens, so it follows the active theme with no local palette.
The score feed
The chip reads a same-origin JSON feed the host mounts on its own webserver:
GET /api/dsh-answer-reviewer/reviews → { entries: [...], at }
Each entry is { messageId, score, threshold, decision, attempt, turn, reason, at },
newest first, capped at the 200 most recent scores. It is registered through
the host's ctx.webServer.register(...) and gated by the composition's
connection fence (Host/Origin check plus the login-token cookie), which is
why the page can read it without any CORS relaxation. The 127.0.0.1:3987
sidecar deliberately does not grant cross-origin reads — doing so would
let any site the user visits read and rewrite their reviewer config — but it
does mirror the same payload at /api/reviews for curl.
Why a separate model?
The review model uses a strict JSON-only grader prompt ({score, reason})
that the agent never sees directly. The findings are injected back as a
user/message with source: { kind: 'plugin', plugin: 'dsh-answer-reviewer' }, so the agent treats them as user input — but
the agent is explicitly told not to mention the review to the user. A
different provider/model reduces the chance that a self-graded check
rubber-stamps its own work, and the explicit numerical gate means a
well-tuned model that returns 80+ will not trigger any re-attempt.
Live config (since 0.4.0)
The plugin starts a tiny node:http server bound to 127.0.0.1:3987
(no external access). Open it in a browser to see the form, the on-disk
config path, and a rolling list of the last 15 review outcomes. The same
JSON API is reachable from the shell:
# Read current effective config + overrides + file path
curl -s http://127.0.0.1:3987/api/config
# Tweak the threshold (only the fields you POST are written)
curl -s -X POST http://127.0.0.1:3987/api/config \
-H 'content-type: application/json' \
-d '{"threshold": 90}'
# Disable the reviewer
curl -s -X POST http://127.0.0.1:3987/api/config \
-H 'content-type: application/json' \
-d '{"enabled": false}'
# Wipe overrides back to defaults (also removes the on-disk file)
curl -s -X DELETE http://127.0.0.1:3987/api/config
# See the last 20 review outcomes
curl -s http://127.0.0.1:3987/api/recent
# See the score of every recent assistant answer, by message id
curl -s http://127.0.0.1:3987/api/reviews
Disable the server with REVIEWER_HTTP=0. Change the port with
REVIEWER_HTTP_PORT=<n>. Move the on-disk file with
REVIEWER_CONFIG_PATH=<abs path>. The next turn picks up the new value
with no host restart.
Config tab (since 0.7.6)
The plugin registers a first-class tab into the shell's conversation.view
slot — the same ring that carries dsh's own 对话 / 轨迹 / 记忆系统 tabs above
the composer:
对话 │ 轨迹 │ 记忆系统 │ Reviewer 配置
╰──────────── active → the main area is the page
Click the tab and the config page takes over the conversation area, exactly like any other conversation view:
┌────────────────────────────────────────────────────────┐
│ dsh-answer-reviewer · 实时配置 新标签打开 127.0.0.1:3987 │
├────────────────────────────────────────────────────────┤
│ │
│ (the config form, iframed) │
│ │
└────────────────────────────────────────────────────────┘
Three deliberate design points:
- It is a real view, not a strip above the composer. The tab is registered
with
order: 40, so it lands after every shipped view (chat0, trajectory10, mnemon30). dsh renders a registered view only while it is the active tab, so switching away unmounts the iframe and its poll timers stop on their own — novisibleprop, no persisted open/closed state. - No collapse control and no overlay. A tab has no "open" and "closed"; the view area is the page. That removes the old dock's whole geometry problem (overlaying rather than pushing the fixed-height composer column) at the source.
- A dead config server gets a hint, not a blank frame. The view probes
/api/healthonce on mount; if the sidecar is down (e.g.REVIEWER_HTTP=0) it replaces the frame with an actionable message pointing at the port and the host log.
This is the primary surface: it needs nothing beyond the core slots
client service, so it is available on every install — no dsh-better-sidebar
required.
The tab, the sidebar tab, and the standalone page all embed the same
URL. Saving through any of them is observable to the others on the very
next GET, because all three read the one ConfigStore living in the host
process.
Side card (since 0.5.0)
If the host profile also installs
dsh-better-sidebar,
the plugin also registers a right-sidebar tab titled "Reviewer 配置".
Open it from the sidebar's + menu (next to Files / Terminal / Browser).
The tab content is the same form as the standalone server, loaded
inside an iframe — saving a value through the tab is observable to the
standalone server (and vice versa) on the very next GET. When the tab
is not active the iframe unmounts so background tabs do not keep
polling.
If dsh-better-sidebar is not installed, the side card is simply hidden —
the config tab and the standalone 127.0.0.1:3987 page still cover
you. Every surface is optional and the plugin works with none of them.
The better-sidebar dependency is soft, by design: the tab is waited for
lazily (ctx.inject(["betterSidebar"], …)) rather than declared in
exports.inject. A hard inject would leave the client entry
pending (waiting for service: betterSidebar) on any host without
dsh-better-sidebar, which the web boot audit reports as
web boot: 1 entry did not activate and the shell renders as a
"Failed to load plugins" banner over the whole main page. See the 0.5.3
changelog entry.
Install
Add the package to your profile's dependencies and to
dsh.profile.bundles, then refresh the host.
// ~/.dsh/profiles/web/package.json
{
"dependencies": {
"dsh-answer-reviewer": "file:/Users/bycall/Downloads/workbuddy/Claw/dsh-answer-reviewer"
},
"dsh": {
"profile": {
"bundles": [
"dsh-answer-reviewer"
]
}
}
}
Then dsh plugin --profile web install and restart the host.
Configuration
| key | type | default | meaning |
|---|---|---|---|
enabled |
boolean |
true |
Kill switch. Set to false to disable without uninstalling. |
threshold |
number |
80 |
Inclusive 1-100 score gate. Below this, the plugin steers. |
maxChallenges |
number |
5 |
Hard cap on steers per (session, turn). Hard-capped at 8. |
maxReviewTokens |
number |
512 |
Output cap for the review model. |
timeoutMs |
number |
60000 |
Wall-clock cap on the review call. |
reviewProvider |
string |
agent's | Provider for the dedicated review model. Must pair with model. |
reviewModel |
string |
agent's | Model id for the dedicated review model. Must pair with provider. |
threshold must be an integer in [1, 100]; out-of-range values fall
back to the default rather than being silently clamped (so a mis-set
config does not quietly change the gate from "80" to "100").
reviewProvider and reviewModel must be supplied together; supplying
one without the other is a configuration error and the plugin will
refuse to mount.
Behaviour
- Subagent turns are skipped. The plugin only reviews the user-facing
agent.
agent.session.header.origin === 'subagent'is filtered out. - Empty / interrupted assistant output is skipped. No review means no steer.
- Review call failures are fail-open. Network errors, timeouts, parse failures, or model-side errors all log a warning and let the turn close normally. A flaky reviewer must never wedge the host.
- Aborted signals short-circuit. If the user paused or cancelled the turn, the review is cancelled mid-flight and the turn is left alone.
- Out-of-band scores fail closed. The review model must return an
integer in
[1, 100]with a non-emptyreason. Anything else (model returns "1.5", "99.9%", "I score this ...", etc.) is treated as a parse failure — failing open would let a misbehaving reviewer buy a pass by returning garbage. - Per-turn counter is capped. Once a turn has been steered
maxChallengestimes, subsequent turn boundaries inside the same turn are allowed to close without review. With the default5you will see at most five steered re-attempts before the agent's reply goes to the user as-is, even if the review model still scores it below the threshold — the cap is a hard budget, not a soft hint. - The agent is told not to mention the review. The steer message explicitly instructs the agent to address the finding silently.
Test
node test/smoke.mjs
The smoke test exercises every pure helper plus the fail-open,
skip-subagent, below-threshold, and cap-exhausted branches of
onTurnStopping with mock objects (mocked BlockAssembler-shaped stream
chunks). It does not boot a dsh host.
Files
lib/index.js— cordisapply, wires the listener and orchestrates the review call. Imports from./internal.js.lib/internal.js— barrel re-export solib/index.jsand tests import from one place.lib/review.js— pure helpers:resolveConfig,extractAssistantText,extractUserPrompts,buildReviewPrompt,parseScore,isScoreAcceptable,buildSteerMessage,createChallengeCounter, and the publicConfigzod schema.lib/config-store.js—createConfigStore(persistent, hot-reloadable),defaultConfigPath, env-var constants for the HTTP server.lib/server.js—startServer(store, opts)— the 127.0.0.1-onlynode:httpinstance (HTML form + JSON API).lib/client.js—window.__ModuleLoader__.loadclient bundle. Registers the score chip and theReviewer 配置tab into the shell'sconversation.chat.assistant-actions/conversation.viewslots; when the host profile also includesdsh-better-sidebar, registers a "Reviewer 配置" side card too. Both mounts are optional and the entry activates with neither.cordis.patch.yml— cordis bundle entry that mounts the plugin.test/smoke.mjs— node ESM smoke test (58 cases).CONFIGURE.md— detailed configuration guide (default vs independent review model, threshold tuning, fail-closed/fail-open matrix).CHANGELOG.md— versioned release history.
Further reading
CONFIGURE.md— answers "how do I configure an independent reviewer model", "how should I pick a threshold", and "what does this plugin do on bad config". Read it before tuningthresholdor wiring up a second provider.CHANGELOG.md— the breaking change in0.3.0is the threshold/maxChallenges split; if you are upgrading from0.2.x, read the migration note there.
评论
评论存放在 GitHub Discussions。用 GitHub 账号登录后可发表评论或点表情。