Skip to content
dsh-market Browse plugins GitHub 中文

yunxiyang/dsh-loop-continue

Resumes an agent turn that ended after narrating its next action without calling a tool: the guard replays the turn log on agent/turn-stopping, asks a judge model one true/false question, and steers the same turn to run one more step.

Stars ★ 2 Category AGI Architecture Exploration Listed 2026-09-11 npm dsh-loop-continue

Install

Inside DeepSeek Harness, with dsh-market

dsh plugin --profile web add dshmarket

Or from the command line

dsh plugin --profile web add dsh-loop-continue

Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.

README

Continue a DeepSeek Harness agent turn when the model only narrated its next action and forgot the tool call.

The agent loop closes a turn as completed the moment a step emits text and no tool call. A model that writes "now I will re-apply the change" and stops looks finished to the loop, even though the work is half done.

This plugin listens on agent/turn-stopping, replays the current turn's own session log, and only asks a judge model one strict true/false question when the shape matches that failure mode. true steers the same turn to run one more step; false (or an unparseable answer) lets the turn close.

Install

Install into a profile with the dsh CLI. Its plugin subcommand forwards to pnpm inside the profile directory, so add/remove behave as usual:

dsh plugin --profile <profile> add dsh-loop-continue

Then restart the profile. The profile's package.json gains the dependency and dsh.profile.bundles entry, and the bundled cordis.patch.yml inserts the guard into the layer stack — no manual patch editing is required.

To remove it:

dsh plugin --profile <profile> remove dsh-loop-continue

Working from a local checkout instead? Point the profile at the directory:

dsh plugin --profile <profile> add link:/path/to/dsh-loop-continue

Source edits under lib/ are picked up only on restart.

Config

field default in card meaning
maxContinuations 10 yes hard cap on steering per turn (no infinite loop)
maxSteps 6 yes newest steps shown to the judge
maxTailChars 1500 yes trailing-text budget split across the first and last halves
judgePrompt built-in yes instruction telling the judge what counts as unfinished
steerText built-in yes message injected when the guard steers
debug false yes log every evaluation (needs a restart)
judgeProvider null no override provider; null/unset = derive from the session route
judgeModel null no override model; null/unset = derive from the session route
judgeMaxTokens 64 no judge output cap
judgeTemperature 0 no judge sampling temperature

Deterministic gates

No model call runs unless the turn both:

  1. ended on a text-only step (no tool call), and
  2. called at least one tool earlier (so a plain one-shot answer is untouched).

This keeps the extra judge call off ordinary finished turns and only spends it where the model plausibly dropped a pending action.

The judge call is self-contained

What goes out for that one call is deliberately small and built from scratch: system is judgePrompt, and messages is a single user message holding renderSummary(...) — the trailing text plus the last maxSteps steps. No tools are sent, and no session history is replayed. The test sends the judge a self-contained request pins that shape.

Reusing the session would be cheaper per token, and that is not the question. The host derives messages incrementally and reuses the frozen message objects (Session.deriveMessages), so appending one user turn leaves the whole prefix byte-identical and the provider's prefix cache does hit. That part works. But the saving is the miss on roughly 1–2k tokens, and the verdict is this plugin's entire output. Three properties are worth more than that:

  • Determinism. temperature: 0 over a fixed summary answers the same way for the same turn shape. A judge reading the whole conversation drifts with it.
  • Focus. The question is narrow — does the trailing text promise an action that no tool call performed. Trailing text plus the last few steps is the signal; dozens of earlier turns dilute it.
  • Cost that tracks the turn, not the session. The call is O(1–2k tokens) whether the conversation is 3 steps or 300.

If the goal is fewer tokens, shrink the input instead: lower maxTailChars or maxSteps. That cuts the call without giving up any of the three.

Stopping early

Steering is not free: each continuation costs a judge round-trip plus one more agent step. So the guard also watches whether a steer worked.

When it steers, it records how far the session log had grown. The next time the same turn stops, it checks whether any assistant step in between actually called a tool. If the model only narrated again, the steer was ignored and the turn is closed instead of spending the remaining budget on the same refusal. A model that answers the steer with a real tool call keeps its normal budget.

Configuring the guard

The guard resolves its config on every evaluation rather than freezing it at mount, so a change reaches the very next turn-stopping check. Two ways to set values:

  • the user settings file, ~/.dsh/settings.yaml, under a loop-continue: key, or
  • the profile's own cordis.patch.yml.

Prefer leaving judgeProvider/judgeModel unset. The guard then judges with the model the current conversation is already running, which is by definition a working route; set them only to judge with a different model on purpose.

Editing the prompts

Both prompt texts ship as defaults and are plain config fields, so you can retune the guard without touching code. The package ships a browser half, so the editor is a card under Settings > Plugins > plugin config. It starts collapsed like every other plugin card; click the header to open it:

Loop Continue - resume an unfinished turn
  judgePrompt - decision policy   [textarea]
  steerText - resume message      [textarea]
  max steering per turn           [3]     0..100
  steps in the summary            [10]    1..100
  trailing-text budget (chars)    [2000]  100..20000
  debug - log every verdict       [ ]

The prompts save on blur, not per keystroke: the Host validates and persists the whole document on every write, and one write per keystroke of a 740-character policy paragraph is wasteful. The numeric fields and the toggle save immediately. Every change except debug applies to the next turn-stopping check with no restart; debug is read when the plugin is loaded, so it needs one.

The same fields are editable in ~/.dsh/settings.yaml, which is also the surface to use when the plugin runs headless with no browser half loaded:

loop-continue:
  judgePrompt: >-
    You inspect one coding-agent turn that just ended. Reply with exactly one
    word: true if the trailing text promises work that no tool call performed,
    false otherwise.
  steerText: >-
    You described an action but did not call any tool. Emit the tool call now.

Both surfaces write the same namespace, so a value set in the card shows up in the file and vice versa.

What each field controls:

  • judgePrompt — the judge's whole decision policy. The judge sees a deterministic summary (tool calls per step, trailing text, the human request) and must answer one word. Tighten it if the guard steers turns that were actually finished; loosen it if it misses dropped actions.
  • steerText — what the model is told when a turn is steered. This is the message the model acts on, so phrase it as an instruction to emit the call.

To revert, delete the fields; the built-ins come back.

Hot mount

cordis.patch.yml ships a plain insert — only id + name, no config and no !!js expressions — so a market hot-mount can add or remove this plugin as a minimal row. No user-specific endpoint, provider, or key is baked into the patch: policy is resolved at runtime as described above.

A patch-layer change is replayed in full on every reload (applyEntryPatches clones the entry list before applying), so a row adds and removes cleanly and never accumulates.

Whether a patch edit takes effect without a restart depends on the host:

  • Under the CLI (dsh profile), runProfile installs an HMR service and registers the profile and user patch files with it, so patch edits are applied live.
  • Under DSH Desktop that path is not taken — the desktop shell composes the profile itself (dsh-app-boot helpers) and never loads cordis-plugin-hmr, so a cordis.patch.yml edit needs a profile restart.

Edits to lib/*.js always need a restart: the dsh HMR service is created with root: [], so no source directory is watched for module replacement.

Development

src/ holds the sources; lib/ holds what npm publishes and what a profile loads. Nothing derives one from the other at install time, so the copy is explicit and checked:

npm install
npm run build     # src/ -> lib/
npm run check     # fails when the two differ
npm test          # vitest, against src/

prepublishOnly runs check then test, so a stale lib/ cannot be published. CI runs the same steps plus a pack assertion.

Content from the project README on GitHub ↗

Comments

Comments live in GitHub Discussions. Sign in with GitHub to post or react.