Install
Inside DeepSeek Harness, with dsh-market
dsh plugin --profile web add dshmarket
Or from the command line
dsh plugin --profile web add dsh-vibe-math
Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.
Screenshots
README
English | 中文
A set of agent presets running inside DeepSeek Harness (
vibe-math-v2/vibe-math-v3/vibe-math-v4/vibe-math-v5), which use multi-agent collaboration to automatically solve mathematical problems and perform multi-agent cross-verification of the conclusions. All four presets share the foundational capabilities of "checkpoint resume, mid-run manual intervention, progress reporting, and natural-language driving", but adopt four generations of different solving architectures: 💡 All four architectures are peers —vibe-math-v2andvibe-math-v3are the classic architectures (mature, usable, actively maintained);vibe-math-v4is the "resident self-organizing collaborative research" architecture andvibe-math-v5is the latest "institute system" (both labelled experimental); choose according to your actual needs (see "How to choose" below).
vibe-math-v2(probability-driven · JSON data layer) · classic:qs.jsonproblem list +Propos/proposition library + probability-driven scheduling + code heuristic scheduling;vibe-math-v3(third generation · paper-style md + planner agent + method library) · classic: all knowledge is stored and extended in Markdown paper/research-report form (Problems/problem list + dependencies + source motivation,Progress/research log,Propos/proposition library,Methods/general theory invention library,Verified/absolutely trustworthy); before scheduling, the planner agent autonomously draws up a plan for the next N steps; theories/frameworks/tools/methods/ideas invented during solving are distilled by the Method Keeper into a reusable method system (as in inventing group theory or functional analysis).vibe-math-v4(fourth generation · resident self-organizing collaborative research) 🧪 Experimental: a group of persistent resident subagents leave messages for one another + hold meetings, and autonomously decide all task arrangements (no central scheduler); each accumulates its own progress/proposition/method/subproblem libraries and consults the others; verification is written toVerified/only when all residents agree (true or false), otherwise it remains in the library with a probability attached; when the context reaches a threshold it automatically/compacts; it stops only when all agree that the original problem has been solved.vibe-math-v5(fifth generation · institute system) 🧪 Experimental · Latest: upgrades the residents into an institute — academicians (leaders / the organizing and coordinating center, responsible for decomposition and assignment, setting priorities, chairing meetings, and supervising progress) + resident researchers (with voting rights, able to autonomously hire/fire their own temp workers) + temp workers (no voting rights); it has a public charter, group chat and meetings, a compare-and-set task board, and real firing; a boolean agreement of ≥ m votes is required to write toVerified/(opposing votes block, abstentions are not counted, and if the threshold is not met it remains in the library with an average probability attached); state is written to the hardened JSON fileState/<institute>.v5state.jsonunder the institute directory (serial writes, a mandatory load before read) — never into the host session log — at zero token cost.
After installing this plugin package (or manually copying the presets), four agent presets appear in DSH's preset selector.
🧭 Navigation: I want to … → start here
- Get started (first run) → Start in five minutes
- Which of the four presets to pick → How to choose among the four presets
- Positioning, flow and division of labour of the four → Architecture Diagrams · Architecture and division of labour
- New: the final paper (produced at closure) → Features · full contract
- Lean formal verification → Lean formal verification · full contract
- What is in the directories → Directory structure
- Tuning parameters → Parameter quick reference
- Checkpoint resume / mid-run intervention → Checkpoint resume & manual intervention
- Known limitations → Known limitations
🧩 Architecture Diagrams (v2 + v3 + v4 + v5)
Static architecture diagrams; for the complete process description see the v1-era architecture notes (historical: the layout changed from v2 on) and the v5 detail diagrams (the full set of v5 detail diagrams); editable generation scripts: the Chinese v2/v3 posters come from the matplotlib scripts v2 / v3 (matplotlib → PNG); the English v2/v3 diagrams come from the zero-dependency Node scripts v2-en / v3-en (→ SVG); v4 / v5 (zero-dependency Node → SVG,
node docs/generate_framework_diagram_v4.mjs; add--lang=enfor示例图/框架图-v4-en.svg). SVG is used from v4 onward: plain text, diff-friendly, crisp at any zoom; when PNG is needed, screenshot with a headless browser (the command is at the top of the generation script).
Vibe Math V2 (probability-driven · JSON data layer) · classic
One-sentence pipeline: qs.json takes problems by priority → Explorer splits out directions (if all are dead ends, re-derive) → one Solver per direction iterates over multiple rounds (lemmas go into Propos/, solutions go back to qs.json, all probabilities <1) → the scheduler picks r (proposition / proposition+proof·disproof / problem+solution) and dispatches ≥3 verifiers for independent review → debate → ruling → at probability=1 it automatically closes out (problem solved, proposition 1/0, priority set to never); state is written to disk throughout, resume continues from the checkpoint, and reportMode can report by file/push/both.
Vibe Math V3 (paper-style md + planner agent + methods library) · classic
One-sentence pipeline: all knowledge is stored and continued as Markdown papers/research reports (Problems/ problem list including dependencies and the source motivation of follow-up problems, Progress/ research log continued by direction and by round, Propos/ proposition library, Methods/ general theory invention library, Verified/ absolutely trustworthy) → before scheduling, the scheduler builds a state brief and calls the planner agent; the planner agent lays out the next N steps in one go (spawn solver/verifier/explorer/method-keeper, interrupt, promote, wait), which are executed after code validation (actions exceeding the concurrency limit are queued and consumed across ticks; a planning failure automatically falls back to the v2-style heuristic) → verifiers review independently → debate → near-consensus ruling (if on the same side and the mean is ≥0.85/≤0.15, take the mean, fixing v2's flat misjudgment) → at probability=1 it closes out and generates a Verified/ card → the solver's methods_used/new_inventions reports are distilled/refined into the methods library by the Method Keeper (which can form system hierarchies and be reused across projects).
Vibe Math V4 (resident self-organizing collaborative research) 🧪 Experimental
The SVG above is generated by a zero-dependency script:
node docs/generate_framework_diagram_v4.mjs --lang=en(pure Node, no Python/matplotlib dependency; generation estimates text width, and any line overflowing its container raises a warning and exits with code 1).
One-sentence pipeline: initially N resident subagents are created (continuable, persistent context) which first brainstorm on their own and produce initial insights/directions → after that all task arrangements are decided autonomously by them leaving messages for each other + holding collective meetings (the framework only provides the message bus/meetings/task board/artifact persistence, and never assigns tasks); each resident persists valuable artifacts into its own Progress/<id>/, Propos/<id>/, Methods/<id>/, Subproblems/<id>/ libraries according to degree of value / planned motivation and use / its own probability estimate, and they can read each other's; verification is initiated by their own deliberation, and only when all residents agree (true or false) is it written to Verified/, otherwise it stays in the library with a probability attached; when a resident's context reaches a threshold (66% by default) it automatically /compacts; they stop only when all of them agree that the original problem is solved; residents can be manually intervened with/added/shut down at any time, and checkpoint resume is supported.
Note: V4 removes v3's central planner and deterministic roles (explorer/solver/verifier/planner/method-keeper) and makes the "researcher" itself the subject. See
vibe-math-v4/实现方案.mdfor details. Keep-alive mechanism (tiered keep-alive A+B + deadlock watchdog): a gang idle for longer thanactivityTimeoutMsreceives a self-driven CHECKPOINT (suggesting it continue solving/send a message/propose a task, rather than "do you want to stop"), and it fills in parallel — branch A fills as much of themaxParallelconcurrency budget as possible in one go (waking several idle residents at the same moment, rather than the serial "wake only r1, then r2 after it finishes"), and mailbox delivery also reaches several idle recipients in parallel; a failed wake automatically re-arms the heartbeat; if the team is idle and has no new artifacts for longer thanstallAutoMeetingMs(6 minutes by default), the framework automatically convenes a synchronous meeting so the residents can decide the next step themselves; if a meeting/verification hangs (still no new speech/votes after more than 2×activityTimeoutMs), the framework automatically abandons that meeting/verification and returns to normal self-organization, so that one broken meeting does not permanently block the whole team; meetings do not preempt verification — meeting requests while verification is under way are held and convened afterwards (keeping the consensus-consistent "truth-seeking" step from being interrupted by coordination discussion) — the framework always only facilitates and never assigns tasks.
Vibe Math V5 (institute system) 🧪 Experimental · Latest
One-sentence positioning: upgrade v4's "a group of residents messaging each other" into an institute — with three classes of staff: academician (leader), resident researcher, and temp worker; with the institute's public charter; with group chat and meetings; with autonomous hiring/firing; and where
any conclusion must be given a Boolean probability of 1 or 0 unanimously by at least m voting members before it can be written to Verified/.
Image sources and all detail diagrams (member lifecycle, one-round sequence, consensus state machine, meeting flow, scheduling priority, state folding, prompt composition, task board, authority matrix): the v5 detail diagrams. The SVG above is generated by a zero-dependency script:
node docs/generate_framework_diagram_v5.mjs --lang=en.
flowchart TB
OFF["👤 Institute office (session root agent / human)<br/>does not research · does not vote · only reports and relays instructions"]
subgraph INST["🏛️ Institute (internal autonomy: roster, organization and assignment all happen among members)"]
ACAD["Academician acad —— leader / center of organization and coordination<br/>L1 institute-wide overview · L2 assignment · L3 priority<br/>L4 chairs meetings · L5 supervision · L6 moving people · L7 external"]
RES["Resident researcher r-n<br/>has voting rights · may autonomously hire/fire its own temp workers"]
TMP["Temp worker t-n<br/>no voting rights · hired temporarily for a specific task"]
end
subgraph FW["⚙️ Framework vibe-v5 —— only a medium (middleware), never assigns tasks"]
M["Message relay · meetings/debates · task board CAS+DAG<br/>m-vote consensus verification · context and liveness · roster and hiring · scheduler"]
end
PROJ["💾 State/<institute>.v5state.json (hardened JSON, authoritative)<br/>12 kinds of events · pure fold applyV5Event · serial writes · recovery = load before read"]
FS["📁 Members/<id>/* · Shared/* · Verified/ · Problems/"]
RULE{{"Truth gate: Boolean unanimity and Boolean votes ≥ m = min(quorumCap, number of registered voting members)"}}
OFF <-->|"vibe_v5_* / /v5 commands ↔ status / report"| M
M <-->|"per-round prompt ↔ single JSON receipt"| ACAD
M <-->|"per-round prompt ↔ single JSON receipt"| RES
M <-->|"per-round prompt ↔ single JSON receipt"| TMP
ACAD -.->|"assign / supervise / chair meetings (in-institute organization, not framework behavior)"| RES
ACAD -.-> TMP
M <--> PROJ
M <--> FS
M --> RULE
One-sentence pipeline: the institute office (session root / human) configure → starts the institute → the academician (center of organization and coordination) decomposes the original problem, assigns tasks, sets priorities, chairs meetings and supervises progress; resident researchers research on their own and hold the vote, and temp workers can be hired as needed (no vote, genuinely dismissible). The framework is only a medium and never assigns tasks; every conclusion needs ≥ m = min(quorumCap, number of registered voting members) Boolean votes, all on the same side, before it is written to Verified/ (an opposing vote blocks; abstentions do not count as votes but count toward the average; below the threshold the object stays in its library with its average probability); meetings and verification are mutually exclusive in both directions; state lives in the hardened JSON State/<institute>.v5state.json under the institute directory (zero token cost, serial writes, a mandatory load before read); the problem is concluded only when all voting members consider it solved.
Positions and authority, the detailed truth rules, operating mechanisms, state and persistence, prompt composition, the institute directory, the tool surface, and the differences from v4: see the v5 detail diagrams §13 "v5 notes moved from the README" (text preserved); parameters are in the parameter quick reference, and the final paper in
docs/final-paper.md.
✨ Features
- Final paper (on by default in all four presets): at closure the run writes a complete paper that only organises evidence it already has, and the paper phase runs before the run is marked complete; the artifacts are
Paper/<id>/{paper.md,paper.tex,paper.pdf,paper.meta.json,paper.log.md}. Parameters:finalPaper(defaulttrue) /paperFormat(defaultboth) /paperLanguage(defaultzh) /paperCompilePdf(defaulttrue) (pluspaperEditoron v4/v5); the manual trigger is/vN paper [lang=] [format=] [editor=] [force]. A PDF needs a LaTeX engine on the host (xelatexpreferred for Chinese); without one the tex+md are still delivered. Full contract and usage:docs/final-paper.md. - Multi-agent automatic solving: the main agent hands the problem to the scheduler, which dispatches explorer / solver / verifier (v2/v3) plus subagents such as planner (planning agent, v3) and method-keeper (method organizing agent, v3) to solve collaboratively; you do not need to operate node by node by hand.
- Multi-agent cross-validation: every conclusion goes to ≥3 "strict reviewers" for independent review → debate (exchange group) → adjudication (v3 defaults to near-consensus adjudication: if on the same side and the mean is ≥0.85/≤0.15, take the mean, so that "0.9 vs 1" is not misjudged as 0.5).
- Paper-style Markdown knowledge base (v3): the problem list (including dependencies between problems, and the causes and plans of descendant problems), research log, propositions, and method library are all written and continued in md paper/research-report style (when a direction is re-derived, the logs of the old direction are automatically archived and kept); only
Verified/and the objects that a verifier judged true/false are absolutely trustworthy, and all other md (including unverified assertions in the method library) serve only as empirical reference. - General theory invention library (v3): the theory systems/frameworks/tools/methods/ideas invented during solving (including empirical summaries) are reported via
methods_used/new_inventions, and the Method Keeper consolidates them intoMethods/method cards (which can form a hierarchy of systems and be reused across projects), forming a systematic method–theory system just like "inventing group theory while solving equations". - Planner-agent scheduling (v3): before scheduling, the planner agent is invoked to autonomously choose the optimal scheduling scheme according to the actual situation (problem dependencies / survival rate / verifiable objects / concurrency budget / results of the last plan), arranging the tasks of each agent for the next N steps in one go; if planning fails, it automatically falls back to heuristics.
- Knowledge accumulation: conclusions that pass verification are promoted into the
Verified/trustworthy knowledge base (v2/v3 additionally have thePropos/proposition library) for reuse by later directions. - Checkpoint resume: the scheduling state, task stack, agent registry, decision queue, verifier historical accuracy, etc. are all written to disk; after a restart,
resumerestores them (v2/v3 use a process epoch to distinguish "pause → resume within the same process" from "restart across processes"; v3's md is itself the narrative breakpoint). - Mid-run manual intervention (and continue): the
auto / manualmodes can be switched at any time; in manual mode, a decision is suspended at key nodes and waits for your approve/reject/override (v3 adds a plan approval gate and a method promotion gate); you can send a message to / interrupt any subagent. - Per-project isolation: each mathematical problem is an independent project folder, with no interference between them, and you can switch at any time.
- Multi-session parallel isolation: a DSH agent preset is a standing mount (all sessions of the same preset share one plugin instance), and inside the plugin all running state is isolated by root session id — two sessions can each run a project at the same time, their respective subagents are correctly attached under their own session, and the scheduler / parameters / decision queue / current project do not interfere with each other (v3 additionally has a project lock, so the same project is scheduled by only one session at a time). The current project is persisted per session (
VibeMath/current.<session id>.json). - Adjustable subagent permissions: you can restrict the tools a subagent is allowed/forbidden to use and the per-round cap on external tool calls, and explicitly tell it that it may read
Verified/,Propos/,Methods/,Reliable/and the progress log. - Configurable:
vibe_math_setting.json(with comments) for customizing default parameters;/vibe setupfor interactive question-and-answer configuration. - Natural-language control: the main agent acts as "assistant + reporter" — you state your needs in plain words, and it calls the tools, reports progress, and configures parameters on its own.
Specific to v4 / v5:
- Persistent self-organization (v4): at the start, N persistent resident subagents are created, after which all task arrangements are decided by those subagents themselves through leaving messages for each other + holding meetings (the framework only provides the message bus/meetings/task board, and never assigns tasks).
- Institute system (v5): on top of v4's self-organization, it introduces the organizational form of a real research institute — the academician (leader) is responsible for decomposition, assignment, prioritization, chairing meetings, and supervising progress; resident researchers have voting rights and can autonomously hire/fire their own temp workers; temp workers have no voting rights; all organizational actions are performed by members of the institute, and the framework still only acts as the medium. See the Vibe Math V5 section above for details.
- Adjustable quorum (v5): for an object to enter
Verified/, ≥ m = min(quorumCap, number of enrolled voters) voters must cast a consistent boolean vote (all1or all0); an opposing vote blocks, and abstentions do not count as votes but do count toward the average probability; if the threshold is not reached, the object is kept in the library with the average probability and the complete debate record, without forcing an adjudication. - Zero-token-cost state persistence (v5): the institute state is written to the hardened JSON file under the institute directory (
State/<institute>.v5state.json) and does not enter the model context; cross-process and same-process recovery go through the same code path (a mandatory load before read). - Genuinely reversible roster (v5): hiring creates a resident child session; firing cancels in-flight turns, releases the child session, reclaims its tasks, and discards undelivered mail; a codename is never reused.
- Human-readable mirror (v4/v5): the roster table, task board, meeting minutes, debate records, and closing records are all written to disk as Markdown and are readable by humans at any time; but the authoritative state is not in these files (in v5 it is
State/<institute>.v5state.json), so manually corrupting them will not break the institute.
🧮 Lean formal verification (shared by the four architectures, adjustable switch)
What it changes is not "being a bit stricter" but the object of review itself. m agents agreeing that "this is right" is still consensus — it cannot rule out shared misunderstanding; passing Lean is machine checking. The only remaining uncertainty is thus reduced to a question a single person can review effectively: are the definitions / objects / conditions / assumptions / conclusions in the Lean code fully consistent with the original text of the proposition (fidelity)?
Switch
formalVerify(same name in all four architectures, default'off'):'off'is a true no-op (no Lean content appears in the prompts, no formalization state is written, and the verification flow and gates are completely unchanged; the five tools stay registered and usable and the persona always lists them, otherwise the switch would be undiscoverable and could not be turned on);'encourage'is encouraged but not mandatory (decide by implementation difficulty whether to formalize; once Lean passes, the focus of review shifts to fidelity; no gate);'require'is mandatory — a true/false conclusion must satisfy "Lean has passed" or "the blocking reason was explicitly recorded", otherwise the adjudication does not take effect (recorded as undecided, reasonformal-required, written into the "formalization TODO"). Related parameters:leanCommand(defaultlean),leanArgs(used withlake env lean),leanTimeoutMs(default 120s).⚠️ A fidelity defect ≠ the proposition is false (important): passing Lean only guarantees that "this piece of code passed the kernel"; it does not guarantee that it says what the proposition means to say. When a voter finds the Lean code and the original proposition inconsistent (too narrow / too broad / a different object / a missing condition): do not vote 0 (0 means "the proposition is false" — that would record "the formalization does not qualify" as "the proposition was disproved", and under v5's all-0 consistency rule it would even write the proposition into
Verified/marked false, so a mechanism meant for truth-seeking would fabricate a wrong negative conclusion); instead vote a value strictly between 0 and 1 (an abstention) and record the deviation with the receiptformal:{decision:'defect', note:'<specific deviation>'}. The framework then revokes the "passed" state of this proof (downgraded toattempted,Verified/Lean/<id>.leandeleted or rewritten as a "withdrawn" note if the host cannot delete it, and written into the "formalization TODO"), and underrequirethis adjudication is not concluded (theencouragesetting has no gate, so you must not claim the framework will force a shelving). Vote 0 only when the voter, independently of this Lean code, can also determine that the proposition is false (and can give independent reasons).The receipt channel and three further hard requirements (members who do not call the Lean tools can also leave a judgment; mandatory under
require): the receipt field is"formal": {"target":"<object id>", "decision":"used|blocked|defect", "file":"Formal/<object id>.lean", "note":"difficulty judgment/blocking reason/specific deviation"}; whendecision='blocked'/'defect',noteis required (if missing the whole entry is rejected),usedonly records the object asattempted, and underoffthis channel is disabled (otherwiseoffwould not be a true no-op). Three more: tool names in the injected text are always given in full (<prefix>_lean_archive, notlean_archive— an abbreviation is not a registered name); before archiving a reusable definition/lemma run it through first, and if it does not run through it must not enter the library; when the toolchain is missing (LEAN_NOT_FOUNDcannot resolve the executable / the host has nosubprocessservice), write the code down, archive it, and state "the host has no Lean toolchain" innote— this counts as an explicit blocking reason and the gate lets it through.Where the artifacts go:
<VibeMath root>/Formal/Lib|Proved/are the cross-project reusable definitions and already-proved lemmas (each with anIndex.md, to be checked before writing a new definition); inside a project,Formal/<object id>.leanis the object's working file (next to anIndex.mdand, underrequire, aTODO.md), andVerified/Lean/<object id>.leanis the archived proof (in v5 these live underProjects/<project>/Institutes/<institute>/).The five tools (one set per architecture, prefix following each one's naming):
<prefix>_lean_runexecutes Lean on the hostsubprocessservice (withleanAsync=trueit enqueues and returns{async:{jobId,state}}; otherwise it returns{ok, exitCode, ms, stdout, stderr}synchronously) and never throws (missing toolchain →LEAN_NOT_FOUND, timeout →LEAN_TIMEOUT, path escape → rejected);<prefix>_lean_archive:kind='def'/'lemma'archives into the cross-projectFormal/Lib|Proved(identical content is deduplicated),kind='proof'writesFormal/<target>.leanand only asettled(ok)job writesVerified/Lean/<target>.leanand marks the object Lean-passed,kind='blocked'records an explicit difficulty judgment/blocking reason (reason required);<prefix>_lean_librebuilds and returns the three indexes, the per-object status and the background jobs — check for duplicates and reuse before writing a new definition;<prefix>_lean_readreads one archived file back verbatim (onlyFormal/Lib|Proved, 64KB cap);<prefix>_lean_jobinspects or waits for a background compile (read-only). For example, v5 isvibe_v5_lean_*, v2/v3 arevibe_math_lean_*, and v4 isvibe_v4_lean_*(each with run / archive / lib / read / job).Boundaries (intentional): the framework does not bundle Lean (no toolchain installation, no dependency downloads; when the toolchain is missing it degrades gracefully and records this faithfully); it does not judge fidelity (that is what agents/humans review and vote on; the framework is only responsible for switching the focus of review to fidelity); Lean passing ≠ the proposition is true — it only means "this piece of formalized code passed the kernel check".
For the complete contract (parameters, paths, state transitions, prompt semantics, gate locations, index format, test requirements), see docs/formal-verification.md. The personas (the prompt the main agent receives) of all four presets fully list these five tools and eight parameters, and the prefix and text blocks are identical line by line (only line 0 may differ); this layer is guarded by audit-persona-surface.test.mjs and audit-persona-sensitivity.mjs — when this feature was added, it was precisely in the four presets that the defect "the tools were registered but the persona never listed them" was found (see the release notes shipped with the package).
🚀 Installation
Two installation methods, choose either one (they can also coexist):
Method A: one-click install as a plugin package (recommended, installs all four presets at once)
On the desktop app, install dsh-vibe-math from Settings → Plugins; on the command line, use a profile:
dsh plugin --profile <your profile> add dsh-vibe-math
# Or install directly from GitHub:
dsh plugin --profile <your profile> add github:ChongCyrus/Vibe-Mathematics
Then start a new session and pick Vibe Math V2 (v2, classic), Vibe Math V3 (v3, classic), Vibe Math V4 (v4, resident self-organization) or Vibe Math V5 (v5, institute system) in the preset picker — the four architectures are peers, choose according to your actual needs (see "How to choose"). The two DSH generations land in different places, and this package adapts to both:
- DSH ≥ 0.1.7 (current): agent presets are declared as composition rows. This package declares all four presets in
cordis.patch.yml(each row hands that preset's full plugin list to the host'sagentPresetsservice), and writes nothing into~/.dsh/.agent-presets/— that directory has not been read since 0.1.7. - DSH ≤ 0.1.6: presets are still directories, and the installer writes the four presets into
~/.dsh/.agent-presets/(vibe-math-v2/…vibe-math-v5/); the declaration rows in the samecordis.patch.ymlregister nothing on those versions, so older hosts do not error at boot.
After upgrading the package version, restart DSH: the managed content is replaced wholesale with the new version's bytes — including files you edited by hand. This is intentional: a preset that is "half old version, half new version" will fail to mount or behave strangely, and you cannot tell from the outside. The hand edits that get replaced are not lost: in the directory form the original text is first backed up to ~/.dsh/.agent-presets/.vibe-math-backup/<old version>/<preset>/, and the file names are listed in the log.
If you want to customize a preset, do not edit managed content — make a copy (the copy action in the preset picker, or declare it under a new id in your own bundle). Note: on current DSH, if you save your own declaration for the same id, this package skips registration and logs one line saying so — your declaration wins.
Method B: manual declaration (customization / secondary development)
- DSH ≥ 0.1.7: put the whole of
vibe-math-vN/agent.cordis.ymlinto a declaration row as itsplugins— either copy thedsh-vibe-math/preset-declarationrow from this package'scordis.patch.yml, or use the host's own@deepseek-ai/dsh-agent-presetrow; install this bundle into your profile. Full rules: the DSH skillediting-cordis-compositions. - DSH ≤ 0.1.6: copy
agent.cordis.yml/preset.yml/vibe-math-vN.jsfromvibe-math-vN/into~/.dsh/.agent-presets/vibe-math-vN/.
Then start a new session and select "Vibe Math V2" / "Vibe Math V3" / "Vibe Math V4" / "Vibe Math V5" in the preset picker; once the session starts it is ready to use: v2/v3 tools are vibe_math_*, v4 is vibe_v4_*, v5 is vibe_v5_*; typing /vibe, /v4, /v5 in the input box gives autocompletion.
After changing a preset definition you must restart the DSH process before starting a new session (a preset's standing mount is cached until the process exits).
DSH version adaptation and dependencies
- Supported range:
0.1.2-alpha.4…0.2.0-rc.2(11 versions declaredcompatible, tested target0.2.0-rc.2;engines.nodeis^22.19.0 || >=24.0.0) — per-version declarations, theengines.dshrange andpeerDependencieslive inpackage.json(dsh.compatibility.dshReleases). - Delivery form: DSH ≥ 0.1.7 uses composition-row declarations (the
agentPresetsservice); DSH ≤ 0.1.6 uses the~/.dsh/.agent-presets/<id>/directory + preset picker — one bundle covers both lines. - Hard constraint: v5 keeps its institute state only in
State/<institute>.v5state.jsonand never in the host session log — on an unknown event type DSH refuses to load the whole session, so it will not open on the next resume. - Everything else (host rows and required services, the startup self-check and capability gate, the live-resident cap, the three 2026 fixes, the
dsh.bundle.patchconstraint, the upgrade path andpeerDependencies): seedocs/COMPAT-AUDIT-ROUND2.md; to upgrade this package usedsh plugin --profile <your profile> add dsh-vibe-math@latest(useadd, notupdate).
🧭 How to choose among the four presets
💡 All four architectures are peers; choose according to your actual needs:
- Choose
vibe-math-v2(probability-driven · JSON data layer) if you:
- prefer structured JSON data (
qs.json/Propos/<category>_Propos.json/Verified/cards), convenient for programmatic retrieval and further processing;- want mature and stable code-heuristic scheduling (priority + probability, predictable behavior, not dependent on the planner agent's "improvisation");
- do not need method library accumulation / paper-style narration, and data being field-oriented is enough.
- Choose
vibe-math-v3(paper-style md + planner agent + method library) if you:
- prefer a paper/research-report-style natural-language knowledge base (the problem list includes dependencies and the source motivation of follow-up problems, the research log is continued by direction and by round, human-readable and freely extendable);
- want scheduling to be planned autonomously by the planner agent as an N-step plan based on the actual situation (more flexible, automatically falls back to heuristics on failure);
- want a general theory invention library — theories/frameworks/tools/methods/ideas invented during solving are accumulated by the Method Keeper into a reusable, systematizable, cross-project-extendable methodology (like "inventing group theory while solving equations");
- accept the trust layering of "only
Verified/is absolutely trustworthy, the rest of the md is empirical reference".Both are mature and usable, continuously maintained, and both support checkpoint resume, manual/automatic intervention, progress reporting, multi-session isolation, proposition promotion, near-consensus/weighted adjudication and other core capabilities; the switching cost is low (the same set of
vibe_math_*tools and/vibecommands, the same parameter system).
- Choose
vibe-math-v4(resident self-organization) if you want a group of persistent resident subagents that message each other and hold meetings, fully self-organizing (no leader, no central scheduling), and can accept a strict threshold of "a conclusion requires unanimity".- Choose
vibe-math-v5(institute system) if you want:
- organized self-organization — like a real research institute, with a leader (academician) responsible for decomposition, assignment, prioritization, chairing meetings and supervising progress, but judgment still belongs to each individual;
- a roster that can grow or shrink — resident researchers + temp workers who can be autonomously hired/dismissed (temp workers have no voting rights, suitable for chores such as checking, trial computation and material organization);
- an adjustable consistency threshold —
m = min(quorumCap, number of voting members)boolean-consistent votes settle the matter (because an opposing vote blocks, the effective threshold is still "every Boolean vote points the same way"; it differs only when fewer thanmmembers cast a Boolean vote);- zero-token-cost state persistence — the institute state is written to a hardened JSON file under the institute directory and does not consume member context budget.
⚠️
vibe-math-v2andvibe-math-v3are the classic architectures;vibe-math-v4andvibe-math-v5are experimental architectures, all four are peers and all are selectable; the oldvibe-math-v1has been removed (this package contains only v2/v3/v4/v5).
| v2 (probability-driven · classic) | v3 (paper-style md · classic) | v4 (resident self-organization · experimental) | v5 (institute system · experimental) | |
|---|---|---|---|---|
| Positioning | Classic (JSON data layer) | Classic (third generation) | Experimental (fourth generation) | Experimental (fifth generation) |
| Core idea | Probability-driven: qs.json problems + Propos/ proposition library, scheduled by "correctness probability / value" |
Paper-style md knowledge base + planner agent scheduling + general theory invention library | Persistent resident subagents self-organizing: message each other + meetings decide all tasks, no central scheduling | Institute: the academician organizes and assigns, members research on their own; a conclusion requires ≥ m boolean-consistent votes; temp workers can be hired as needed |
| Data | qs/qs.json + Propos/<category>_Propos.json + Reliable/ |
Problems/ + Progress/ + Propos/ + Methods/ (all md, soft-spec anchors + free narration) + Verified/ |
Problems/ + `Progress |
Propos |
| Roles | explorer → per-direction solver → verifier | planner (planner agent) → explorer → per-direction solver → verifier → method-keeper (method organization agent) | N resident researchers (continuable), no fixed roles | academician acad (leader) + resident researcher r-n (with voting rights) + temp worker t-n (no voting rights, hireable and dismissible) + institute office (does not research and does not vote) |
| Scheduling | Code heuristics (priority + probability) | The planner agent produces an N-step plan (executed after validation, falls back to heuristics on failure) | No central scheduling: tasks arise from residents messaging each other / holding meetings (the framework is only a medium and does not assign) | The framework still does not assign; the academician decomposes/assigns/prioritizes/supervises, and members may object with reasons; the framework only relays, keeps the task board, holds meetings and counts |
| Closing rule | A solution/proof reaching probability 1 closes it; never is never scheduled |
Same as v2 (near-consensus adjudication fixes flat misjudgment) | Writes to Verified/ only when all residents agree (true or false), otherwise leaves it in the library with a probability |
Writes to Verified/ only when boolean votes ≥ m = min(quorumCap, number of voting members) and all are 1 or all are 0; opposing votes block; abstentions do not count as votes but count toward the average; (can be switched back to the v4 criterion) |
| Stopping | All solved or stuck | All solved / no candidates | Stops only when all residents unanimously agree the original problem is solved | Same as v4: concluded only when all voting members unanimously agree the original problem is solved |
| Context | None | None | Resident context automatically runs /compact on reaching the threshold (adjustable) |
Same as v4 (threshold/round count adjustable; after compaction the charter still takes effect in the persona) |
| Ad hoc capabilities | Automatic promotion of proposition "value/criticality" to the problem list; reportMode file/push/both; priorityAdjust |
Method library accumulation loop (methods_used/new_inventions → Method Keeper); plan approval gate/method promotion gate; project lock; follow-up problem "source and motivation" as a first-class citizen |
Residents accumulate individually + read each other; unanimous verification; add/close residents at any time, intervene by message; checkpoint resume | All v4 capabilities, plus: real hiring/dismissal (releases sub-sessions, reclaims tasks); compare-and-set task board + dependency DAG; strict mutual exclusion between meetings and verification; roster mirror and conclusion records; zero-token-cost state |
All four support: checkpoint resume (vibe_math_resume / vibe_v4_resume / vibe_v5_resume), manual intervention and pause/resume,
per-project isolation, subagent permission control, and natural-language driving. All four architectures are peers — choose v2 if you prefer structured JSON data and deterministic scheduling,
choose v3 if you prefer paper-style md, the planner agent and the theory invention library; v4 is fully self-organizing resident collaborative research, and v5 is an institute system with "a leader + a roster that can grow or shrink + an adjustable threshold".
🧠 Architecture and division of labour (all four in parallel)
V2 (probability-driven · classic)
The framework = main agent + code scheduler + explorer / solver / verifier subagents. The scheduler is the sole master control and the sole file writer (subagents only return structured JSON and never write files); it decides the next step by code heuristics over "correctness probability + value/criticality + priority"; each verification object goes to ≥3 independent "harsh reviewers" for review → debate → adjudication (a solution/proof reaching probability 1 closes it; never is never scheduled). Pick it if you prefer structured JSON (qs.json / Propos/) and predictable, deterministic scheduling that does not depend on a planner agent.
Division of labour in one sentence: the main agent handles "talking to people", the scheduler handles "execution and boundary-keeping", and the subagents handle "thinking".
V3 (paper-style md + planner agent + method library) · classic
The framework = main agent + code scheduler + planner agent + explorer / solver / verifier / method-keeper subagents. Each time a dispatch is prepared, the planner agent reads the state brief (problem dependencies, survival rate, verifiable objects, concurrency budget, result of the last plan) and autonomously draws up an N-step plan, which the scheduler validates before executing (falling back to heuristics on failure); theories/frameworks/tools/methods/ideas invented during solving are reported via methods_used/new_inventions and distilled by the Method Keeper into new method cards in the Methods/ general theory invention library. Pick it if you prefer a paper-style md knowledge base, want more flexible scheduling, and want a systematizable, cross-project method library.
Division of labour in one sentence: the main agent handles "talking to people", the planner agent handles "setting the plan", the scheduler handles "execution and boundary-keeping", the subagents handle "thinking", and the Method Keeper handles "depositing inventions into theory".
V4 (resident self-organization · experimental)
The principal is N persistent resident subagents (continuable): there is no central scheduling and no leader — task arrangements emerge from residents messaging each other + holding meetings (the framework only provides the message bus/meetings/task board and never assigns); each accumulates its own Progress/Propos/Methods/Subproblems library owned by resident id and may read the others'; verification is written to Verified/ only when all residents agree (true or false), otherwise it stays in the library with a probability; when the context reaches a threshold it automatically /compacts; it stops only when all agree that the original problem is solved (vibe_v4_resume resumes from a checkpoint). Pick it if you want fully self-organizing research and can accept the strict "unanimity before a conclusion" threshold.
V5 (institute system · experimental · latest)
It upgrades v4's residents into an institute: academician (leader / organizing and coordinating centre: decomposition, assignment, prioritization, chairing meetings, supervising progress) + resident researchers (voting rights, may autonomously hire/fire their own temp workers) + temp workers (no voting rights); the framework still only relays, keeps the task board, holds meetings and counts — it never assigns tasks. The conclusion threshold is ≥ m = min(quorumCap, number of enrolled voting members) boolean-consistent votes (opposing votes block; abstentions do not count as votes but count toward the average; below the threshold the object stays in the library with its average probability); state lives in State/<institute>.v5state.json (zero token cost); meetings and verification are strictly mutually exclusive. Pick it if you want "organized self-organization", a roster that can grow or shrink, and an adjustable threshold.
For the complete v5 architecture (member lifecycle, one-round timeline, consensus state machine, meeting flow, scheduling priority, state folding, prompt composition, task board, authority matrix), see the v5 detail diagrams; for the textual specification see the v5 specification.
📁 Directory Structure
Only the top-level and high-value paths are listed; for the full meaning of every file see each preset's
实现方案.md(v2 · v3 · v4 · v5) and the v5 detail diagrams.
v2 (probability-driven · classic)
<VibeMath>/Projects/<project>/
├─ qs/qs.json # problem list (overview/solved/solutions + correctness probability/priority)
├─ Propos/<category>_Propos.json # proposition library (boolean estimate/proof·disproof/value·criticality)
├─ Verified/ # concluded-fact index
├─ Reliable/ # trusted references (read-only, placed by the user)
└─ VibeMath_State/ # scheduler-private persistent state (for checkpoint resume)
v3 (paper-style md + planner agent + method library) · classic
<VibeMath>/
├─ Methods/ # [global] cross-project general theory invention library (promoted from project level)
├─ Formal/{Lib,Proved}/ # cross-project re
…
Comments
Comments live in GitHub Discussions. Sign in with GitHub to post or react.