Skip to content
dsh-market Browse plugins GitHub 中文

ChongCyrus/Vibe-Mathematics

Installs four multi-agent math-research presets for DSH: v2 (probability-driven pipeline with multi-verifier debate), v3 (paper-style Markdown knowledge base with a planner agent and a reusable method library), v4 (persistent self-organizing residents that message and meet), and v5 (a research institute with an academician who decomposes and assigns work, voting researchers, temp workers, group chat and a compare-and-set task board); all four support checkpoint resume, human intervention, and an optional Lean formal-verification switch (off/encourage/require) whose passing proof turns the vote into a fidelity check of the Lean statements.

Stars ★ 33 Category Workflow & Automation Listed 2026-08-15 npm dsh-vibe-math

Install

Inside DeepSeek Harness, with dsh-market

dsh plugin --profile web add dshmarket

Or from the command line

dsh plugin --profile web add dsh-vibe-math

Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.

Screenshots

README

English | 中文

A set of agent presets running inside DeepSeek Harness (vibe-math-v2 / vibe-math-v3 / vibe-math-v4 / vibe-math-v5), which use multi-agent collaboration to automatically solve mathematical problems and perform multi-agent cross-verification of the conclusions. All four presets share the foundational capabilities of "checkpoint resume, mid-run manual intervention, progress reporting, and natural-language driving", but adopt four generations of different solving architectures: 💡 All four architectures are peers — vibe-math-v2 and vibe-math-v3 are the classic architectures (mature, usable, actively maintained); vibe-math-v4 is the "resident self-organizing collaborative research" architecture and vibe-math-v5 is the latest "institute system" (both labelled experimental); choose according to your actual needs (see "How to choose" below).

  • vibe-math-v2 (probability-driven · JSON data layer) · classic: qs.json problem list + Propos/ proposition library + probability-driven scheduling + code heuristic scheduling;
  • vibe-math-v3 (third generation · paper-style md + planner agent + method library) · classic: all knowledge is stored and extended in Markdown paper/research-report form (Problems/ problem list + dependencies + source motivation, Progress/ research log, Propos/ proposition library, Methods/ general theory invention library, Verified/ absolutely trustworthy); before scheduling, the planner agent autonomously draws up a plan for the next N steps; theories/frameworks/tools/methods/ideas invented during solving are distilled by the Method Keeper into a reusable method system (as in inventing group theory or functional analysis).
  • vibe-math-v4 (fourth generation · resident self-organizing collaborative research) 🧪 Experimental: a group of persistent resident subagents leave messages for one another + hold meetings, and autonomously decide all task arrangements (no central scheduler); each accumulates its own progress/proposition/method/subproblem libraries and consults the others; verification is written to Verified/ only when all residents agree (true or false), otherwise it remains in the library with a probability attached; when the context reaches a threshold it automatically /compacts; it stops only when all agree that the original problem has been solved.
  • vibe-math-v5 (fifth generation · institute system) 🧪 Experimental · Latest: upgrades the residents into an institute — academicians (leaders / the organizing and coordinating center, responsible for decomposition and assignment, setting priorities, chairing meetings, and supervising progress) + resident researchers (with voting rights, able to autonomously hire/fire their own temp workers) + temp workers (no voting rights); it has a public charter, group chat and meetings, a compare-and-set task board, and real firing; a boolean agreement of ≥ m votes is required to write to Verified/ (opposing votes block, abstentions are not counted, and if the threshold is not met it remains in the library with an average probability attached); state is written to the hardened JSON file State/<institute>.v5state.json under the institute directory (serial writes, a mandatory load before read) — never into the host session log — at zero token cost.

After installing this plugin package (or manually copying the presets), four agent presets appear in DSH's preset selector.


🧭 Navigation: I want to … → start here


🧩 Architecture Diagrams (v2 + v3 + v4 + v5)

Static architecture diagrams; for the complete process description see the v1-era architecture notes (historical: the layout changed from v2 on) and the v5 detail diagrams (the full set of v5 detail diagrams); editable generation scripts: the Chinese v2/v3 posters come from the matplotlib scripts v2 / v3 (matplotlib → PNG); the English v2/v3 diagrams come from the zero-dependency Node scripts v2-en / v3-en (→ SVG); v4 / v5 (zero-dependency Node → SVG, node docs/generate_framework_diagram_v4.mjs; add --lang=en for 示例图/框架图-v4-en.svg). SVG is used from v4 onward: plain text, diff-friendly, crisp at any zoom; when PNG is needed, screenshot with a headless browser (the command is at the top of the generation script).

Vibe Math V2 (probability-driven · JSON data layer) · classic

Vibe Math V2 architecture diagram

One-sentence pipeline: qs.json takes problems by priority → Explorer splits out directions (if all are dead ends, re-derive) → one Solver per direction iterates over multiple rounds (lemmas go into Propos/, solutions go back to qs.json, all probabilities <1) → the scheduler picks r (proposition / proposition+proof·disproof / problem+solution) and dispatches ≥3 verifiers for independent review → debate → ruling → at probability=1 it automatically closes out (problem solved, proposition 1/0, priority set to never); state is written to disk throughout, resume continues from the checkpoint, and reportMode can report by file/push/both.

Vibe Math V3 (paper-style md + planner agent + methods library) · classic

Vibe Math V3 architecture diagram

One-sentence pipeline: all knowledge is stored and continued as Markdown papers/research reports (Problems/ problem list including dependencies and the source motivation of follow-up problems, Progress/ research log continued by direction and by round, Propos/ proposition library, Methods/ general theory invention library, Verified/ absolutely trustworthy) → before scheduling, the scheduler builds a state brief and calls the planner agent; the planner agent lays out the next N steps in one go (spawn solver/verifier/explorer/method-keeper, interrupt, promote, wait), which are executed after code validation (actions exceeding the concurrency limit are queued and consumed across ticks; a planning failure automatically falls back to the v2-style heuristic) → verifiers review independently → debate → near-consensus ruling (if on the same side and the mean is ≥0.85/≤0.15, take the mean, fixing v2's flat misjudgment) → at probability=1 it closes out and generates a Verified/ card → the solver's methods_used/new_inventions reports are distilled/refined into the methods library by the Method Keeper (which can form system hierarchies and be reused across projects).

Vibe Math V4 (resident self-organizing collaborative research) 🧪 Experimental

Vibe Math V4 architecture diagram

The SVG above is generated by a zero-dependency script: node docs/generate_framework_diagram_v4.mjs --lang=en (pure Node, no Python/matplotlib dependency; generation estimates text width, and any line overflowing its container raises a warning and exits with code 1).

One-sentence pipeline: initially N resident subagents are created (continuable, persistent context) which first brainstorm on their own and produce initial insights/directions → after that all task arrangements are decided autonomously by them leaving messages for each other + holding collective meetings (the framework only provides the message bus/meetings/task board/artifact persistence, and never assigns tasks); each resident persists valuable artifacts into its own Progress/<id>/, Propos/<id>/, Methods/<id>/, Subproblems/<id>/ libraries according to degree of value / planned motivation and use / its own probability estimate, and they can read each other's; verification is initiated by their own deliberation, and only when all residents agree (true or false) is it written to Verified/, otherwise it stays in the library with a probability attached; when a resident's context reaches a threshold (66% by default) it automatically /compacts; they stop only when all of them agree that the original problem is solved; residents can be manually intervened with/added/shut down at any time, and checkpoint resume is supported.

Note: V4 removes v3's central planner and deterministic roles (explorer/solver/verifier/planner/method-keeper) and makes the "researcher" itself the subject. See vibe-math-v4/实现方案.md for details. Keep-alive mechanism (tiered keep-alive A+B + deadlock watchdog): a gang idle for longer than activityTimeoutMs receives a self-driven CHECKPOINT (suggesting it continue solving/send a message/propose a task, rather than "do you want to stop"), and it fills in parallel — branch A fills as much of the maxParallel concurrency budget as possible in one go (waking several idle residents at the same moment, rather than the serial "wake only r1, then r2 after it finishes"), and mailbox delivery also reaches several idle recipients in parallel; a failed wake automatically re-arms the heartbeat; if the team is idle and has no new artifacts for longer than stallAutoMeetingMs (6 minutes by default), the framework automatically convenes a synchronous meeting so the residents can decide the next step themselves; if a meeting/verification hangs (still no new speech/votes after more than 2×activityTimeoutMs), the framework automatically abandons that meeting/verification and returns to normal self-organization, so that one broken meeting does not permanently block the whole team; meetings do not preempt verification — meeting requests while verification is under way are held and convened afterwards (keeping the consensus-consistent "truth-seeking" step from being interrupted by coordination discussion) — the framework always only facilitates and never assigns tasks.


Vibe Math V5 (institute system) 🧪 Experimental · Latest

One-sentence positioning: upgrade v4's "a group of residents messaging each other" into an institute — with three classes of staff: academician (leader), resident researcher, and temp worker; with the institute's public charter; with group chat and meetings; with autonomous hiring/firing; and where any conclusion must be given a Boolean probability of 1 or 0 unanimously by at least m voting members before it can be written to Verified/.

Vibe Math V5 architecture diagram

Image sources and all detail diagrams (member lifecycle, one-round sequence, consensus state machine, meeting flow, scheduling priority, state folding, prompt composition, task board, authority matrix): the v5 detail diagrams. The SVG above is generated by a zero-dependency script: node docs/generate_framework_diagram_v5.mjs --lang=en.

flowchart TB
    OFF["👤 Institute office (session root agent / human)<br/>does not research · does not vote · only reports and relays instructions"]
    subgraph INST["🏛️ Institute (internal autonomy: roster, organization and assignment all happen among members)"]
        ACAD["Academician acad —— leader / center of organization and coordination<br/>L1 institute-wide overview · L2 assignment · L3 priority<br/>L4 chairs meetings · L5 supervision · L6 moving people · L7 external"]
        RES["Resident researcher r-n<br/>has voting rights · may autonomously hire/fire its own temp workers"]
        TMP["Temp worker t-n<br/>no voting rights · hired temporarily for a specific task"]
    end
    subgraph FW["⚙️ Framework vibe-v5 —— only a medium (middleware), never assigns tasks"]
        M["Message relay · meetings/debates · task board CAS+DAG<br/>m-vote consensus verification · context and liveness · roster and hiring · scheduler"]
    end
    PROJ["💾 State/&lt;institute&gt;.v5state.json (hardened JSON, authoritative)<br/>12 kinds of events · pure fold applyV5Event · serial writes · recovery = load before read"]
    FS["📁 Members/&lt;id&gt;/* · Shared/* · Verified/ · Problems/"]
    RULE{{"Truth gate: Boolean unanimity and Boolean votes ≥ m = min(quorumCap, number of registered voting members)"}}
    OFF <-->|"vibe_v5_* / /v5 commands ↔ status / report"| M
    M <-->|"per-round prompt ↔ single JSON receipt"| ACAD
    M <-->|"per-round prompt ↔ single JSON receipt"| RES
    M <-->|"per-round prompt ↔ single JSON receipt"| TMP
    ACAD -.->|"assign / supervise / chair meetings (in-institute organization, not framework behavior)"| RES
    ACAD -.-> TMP
    M <--> PROJ
    M <--> FS
    M --> RULE

One-sentence pipeline: the institute office (session root / human) configure → starts the institute → the academician (center of organization and coordination) decomposes the original problem, assigns tasks, sets priorities, chairs meetings and supervises progress; resident researchers research on their own and hold the vote, and temp workers can be hired as needed (no vote, genuinely dismissible). The framework is only a medium and never assigns tasks; every conclusion needs ≥ m = min(quorumCap, number of registered voting members) Boolean votes, all on the same side, before it is written to Verified/ (an opposing vote blocks; abstentions do not count as votes but count toward the average; below the threshold the object stays in its library with its average probability); meetings and verification are mutually exclusive in both directions; state lives in the hardened JSON State/<institute>.v5state.json under the institute directory (zero token cost, serial writes, a mandatory load before read); the problem is concluded only when all voting members consider it solved.

Positions and authority, the detailed truth rules, operating mechanisms, state and persistence, prompt composition, the institute directory, the tool surface, and the differences from v4: see the v5 detail diagrams §13 "v5 notes moved from the README" (text preserved); parameters are in the parameter quick reference, and the final paper in docs/final-paper.md.


✨ Features

  • Final paper (on by default in all four presets): at closure the run writes a complete paper that only organises evidence it already has, and the paper phase runs before the run is marked complete; the artifacts are Paper/<id>/{paper.md,paper.tex,paper.pdf,paper.meta.json,paper.log.md}. Parameters: finalPaper (default true) / paperFormat (default both) / paperLanguage (default zh) / paperCompilePdf (default true) (plus paperEditor on v4/v5); the manual trigger is /vN paper [lang=] [format=] [editor=] [force]. A PDF needs a LaTeX engine on the host (xelatex preferred for Chinese); without one the tex+md are still delivered. Full contract and usage: docs/final-paper.md.
  • Multi-agent automatic solving: the main agent hands the problem to the scheduler, which dispatches explorer / solver / verifier (v2/v3) plus subagents such as planner (planning agent, v3) and method-keeper (method organizing agent, v3) to solve collaboratively; you do not need to operate node by node by hand.
  • Multi-agent cross-validation: every conclusion goes to ≥3 "strict reviewers" for independent review → debate (exchange group) → adjudication (v3 defaults to near-consensus adjudication: if on the same side and the mean is ≥0.85/≤0.15, take the mean, so that "0.9 vs 1" is not misjudged as 0.5).
  • Paper-style Markdown knowledge base (v3): the problem list (including dependencies between problems, and the causes and plans of descendant problems), research log, propositions, and method library are all written and continued in md paper/research-report style (when a direction is re-derived, the logs of the old direction are automatically archived and kept); only Verified/ and the objects that a verifier judged true/false are absolutely trustworthy, and all other md (including unverified assertions in the method library) serve only as empirical reference.
  • General theory invention library (v3): the theory systems/frameworks/tools/methods/ideas invented during solving (including empirical summaries) are reported via methods_used/new_inventions, and the Method Keeper consolidates them into Methods/ method cards (which can form a hierarchy of systems and be reused across projects), forming a systematic method–theory system just like "inventing group theory while solving equations".
  • Planner-agent scheduling (v3): before scheduling, the planner agent is invoked to autonomously choose the optimal scheduling scheme according to the actual situation (problem dependencies / survival rate / verifiable objects / concurrency budget / results of the last plan), arranging the tasks of each agent for the next N steps in one go; if planning fails, it automatically falls back to heuristics.
  • Knowledge accumulation: conclusions that pass verification are promoted into the Verified/ trustworthy knowledge base (v2/v3 additionally have the Propos/ proposition library) for reuse by later directions.
  • Checkpoint resume: the scheduling state, task stack, agent registry, decision queue, verifier historical accuracy, etc. are all written to disk; after a restart, resume restores them (v2/v3 use a process epoch to distinguish "pause → resume within the same process" from "restart across processes"; v3's md is itself the narrative breakpoint).
  • Mid-run manual intervention (and continue): the auto / manual modes can be switched at any time; in manual mode, a decision is suspended at key nodes and waits for your approve/reject/override (v3 adds a plan approval gate and a method promotion gate); you can send a message to / interrupt any subagent.
  • Per-project isolation: each mathematical problem is an independent project folder, with no interference between them, and you can switch at any time.
  • Multi-session parallel isolation: a DSH agent preset is a standing mount (all sessions of the same preset share one plugin instance), and inside the plugin all running state is isolated by root session id — two sessions can each run a project at the same time, their respective subagents are correctly attached under their own session, and the scheduler / parameters / decision queue / current project do not interfere with each other (v3 additionally has a project lock, so the same project is scheduled by only one session at a time). The current project is persisted per session (VibeMath/current.<session id>.json).
  • Adjustable subagent permissions: you can restrict the tools a subagent is allowed/forbidden to use and the per-round cap on external tool calls, and explicitly tell it that it may read Verified/, Propos/, Methods/, Reliable/ and the progress log.
  • Configurable: vibe_math_setting.json (with comments) for customizing default parameters; /vibe setup for interactive question-and-answer configuration.
  • Natural-language control: the main agent acts as "assistant + reporter" — you state your needs in plain words, and it calls the tools, reports progress, and configures parameters on its own.

Specific to v4 / v5:

  • Persistent self-organization (v4): at the start, N persistent resident subagents are created, after which all task arrangements are decided by those subagents themselves through leaving messages for each other + holding meetings (the framework only provides the message bus/meetings/task board, and never assigns tasks).
  • Institute system (v5): on top of v4's self-organization, it introduces the organizational form of a real research institute — the academician (leader) is responsible for decomposition, assignment, prioritization, chairing meetings, and supervising progress; resident researchers have voting rights and can autonomously hire/fire their own temp workers; temp workers have no voting rights; all organizational actions are performed by members of the institute, and the framework still only acts as the medium. See the Vibe Math V5 section above for details.
  • Adjustable quorum (v5): for an object to enter Verified/, ≥ m = min(quorumCap, number of enrolled voters) voters must cast a consistent boolean vote (all 1 or all 0); an opposing vote blocks, and abstentions do not count as votes but do count toward the average probability; if the threshold is not reached, the object is kept in the library with the average probability and the complete debate record, without forcing an adjudication.
  • Zero-token-cost state persistence (v5): the institute state is written to the hardened JSON file under the institute directory (State/<institute>.v5state.json) and does not enter the model context; cross-process and same-process recovery go through the same code path (a mandatory load before read).
  • Genuinely reversible roster (v5): hiring creates a resident child session; firing cancels in-flight turns, releases the child session, reclaims its tasks, and discards undelivered mail; a codename is never reused.
  • Human-readable mirror (v4/v5): the roster table, task board, meeting minutes, debate records, and closing records are all written to disk as Markdown and are readable by humans at any time; but the authoritative state is not in these files (in v5 it is State/<institute>.v5state.json), so manually corrupting them will not break the institute.

🧮 Lean formal verification (shared by the four architectures, adjustable switch)

What it changes is not "being a bit stricter" but the object of review itself. m agents agreeing that "this is right" is still consensus — it cannot rule out shared misunderstanding; passing Lean is machine checking. The only remaining uncertainty is thus reduced to a question a single person can review effectively: are the definitions / objects / conditions / assumptions / conclusions in the Lean code fully consistent with the original text of the proposition (fidelity)?

  • Switch formalVerify (same name in all four architectures, default 'off'): 'off' is a true no-op (no Lean content appears in the prompts, no formalization state is written, and the verification flow and gates are completely unchanged; the five tools stay registered and usable and the persona always lists them, otherwise the switch would be undiscoverable and could not be turned on); 'encourage' is encouraged but not mandatory (decide by implementation difficulty whether to formalize; once Lean passes, the focus of review shifts to fidelity; no gate); 'require' is mandatory — a true/false conclusion must satisfy "Lean has passed" or "the blocking reason was explicitly recorded", otherwise the adjudication does not take effect (recorded as undecided, reason formal-required, written into the "formalization TODO"). Related parameters: leanCommand (default lean), leanArgs (used with lake env lean), leanTimeoutMs (default 120s).

  • ⚠️ A fidelity defect ≠ the proposition is false (important): passing Lean only guarantees that "this piece of code passed the kernel"; it does not guarantee that it says what the proposition means to say. When a voter finds the Lean code and the original proposition inconsistent (too narrow / too broad / a different object / a missing condition): do not vote 0 (0 means "the proposition is false" — that would record "the formalization does not qualify" as "the proposition was disproved", and under v5's all-0 consistency rule it would even write the proposition into Verified/ marked false, so a mechanism meant for truth-seeking would fabricate a wrong negative conclusion); instead vote a value strictly between 0 and 1 (an abstention) and record the deviation with the receipt formal:{decision:'defect', note:'<specific deviation>'}. The framework then revokes the "passed" state of this proof (downgraded to attempted, Verified/Lean/<id>.lean deleted or rewritten as a "withdrawn" note if the host cannot delete it, and written into the "formalization TODO"), and under require this adjudication is not concluded (the encourage setting has no gate, so you must not claim the framework will force a shelving). Vote 0 only when the voter, independently of this Lean code, can also determine that the proposition is false (and can give independent reasons).

  • The receipt channel and three further hard requirements (members who do not call the Lean tools can also leave a judgment; mandatory under require): the receipt field is "formal": {"target":"<object id>", "decision":"used|blocked|defect", "file":"Formal/<object id>.lean", "note":"difficulty judgment/blocking reason/specific deviation"}; when decision='blocked'/'defect', note is required (if missing the whole entry is rejected), used only records the object as attempted, and under off this channel is disabled (otherwise off would not be a true no-op). Three more: tool names in the injected text are always given in full (<prefix>_lean_archive, not lean_archive — an abbreviation is not a registered name); before archiving a reusable definition/lemma run it through first, and if it does not run through it must not enter the library; when the toolchain is missing (LEAN_NOT_FOUND cannot resolve the executable / the host has no subprocess service), write the code down, archive it, and state "the host has no Lean toolchain" in note — this counts as an explicit blocking reason and the gate lets it through.

  • Where the artifacts go: <VibeMath root>/Formal/Lib|Proved/ are the cross-project reusable definitions and already-proved lemmas (each with an Index.md, to be checked before writing a new definition); inside a project, Formal/<object id>.lean is the object's working file (next to an Index.md and, under require, a TODO.md), and Verified/Lean/<object id>.lean is the archived proof (in v5 these live under Projects/<project>/Institutes/<institute>/).

  • The five tools (one set per architecture, prefix following each one's naming): <prefix>_lean_run executes Lean on the host subprocess service (with leanAsync=true it enqueues and returns {async:{jobId,state}}; otherwise it returns {ok, exitCode, ms, stdout, stderr} synchronously) and never throws (missing toolchain → LEAN_NOT_FOUND, timeout → LEAN_TIMEOUT, path escape → rejected); <prefix>_lean_archive: kind='def'/'lemma' archives into the cross-project Formal/Lib|Proved (identical content is deduplicated), kind='proof' writes Formal/<target>.lean and only a settled(ok) job writes Verified/Lean/<target>.lean and marks the object Lean-passed, kind='blocked' records an explicit difficulty judgment/blocking reason (reason required); <prefix>_lean_lib rebuilds and returns the three indexes, the per-object status and the background jobs — check for duplicates and reuse before writing a new definition; <prefix>_lean_read reads one archived file back verbatim (only Formal/Lib|Proved, 64KB cap); <prefix>_lean_job inspects or waits for a background compile (read-only). For example, v5 is vibe_v5_lean_*, v2/v3 are vibe_math_lean_*, and v4 is vibe_v4_lean_* (each with run / archive / lib / read / job).

  • Boundaries (intentional): the framework does not bundle Lean (no toolchain installation, no dependency downloads; when the toolchain is missing it degrades gracefully and records this faithfully); it does not judge fidelity (that is what agents/humans review and vote on; the framework is only responsible for switching the focus of review to fidelity); Lean passing ≠ the proposition is true — it only means "this piece of formalized code passed the kernel check".

For the complete contract (parameters, paths, state transitions, prompt semantics, gate locations, index format, test requirements), see docs/formal-verification.md. The personas (the prompt the main agent receives) of all four presets fully list these five tools and eight parameters, and the prefix and text blocks are identical line by line (only line 0 may differ); this layer is guarded by audit-persona-surface.test.mjs and audit-persona-sensitivity.mjs — when this feature was added, it was precisely in the four presets that the defect "the tools were registered but the persona never listed them" was found (see the release notes shipped with the package).


🚀 Installation

Two installation methods, choose either one (they can also coexist):

Method A: one-click install as a plugin package (recommended, installs all four presets at once)

On the desktop app, install dsh-vibe-math from Settings → Plugins; on the command line, use a profile:

dsh plugin --profile <your profile> add dsh-vibe-math
# Or install directly from GitHub:
dsh plugin --profile <your profile> add github:ChongCyrus/Vibe-Mathematics

Then start a new session and pick Vibe Math V2 (v2, classic), Vibe Math V3 (v3, classic), Vibe Math V4 (v4, resident self-organization) or Vibe Math V5 (v5, institute system) in the preset picker — the four architectures are peers, choose according to your actual needs (see "How to choose"). The two DSH generations land in different places, and this package adapts to both:

  • DSH ≥ 0.1.7 (current): agent presets are declared as composition rows. This package declares all four presets in cordis.patch.yml (each row hands that preset's full plugin list to the host's agentPresets service), and writes nothing into ~/.dsh/.agent-presets/ — that directory has not been read since 0.1.7.
  • DSH ≤ 0.1.6: presets are still directories, and the installer writes the four presets into ~/.dsh/.agent-presets/ (vibe-math-v2/ … vibe-math-v5/); the declaration rows in the same cordis.patch.yml register nothing on those versions, so older hosts do not error at boot.

After upgrading the package version, restart DSH: the managed content is replaced wholesale with the new version's bytes — including files you edited by hand. This is intentional: a preset that is "half old version, half new version" will fail to mount or behave strangely, and you cannot tell from the outside. The hand edits that get replaced are not lost: in the directory form the original text is first backed up to ~/.dsh/.agent-presets/.vibe-math-backup/<old version>/<preset>/, and the file names are listed in the log. If you want to customize a preset, do not edit managed content — make a copy (the copy action in the preset picker, or declare it under a new id in your own bundle). Note: on current DSH, if you save your own declaration for the same id, this package skips registration and logs one line saying so — your declaration wins.

Method B: manual declaration (customization / secondary development)

  • DSH ≥ 0.1.7: put the whole of vibe-math-vN/agent.cordis.yml into a declaration row as its plugins — either copy the dsh-vibe-math/preset-declaration row from this package's cordis.patch.yml, or use the host's own @deepseek-ai/dsh-agent-preset row; install this bundle into your profile. Full rules: the DSH skill editing-cordis-compositions.
  • DSH ≤ 0.1.6: copy agent.cordis.yml / preset.yml / vibe-math-vN.js from vibe-math-vN/ into ~/.dsh/.agent-presets/vibe-math-vN/.

Then start a new session and select "Vibe Math V2" / "Vibe Math V3" / "Vibe Math V4" / "Vibe Math V5" in the preset picker; once the session starts it is ready to use: v2/v3 tools are vibe_math_*, v4 is vibe_v4_*, v5 is vibe_v5_*; typing /vibe, /v4, /v5 in the input box gives autocompletion.

After changing a preset definition you must restart the DSH process before starting a new session (a preset's standing mount is cached until the process exits).

DSH version adaptation and dependencies

  • Supported range: 0.1.2-alpha.4 … 0.2.0-rc.2 (11 versions declared compatible, tested target 0.2.0-rc.2; engines.node is ^22.19.0 || >=24.0.0) — per-version declarations, the engines.dsh range and peerDependencies live in package.json (dsh.compatibility.dshReleases).
  • Delivery form: DSH ≥ 0.1.7 uses composition-row declarations (the agentPresets service); DSH ≤ 0.1.6 uses the ~/.dsh/.agent-presets/<id>/ directory + preset picker — one bundle covers both lines.
  • Hard constraint: v5 keeps its institute state only in State/<institute>.v5state.json and never in the host session log — on an unknown event type DSH refuses to load the whole session, so it will not open on the next resume.
  • Everything else (host rows and required services, the startup self-check and capability gate, the live-resident cap, the three 2026 fixes, the dsh.bundle.patch constraint, the upgrade path and peerDependencies): see docs/COMPAT-AUDIT-ROUND2.md; to upgrade this package use dsh plugin --profile <your profile> add dsh-vibe-math@latest (use add, not update).

🧭 How to choose among the four presets

💡 All four architectures are peers; choose according to your actual needs:

  • Choose vibe-math-v2 (probability-driven · JSON data layer) if you:
    • prefer structured JSON data (qs.json / Propos/<category>_Propos.json / Verified/ cards), convenient for programmatic retrieval and further processing;
    • want mature and stable code-heuristic scheduling (priority + probability, predictable behavior, not dependent on the planner agent's "improvisation");
    • do not need method library accumulation / paper-style narration, and data being field-oriented is enough.
  • Choose vibe-math-v3 (paper-style md + planner agent + method library) if you:
    • prefer a paper/research-report-style natural-language knowledge base (the problem list includes dependencies and the source motivation of follow-up problems, the research log is continued by direction and by round, human-readable and freely extendable);
    • want scheduling to be planned autonomously by the planner agent as an N-step plan based on the actual situation (more flexible, automatically falls back to heuristics on failure);
    • want a general theory invention library — theories/frameworks/tools/methods/ideas invented during solving are accumulated by the Method Keeper into a reusable, systematizable, cross-project-extendable methodology (like "inventing group theory while solving equations");
    • accept the trust layering of "only Verified/ is absolutely trustworthy, the rest of the md is empirical reference".

Both are mature and usable, continuously maintained, and both support checkpoint resume, manual/automatic intervention, progress reporting, multi-session isolation, proposition promotion, near-consensus/weighted adjudication and other core capabilities; the switching cost is low (the same set of vibe_math_* tools and /vibe commands, the same parameter system).

  • Choose vibe-math-v4 (resident self-organization) if you want a group of persistent resident subagents that message each other and hold meetings, fully self-organizing (no leader, no central scheduling), and can accept a strict threshold of "a conclusion requires unanimity".
  • Choose vibe-math-v5 (institute system) if you want:
    • organized self-organization — like a real research institute, with a leader (academician) responsible for decomposition, assignment, prioritization, chairing meetings and supervising progress, but judgment still belongs to each individual;
    • a roster that can grow or shrink — resident researchers + temp workers who can be autonomously hired/dismissed (temp workers have no voting rights, suitable for chores such as checking, trial computation and material organization);
    • an adjustable consistency threshold — m = min(quorumCap, number of voting members) boolean-consistent votes settle the matter (because an opposing vote blocks, the effective threshold is still "every Boolean vote points the same way"; it differs only when fewer than m members cast a Boolean vote);
    • zero-token-cost state persistence — the institute state is written to a hardened JSON file under the institute directory and does not consume member context budget.

⚠️ vibe-math-v2 and vibe-math-v3 are the classic architectures; vibe-math-v4 and vibe-math-v5 are experimental architectures, all four are peers and all are selectable; the old vibe-math-v1 has been removed (this package contains only v2/v3/v4/v5).

v2 (probability-driven · classic) v3 (paper-style md · classic) v4 (resident self-organization · experimental) v5 (institute system · experimental)
Positioning Classic (JSON data layer) Classic (third generation) Experimental (fourth generation) Experimental (fifth generation)
Core idea Probability-driven: qs.json problems + Propos/ proposition library, scheduled by "correctness probability / value" Paper-style md knowledge base + planner agent scheduling + general theory invention library Persistent resident subagents self-organizing: message each other + meetings decide all tasks, no central scheduling Institute: the academician organizes and assigns, members research on their own; a conclusion requires ≥ m boolean-consistent votes; temp workers can be hired as needed
Data qs/qs.json + Propos/<category>_Propos.json + Reliable/ Problems/ + Progress/ + Propos/ + Methods/ (all md, soft-spec anchors + free narration) + Verified/ Problems/ + `Progress Propos
Roles explorer → per-direction solver → verifier planner (planner agent) → explorer → per-direction solver → verifier → method-keeper (method organization agent) N resident researchers (continuable), no fixed roles academician acad (leader) + resident researcher r-n (with voting rights) + temp worker t-n (no voting rights, hireable and dismissible) + institute office (does not research and does not vote)
Scheduling Code heuristics (priority + probability) The planner agent produces an N-step plan (executed after validation, falls back to heuristics on failure) No central scheduling: tasks arise from residents messaging each other / holding meetings (the framework is only a medium and does not assign) The framework still does not assign; the academician decomposes/assigns/prioritizes/supervises, and members may object with reasons; the framework only relays, keeps the task board, holds meetings and counts
Closing rule A solution/proof reaching probability 1 closes it; never is never scheduled Same as v2 (near-consensus adjudication fixes flat misjudgment) Writes to Verified/ only when all residents agree (true or false), otherwise leaves it in the library with a probability Writes to Verified/ only when boolean votes ≥ m = min(quorumCap, number of voting members) and all are 1 or all are 0; opposing votes block; abstentions do not count as votes but count toward the average; (can be switched back to the v4 criterion)
Stopping All solved or stuck All solved / no candidates Stops only when all residents unanimously agree the original problem is solved Same as v4: concluded only when all voting members unanimously agree the original problem is solved
Context None None Resident context automatically runs /compact on reaching the threshold (adjustable) Same as v4 (threshold/round count adjustable; after compaction the charter still takes effect in the persona)
Ad hoc capabilities Automatic promotion of proposition "value/criticality" to the problem list; reportMode file/push/both; priorityAdjust Method library accumulation loop (methods_used/new_inventions → Method Keeper); plan approval gate/method promotion gate; project lock; follow-up problem "source and motivation" as a first-class citizen Residents accumulate individually + read each other; unanimous verification; add/close residents at any time, intervene by message; checkpoint resume All v4 capabilities, plus: real hiring/dismissal (releases sub-sessions, reclaims tasks); compare-and-set task board + dependency DAG; strict mutual exclusion between meetings and verification; roster mirror and conclusion records; zero-token-cost state

All four support: checkpoint resume (vibe_math_resume / vibe_v4_resume / vibe_v5_resume), manual intervention and pause/resume, per-project isolation, subagent permission control, and natural-language driving. All four architectures are peers — choose v2 if you prefer structured JSON data and deterministic scheduling, choose v3 if you prefer paper-style md, the planner agent and the theory invention library; v4 is fully self-organizing resident collaborative research, and v5 is an institute system with "a leader + a roster that can grow or shrink + an adjustable threshold".


🧠 Architecture and division of labour (all four in parallel)

V2 (probability-driven · classic)

The framework = main agent + code scheduler + explorer / solver / verifier subagents. The scheduler is the sole master control and the sole file writer (subagents only return structured JSON and never write files); it decides the next step by code heuristics over "correctness probability + value/criticality + priority"; each verification object goes to ≥3 independent "harsh reviewers" for review → debate → adjudication (a solution/proof reaching probability 1 closes it; never is never scheduled). Pick it if you prefer structured JSON (qs.json / Propos/) and predictable, deterministic scheduling that does not depend on a planner agent.

Division of labour in one sentence: the main agent handles "talking to people", the scheduler handles "execution and boundary-keeping", and the subagents handle "thinking".

V3 (paper-style md + planner agent + method library) · classic

The framework = main agent + code scheduler + planner agent + explorer / solver / verifier / method-keeper subagents. Each time a dispatch is prepared, the planner agent reads the state brief (problem dependencies, survival rate, verifiable objects, concurrency budget, result of the last plan) and autonomously draws up an N-step plan, which the scheduler validates before executing (falling back to heuristics on failure); theories/frameworks/tools/methods/ideas invented during solving are reported via methods_used/new_inventions and distilled by the Method Keeper into new method cards in the Methods/ general theory invention library. Pick it if you prefer a paper-style md knowledge base, want more flexible scheduling, and want a systematizable, cross-project method library.

Division of labour in one sentence: the main agent handles "talking to people", the planner agent handles "setting the plan", the scheduler handles "execution and boundary-keeping", the subagents handle "thinking", and the Method Keeper handles "depositing inventions into theory".

V4 (resident self-organization · experimental)

The principal is N persistent resident subagents (continuable): there is no central scheduling and no leader — task arrangements emerge from residents messaging each other + holding meetings (the framework only provides the message bus/meetings/task board and never assigns); each accumulates its own Progress/Propos/Methods/Subproblems library owned by resident id and may read the others'; verification is written to Verified/ only when all residents agree (true or false), otherwise it stays in the library with a probability; when the context reaches a threshold it automatically /compacts; it stops only when all agree that the original problem is solved (vibe_v4_resume resumes from a checkpoint). Pick it if you want fully self-organizing research and can accept the strict "unanimity before a conclusion" threshold.

V5 (institute system · experimental · latest)

It upgrades v4's residents into an institute: academician (leader / organizing and coordinating centre: decomposition, assignment, prioritization, chairing meetings, supervising progress) + resident researchers (voting rights, may autonomously hire/fire their own temp workers) + temp workers (no voting rights); the framework still only relays, keeps the task board, holds meetings and counts — it never assigns tasks. The conclusion threshold is ≥ m = min(quorumCap, number of enrolled voting members) boolean-consistent votes (opposing votes block; abstentions do not count as votes but count toward the average; below the threshold the object stays in the library with its average probability); state lives in State/<institute>.v5state.json (zero token cost); meetings and verification are strictly mutually exclusive. Pick it if you want "organized self-organization", a roster that can grow or shrink, and an adjustable threshold.

For the complete v5 architecture (member lifecycle, one-round timeline, consensus state machine, meeting flow, scheduling priority, state folding, prompt composition, task board, authority matrix), see the v5 detail diagrams; for the textual specification see the v5 specification.


📁 Directory Structure

Only the top-level and high-value paths are listed; for the full meaning of every file see each preset's 实现方案.md (v2 · v3 · v4 · v5) and the v5 detail diagrams.

v2 (probability-driven · classic)

<VibeMath>/Projects/<project>/
├─ qs/qs.json                   # problem list (overview/solved/solutions + correctness probability/priority)
├─ Propos/<category>_Propos.json # proposition library (boolean estimate/proof·disproof/value·criticality)
├─ Verified/                    # concluded-fact index
├─ Reliable/                    # trusted references (read-only, placed by the user)
└─ VibeMath_State/              # scheduler-private persistent state (for checkpoint resume)

v3 (paper-style md + planner agent + method library) · classic

<VibeMath>/
├─ Methods/                     # [global] cross-project general theory invention library (promoted from project level)
├─ Formal/{Lib,Proved}/         # cross-project re

…

Content from the project README on GitHub ↗

Comments

Comments live in GitHub Discussions. Sign in with GitHub to post or react.