Skip to content
dsh-market Browse plugins GitHub 中文

fenglin-ai/dsh-funasr-voice

Offline voice input for the DSH Web UI: mic to local FunASR (SenseVoiceSmall), one-click install, no cloud.

Stars ★ 1 Category Voice & Audio Listed 2026-08-24 npm dsh-funasr-voice

Install

Inside DeepSeek Harness, with dsh-market

dsh plugin --profile web add dshmarket

Or from the command line

dsh plugin --profile web add dsh-funasr-voice

Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.

README

中文 | English

A local, offline voice input plugin for the DeepSeek Harness (DSH) Web UI. The browser captures the microphone, the host half lazily launches a local FunASR + SenseVoiceSmall recognizer, and the recognized text is written into the composer in real time. No cloud, no network, no uploads.

⚠️ macOS users: DSH Desktop 2.0.2 has a microphone-permission bug that yields a silent stream and looks like "it records but no text appears" (a silent failure). You must patch DSH Desktop first — see "Known bug & fix" below.

Features

  • Click to start, click again to stop — continuous dictation with VAD auto-segmentation; text appears as you speak.
  • One-click install — click Install in settings to create a venv, install funasr + torch, and download the SenseVoiceSmall model (~2–3 GB) automatically.
  • Recording timer — the elapsed time is shown beside the mic button.
  • Fully offline — ASR runs on the local CPU (SenseVoice RTF ~0.07, ~10x real time).
  • Auto-fill / auto-send — recognized text goes into the draft; optional auto-send after recognition.
  • Leading/trailing SenseVoice event & emotion emojis (👏🎼😊 …) are stripped.

Screenshots

Voice input (click the mic → speak → text fills the composer, with a timer):

Voice input

Configuration (Settings → General → Voice input, with auto-detect and manual paths):

Configuration

How it works

browser mic ──AudioWorklet──▶ VAD segmentation ──POST /funasr-voice/asr (raw f32 PCM @16 kHz)
                                           │ host half proxies
                                           ▼
                            local FunASR Python service (SenseVoiceSmall)
                                           │
             composer ◀── recognized text ─┘
  • The browser only records and does endpoint detection; ASR runs entirely in a local Python process on the host.
  • The host half lazily launches the FunASR service (first recognition takes ~2–3 s to load the model, then it stays resident).
  • Audio capture uses AudioWorklet (not the deprecated ScriptProcessorNode).

Requirements

Dependency Notes
DeepSeek Harness Desktop 2.0.2 (the version this plugin was developed against)
Python venv + FunASR 1.4.x + torch the local ASR engine
SenseVoiceSmall model loaded locally, offline
Node ≥ 20 host half

⚠️ DSH Desktop 2.0.2 has a microphone-permission bug — without a fix the mic yields a silent stream and recognition never works. See "Known bug & fix" below.

One-click install (recommended)

Open DSH Settings → General → Voice input and click Install. The plugin will:

  1. detect python / python3 and pick a usable one as the venv base;
  2. create a venv under the install directory (default ~/.dsh/funasr-voice) and install torch + funasr;
  3. download the SenseVoiceSmall model to ~/.dsh/funasr-voice/models/SenseVoiceSmall (~1 GB).

Progress and logs are shown live, and the paths are filled in automatically. On Desktop, clicking Install first opens a native directory picker; on Web you type the directory instead.

Alternative: manual install

If you already have a FunASR environment or want to control the location yourself. Example for macOS:

# 1. create a dedicated venv
python3 -m venv ~/Documents/ASR/funasr/.venv
source ~/Documents/ASR/funasr/.venv/bin/activate

# 2. install FunASR (pulls torch, modelscope, etc.; several GB of disk)
pip install --upgrade pip
pip install funasr

# 3. download the SenseVoiceSmall model (~1 GB; pick one)
mkdir -p ~/Documents/ASR/funasr/models
cd ~/Documents/ASR/funasr/models

#    A: ModelScope (faster inside China)
python -c "from modelscope import snapshot_download; snapshot_download('iic/SenseVoiceSmall', local_dir='SenseVoiceSmall')"

#    B: HuggingFace
pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('FunAudioLLM/SenseVoiceSmall', local_dir='SenseVoiceSmall')"

Afterwards:

  • python path: ~/Documents/ASR/funasr/.venv/bin/python
  • model directory: ~/Documents/ASR/funasr/models/SenseVoiceSmall (must contain model.pt and config.yaml)

Fill these into the plugin's python / model config (below), or use the Auto-detect button in the plugin settings.

Install

  1. Install the plugin:

    dsh plugin --profile <your-profile> add <this-repo-or-path>
    
  2. Restart DSH, open Settings → General → Voice input, and click Install to set up FunASR + the model automatically (or use the manual path above).

  3. Click the mic icon on the right side of the composer. Grant microphone permission when asked.

Configuration

Two ways to configure python / model:

① GUI (recommended) — DSH Settings → General → Voice input, click Install to set up FunASR + the model and fill the paths automatically; the button then shows "✓ Ready". Persisted to ~/.dsh/dsh-funasr-voice.json, which takes precedence over YAML config.

② YAML — the plugin source ships no personal paths (python: python3, model empty by default). Override them in your profile's cordis.patch.yml:

# ~/.dsh/profiles/<your-profile>/cordis.patch.yml
- id: dsh-funasr-voice
  config:
    python: /your/FunASR/.venv/bin/python
    model: /your/SenseVoiceSmall/directory

Configurable fields:

Field Default Notes
enabled true master switch
basePath /funasr-voice HTTP route prefix
python python3 FunASR venv python (required)
model (empty) SenseVoiceSmall directory (required)
port 18765 local ASR service port
language auto auto | zh | en | ja | ko | yue
useItn true inverse text normalization
autoSend false auto-send after recognition

Known bug & fix: DSH Desktop 2.0.2 microphone permission

Symptom — DSH Desktop 2.0.2's Electron main process does not handle the getUserMedia media permission. The browser receives a silent stream and macOS never shows the "want to access the microphone" prompt — the plugin looks like "it records but recognizes nothing, no text appears".

Root cause — dsh-plugin-desktop/src/main.ts is missing a session.setPermissionRequestHandler and never calls systemPreferences.askForMediaAccess('microphone') to trigger the macOS prompt.

Fix — patch the DSH Desktop source (dsh-plugin-desktop/src/main.ts, after app.whenReady()):

import { session, systemPreferences } from 'electron'

session.defaultSession.setPermissionRequestHandler((_wc, permission, callback, details) => {
  if (permission !== 'media') {
    callback(true)
    return
  }
  const mediaTypes = 'mediaTypes' in details ? details.mediaTypes : undefined
  const checks: Promise<boolean>[] = []
  if (!mediaTypes || mediaTypes.includes('audio')) checks.push(systemPreferences.askForMediaAccess('microphone'))
  if (!mediaTypes || mediaTypes.includes('video')) checks.push(systemPreferences.askForMediaAccess('camera'))
  void Promise.all(checks).then(results => callback(results.every(Boolean)))
})

Rebuild and relaunch; clicking the mic then shows the system permission dialog. Alternatively, wait for an official DSH Desktop release that fixes this.

Structure

.
├── lib/index.js       # host half: launch/proxy FunASR + HTTP routes
├── lib/client.js      # client half: mic button + VAD + timer + settings
├── asr_server.py      # FunASR (SenseVoice) HTTP service, launched by the host
├── cordis.patch.yml   # bundle patch (configuration)
└── package.json

Troubleshooting

  • "server exited during startup" — the python / model paths are wrong; check the DSH log for the captured stderr.
  • Records but no text, no permission prompt — see "Known bug & fix" above.
  • Slow first recognition — model + torch initialization takes ~2–3 s, then stays resident.

License

MIT

Content from the project README on GitHub ↗

Comments

Comments live in GitHub Discussions. Sign in with GitHub to post or react.