DeepSeek Harness Plugin Hub

Publish and manage complete Harness Profiles. Discover Plugins for your next setup.

Explore

PluginsPresetsDocsNews

Community

Publish a pluginContactReport an issue

Resources

Plugin Hub on GitHubDeepSeek HarnessSystem statusPrivacy notice
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

Independent and unofficial. Not affiliated with, authorized by, or endorsed by DeepSeek.

Voice — DSH Plugin for DeepSeek Harness
DeepSeek Harness Plugin Hub
ProfilesPluginsCategoriesNewsDocsSign inManage Profiles
ProfilesPluginsCategoriesNewsDocsSign in
← Plugins
V

@nn12138/dsh-voice

Voice

Voice input plugin for DeepSeek Harness: microphone → local/browser speech recognition → text submitted as a normal chat message (input-only, preset-agnostic)

The plugin will be installed here. Keep web if you are unsure.

npx -y @deepseek-ai/dsh plugin --profile web add @nn12138/dsh-voice@0.2.6
READMECompatibilityVersions

Compatibility and provenance

Voice is published as @nn12138/dsh-voice and currently resolves to version 0.2.6. The Hub verifies its manifest and preserves the exact installation source for reproducible installs.

DSH compatibility
*
Runtime surfaces
web
Release source
npm
Registry updated
8/21/2026

Versions

0.2.6stable
8/21/2026
0.2.5stable
8/21/2026

Related plugins

Loading related plugins…

Latest
0.2.6
DSH
*
HMR
Process restart
Tree shaking
Safe tree shaking not declared
Unpacked size
161.3 kB
Files
44
Surface
web
License
MIT
Source
npm
GitHub
★ 7
Weekly downloads
0
Last push
8/21/2026
View source ↗
README badge

Click the badge to copy Markdown for your README.

Do you maintain this Plugin?Claim benefit · Priority security scan

Verify the GitHub repository declared in package.json to manage this listing. After you claim it, Hub will prioritize a security scan of the current version and publish the result when it passes.

Claim this Plugin →
Report an issue

Related plugins

More verified plugins in vision-media.

Tool Describe Image@linxin666/dsh-tool-describe-imageModel-facing describe_image tool for the dsh web GUI: gives a text-only model image understanding by asking a vision-language model at an OpenAI-compatible endpoint to describe one image (local path, http(s) URL, or attachment reference). Hot-pluggable — Modlens@liustack/modlensPlug-in vision for text-only LLMs, powered by the free Antigravity CLIVision Toolkit@anionex/dsh-vision-toolkitDeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, OCR, grounding, UI restoration, pixel diff, Artifacts, and Web UI.Deepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio and 27 editable video templates as a native DeepSeek Harness conversation view.

README

dsh-voice 🎤

English | 中文

A voice input plugin for DeepSeek Harness: click 🎤 in the web UI (or press a hotkey), speak, and the recognized text is submitted as a normal chat message. Input only — it never touches the agent preset/persona, so it behaves like "another input method" in every mode.

Features

  • 🎤 Voice input: microphone button (platform design-system UI) + configurable global hotkey (Ctrl+Space by default)
  • ⚡ Live recognition: streaming partial transcripts are echoed while you speak; VAD finalization commits on stop (0.6s tail padding keeps sentence endings)
  • 🧠 Adaptive dual engine: host-native ASR (sherpa-onnx-node zipformer2 + silero VAD, offline/private) with automatic fallback to browser Web Speech (zero extra dependencies)
  • 🔌 Preset-agnostic: does not touch the persona/system prompt; works with code/standard/minimal/custom presets
  • 📦 Optional models: zero-config out of the box; run dsh-voice-models when you want native offline recognition

Installation

# Plugin
dsh plugin --profile web add @nn12138/dsh-voice

# Optional: offline native recognition (the plugin does not auto-install this runtime)
dsh plugin --profile web add sherpa-onnx-node
dsh-voice-models            # one-shot model download (~100MB) → ./dsh-voice-models

# Optional: configuration (edit ~/.dsh/profiles/web/cordis.patch.yml)
- id: voice
  config:
    modelDir: './dsh-voice-models'   # native ASR model directory
    hotkey: 'ctrl+space'             # global hotkey
    vadThreshold: 0.3                # lower = less clipping at sentence boundaries
    tailPadSeconds: 0.6              # tail-padding duration
    engine: auto                     # auto (default) | native | browser

The row-level config is received by the host half. engine and hotkey are synced to the browser half over the /voice.config loopback RPC, so there is no separate client config to write. auto probes host native capability: with a model it uses native; without one it falls back to Web Speech, so zero-config users keep working. Restart dsh web after changing the config.

dsh web   # 🎤 button appears on the left of the composer, or press Ctrl+Space

See USAGE.md and INSTALL.md (Chinese) for details.

How it works

Browser captures mic audio (auto-resampled to 16 kHz)
  → PCM base64 chunks (256 ms) → /voice RPC channel (loopback)
  → host: silero VAD + zipformer2 streaming decode
  → partials returned per chunk (live echo) / finals committed
    (VAD segmentation + 0.6s tail padding)
  → conversation service submits the text (same path as typing)

Engine selection: the host resolves the effective engine (config + model-load result) and the client consumes it via /voice.ping — native unavailable falls back to browser Web Speech. /voice.config carries the row-level engine/hotkey from host to client.

Development

pnpm install --ignore-workspace        # standalone deps (no DSH monorepo needed); no install-time scripts
pnpm --ignore-workspace test           # unit tests (including real-model smoke tests)
pnpm --ignore-workspace typecheck      # type check
pnpm --ignore-workspace build          # build (tsc host half + tsdown client half)
pnpm --ignore-workspace verify:package # pack-level checks: scripts-free install, complete runtime files

The build only ever runs on the publisher's side (prepack — at npm pack/npm publish time) and in CI before release; consumers installing @nn12138/dsh-voice from the registry never execute lifecycle scripts, so --ignore-scripts installs are complete and usable (see issue #2).

Real-model smoke tests look for the local voxelf assets and skip when absent; override with: DSH_VOICE_MODEL_DIR (model directory) / DSH_VOICE_TEST_WAV (test wav) / DSH_VOICE_DOWNLOADED_MODELS (downloaded model directory).

Layout: src/index.ts (host half) / src/client/ (browser half) / src/core/ (recognition core) / tools/ (wire-protocol smoke tools + model downloader).

License

MIT