DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Voice — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins

@goodandready/dsh-voice

Voice

DeepSeek Harness 的语音输入:根据停顿将听写内容分块,并支持语音消息;每种输入都有自己的提供商回退链(Deepgram、Groq、HuggingFace、本地 whisper.cpp,以及你自己的任何 OpenAI-compatible API)。

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add @goodandready/dsh-voice@0.8.33
README兼容性版本

兼容性与来源证明

Voice 以 @goodandready/dsh-voice 发布,当前版本为 0.8.33。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
web
发布来源
npm
Registry 更新时间
2026/9/20

版本

0.8.33stable
2026/9/19
0.8.32stable
2026/9/19
0.8.31stable
2026/9/18
查看其余 42 个版本收起版本
0.8.30stable
2026/9/17
0.8.29stable
2026/9/17
0.8.28stable
2026/9/15
0.8.27stable
2026/9/15
0.8.26stable
2026/9/13
0.8.25stable
2026/9/12
0.8.24stable
2026/9/12
0.8.23stable
2026/9/11
0.8.22stable
2026/9/11
0.8.21stable
2026/9/11
0.8.20stable
2026/9/10
0.8.19stable
2026/9/9
0.8.18stable
2026/9/7
0.8.17stable
2026/9/6
0.8.16stable
2026/9/5
0.8.15stable
2026/9/4
0.8.14stable
2026/9/2
0.8.13stable
2026/9/2

相关插件

正在加载相关插件…

最新版
0.8.33
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
276.2 kB
文件数
19
Surface
web
许可证
MIT
发布源
npm
GitHub
★ 6
周下载
1,364
安全扫描
✓ v0.8.33 扫描通过
最近提交
2026/9/19
查看源码 ↗项目主页 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题
0.8.12stable
2026/9/2
0.8.11stable
2026/9/2
0.8.10stable
2026/9/2
0.8.9stable
2026/8/30
0.8.8stable
2026/8/28
0.8.7stable
2026/8/27
0.8.6stable
2026/8/26
0.8.5stable
2026/8/26
0.8.4stable
2026/8/26
0.8.3stable
2026/8/26
0.8.1stable
2026/8/25
0.8.2stable
2026/8/25
0.8.0stable
2026/8/25
0.7.3stable
2026/8/23
0.7.2stable
2026/8/23
0.7.1stable
2026/8/23
0.7.0stable
2026/8/23
0.6.1stable
2026/8/21
0.6.0stable
2026/8/20
0.5.0stable
2026/8/20
0.4.2stable
2026/8/20
0.4.1stable
2026/8/20
0.4.0stable
2026/8/20
0.3.0stable
2026/8/19

相关插件

继续浏览 vision-media 分类下经过校验的插件。

Tool Describe Image@linxin666/dsh-tool-describe-image面向模型的 describe_image 工具,用于 dsh Web GUI:通过在兼容 OpenAI 的端点调用视觉语言模型,为文本模型提供图像理解能力,以描述一张图像(本地路径、http(s) URL 或附件引用)。可热插拔 —Deepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio,以及 27 个可编辑视频模板,以原生 DeepSeek Harness 对话视图呈现。Imagegen@dickpy/dsh-imagegendsh Web GUI 的 AI 图像生成插件:通过可配置的提供商渠道实现文生图和图生图(gpt-image-2 / grok-imagine-image / nanobanana series / seedream-5.0-pro / dall-e-3,支持原生 xAI Grok Imagine、Google Nano Banana aVoice Modedsh-voice-mode适用于 DeepSeek Harness 的全双工语音插件:本地 zipformer2 流式 ASR(无需 API 密钥)→ 可编辑草稿;Edge TTS 或本地 VITS / Kokoro 朗读并显示实时字幕;真正的抢话;强化的 HTTP 接口 + 模型 SHA256 固定;兼容所有 ds

README

📦 @goodandready/dsh-voice

Zero-Latency Streaming Dictation & Multi-Provider Voice Input for DeepSeek Harness

🇬🇧 English • 🇷🇺 Русский • 🇨🇳 中文说明

⭐ If you like this plugin, please star it on GitHub — it shows me that the plugin is useful to you and motivates me to keep developing it.

🐛 If you find a bug or would like to request a feature, open a GitHub issue in any language — I will review your proposal and implement useful suggestions in a future plugin version.

⚡ Overview

dsh-voice brings voice superpowers to the DeepSeek Harness Web UI. Whether you need hands-free real-time streaming dictation segmented on natural breath pauses or crisp voice notes with keyboard/mouse Push-to-Talk gestures, dsh-voice ensures your audio is never lost thanks to automatic multi-provider fallback chains.

graph LR
    subgraph Client [Browser Web UI]
        Mic[🎙️ Dictation Mic] -->|VAD Cut on Pause| Stream[Audio Chunks]
        Wave[🌊 Voice Message] -->|Hold / Release| PTT[Push-to-Talk]
    end

    subgraph Host [DSH Host Backend]
        Stream --> FFMPEG[ffmpeg 16kHz Transcoder]
        PTT --> FFMPEG
        FFMPEG --> Chain{Fallback Chain}
        
        Chain -->|1st Priority| P1[Deepgram / Nova-2]
        Chain -.->|On Rate Limit / 429| P2[Groq / Whisper Turbo]
        Chain -.->|On Failure| P3[Local whisper.cpp / Offline]
    end

    subgraph Output [Target]
        P1 --> Composer[💬 Web Composer / Chat]
        P2 --> Composer
        P3 --> Composer
    end

    style Client fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
    style Host fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
    style Output fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4

✨ Key Features

  • 🎙️ Streaming Dictation with VAD: Speech is automatically sliced at natural pauses (vadSilenceMs, default 700ms) and typed into the composer in real time.
  • 🌊 Voice Notes with Cancel Window: Record your thought and have it automatically dispatched to the agent after a safety countdown (autoSendMs, default 4000ms).
  • 🎮 Tactile Push-to-Talk:
    • Mouse: Hold the wave button — releasing sends the message; dragging pointer away discards.
    • Keyboard: Hold Ctrl (or custom hotkey) for hands-free speaking; press Esc to cancel.
  • ⚡ Zero-Latency In-Browser Captions (browser): Chrome Web Speech API recognition runs 100% locally with live floating captions as you speak.
  • 🛡️ Ironclad Multi-Provider Fallbacks: If your primary cloud provider runs out of credits or hits a 429 rate limit, requests seamlessly fail over down the chain.
  • 🧠 Context Glossary Injection: Automatically extracts code variables and identifiers from your composer draft to steer STT model accuracy on technical jargon.
  • 🎵 Embedded Audio Player: Preview, scrubber, and playback of your recorded voice message directly in chat and the composer dock.
  • 🔇 Hardware Noise Suppression Toggle: Configurable in settings to toggle browser-level noise suppression, echo cancellation, and auto gain control.
  • 📊 Provider Latency & Health Dashboard: Live visual telemetry of provider latency (ms), success rates, and errors directly within the settings UI.
  • 🔒 Zero API Key Leakage: Keys are resolved on the host via ctx.credentials (credentialRef) and never transmitted to browser clients.
  • 🖥️ Offline Local Whisper Server: Automatically boots and manages whisper.cpp (whisper-server) with on-the-fly ffmpeg transcode.
  • ⚡ SenseVoice-ONNX / Sherpa-ONNX (0.8.11, updated 0.8.32): Ultra-fast (~50–100ms) non-autoregressive local STT engine with 1-click automatic model installation to ~/.dsh/models/sensevoice.
  • 🎙️ Gated Turn-Taking & Echo Prevention (0.8.32): Automatically mutes the mic while the assistant is speaking (dsh:tts:start/stop) and supports speech-triggered Barge-In.

📝 What's New in 0.8.32

10-point feature epic & SenseVoice 1-Click Installer (Issue #129):

AreaFeatureDescription
Local ASR1-Click SenseVoice-SmallAutomated download and extraction of model.int8.onnx and tokens.txt directly to ~/.dsh/models/sensevoice via host loopback endpoint.
Turn-TakingGated ModeMicrophone input is muted when assistant audio starts (dsh:tts:start) and unmuted on dsh:tts:stop, eliminating acoustic feedback.
InterruptionBarge-InSpeaking immediately signals assistant TTS to pause and abort current speech playback.
FeedbackGhost Interim PreviewSemi-transparent live text preview in composer pill during recognition before sentence finalizing.
CountdownSVG Silence RingCircular SVG progress ring visually ticking down pending auto-send delay.
ControlHands-free ActionsSpoken control words ("send", "cancel", "clear", "new line") trigger actions instead of becoming message text.
PromptsStructured PromptsSpoken replies automatically match and submit options in active harness interactive prompts.
VocabularyIT Jargon NormalizerAuto-corrects spoken developer slang to canonical spelling (GitHub, Docker, Kubernetes, etc.) with custom dictionary in Settings.
ConfigurationIndividual TogglesDedicated switches for every enhancement in Settings card with native --dsw-alias-* token styling.

📝 What's New in 0.8.19

Quality batch after the v0.8.18 review (Gitea #79–#87, PR #88):

AreaChange
Settings placeholdersModel/path hints resolve through locale at render time — no frozen i18n keys (#79)
Model field hintswhisperModel / sensevoiceModel show path-to-model copy, not binary-autostart hints (#82)
VisualizerLiquid wave / dynamic orb read DSH theme tokens instead of hardcoded hex (#80)
Style isolationInjected CSS uses data-dsh-plugin="dsh-voice" so neighbour HMR cleanup cannot strip styles (#81)
Source languageCode comments, errors, and tests are English; EN edit-command phrases added. Spoken RU STT patterns remain for recognition (#86)
LocalesChanged: the full inline ru dictionary was removed. English is the only bundled locale. Install the translation plugin for Russian UI (#85)
Client sourceBrowser client is built from ordered lib/client-src/*.js fragments via npm run build:client (#87)
Process docsAdded docs/design/DESIGN.md, index.md, project AGENTS.md, docs/testing/unit.md (#83)

[!IMPORTANT] v0.8.19 locale behavior: without a translation plugin the Web UI stays in English. Session/edit spoken command phrases still match Russian speech for STT where documented; interface labels do not ship a second dictionary.


🎮 Four Ways to Speak

ModeGesture / TriggerBehavior
DictationClick 🎙️ MicSpeech is sliced on pauses (vadSilenceMs) and typed live into composer
Voice MessageClick 🌊 WaveRecords until stopped, then sends after cancel window (autoSendMs)
Mouse PTTHold 🌊 WaveRecords while held; release sends message, drag off button to discard
Keyboard PTTHold CtrlHands-free recording; release sends message, press Esc to discard

[!TIP] You can customize the keyboard modifier in settings (hotkey: Control, Alt, Shift, or any KeyboardEvent.code).


🛠️ Supported Providers Matrix

Provider KeyService BackendDefault ModelCredential RefFeatures & Notes
browserWeb Speech APINative BrowserNoneZero latency, floating live captions in Chrome
deepgramDeepgram APInova-2DEEPGRAM_API_KEYUltra-fast cloud transcription
groqGroq Whisperwhisper-large-v3-turboGROQ_API_KEYNear-instant inference speed
hfHuggingFace Inferenceopenai/whisper-large-v3HF_TOKENHigh-accuracy open Whisper
local-whisperLocal whisper.cppServer definedNone100% private, offline, no internet needed
sensevoiceSenseVoice-ONNX / Sherpa-ONNXSenseVoiceSmallNoneUltra-fast (~50ms) local non-autoregressive STT

🚀 Ready-Made Presets (Plug & Play)

Just specify the name in your fallback chain and add the corresponding API key:

  • openai (whisper-1) → OPENAI_API_KEY
  • siliconflow (SenseVoiceSmall) → SILICONFLOW_API_KEY
  • mistral (voxtral-mini-latest) → MISTRAL_API_KEY
  • openrouter (google/gemini-2.5-flash) → OPENROUTER_API_KEY
  • deepinfra (whisper-large-v3-turbo) → DEEPINFRA_API_KEY
  • fireworks (whisper-v3-turbo) → FIREWORKS_API_KEY

📦 Quick Installation

dsh plugin --profile web add @goodandready/dsh-voice

[!IMPORTANT] Restart DSH Web UI after installation (systemctl --user restart dsh-web) and refresh your browser tab.


⚙️ Configuration

Open Settings → Plugins → Plugin settings → Voice in the Web UI:

- id: dsh-voice
  config:
    dictation:
      language: ru
      vadSilenceMs: 700
      chain:
        - provider: deepgram
        - provider: groq
        - provider: local-whisper
    message:
      language: ru
      autoSendMs: 4000
      chain:
        - provider: openai
        - provider: local-whisper
    hotkey: Control
    autoStart: true
    whisperModel: /models/ggml-medium-q8_0.bin

🤖 Agent Tool & HTTP API

Agent Tool (transcribe_audio)

Registers transcribe_audio(file_path, language?) in ctx.tools, allowing agents to analyze audio files, interview recordings, and voice notes directly from disk.

Internal HTTP Endpoints

  • POST /dsh-voice/transcribe — { dataBase64, mimeType, mode } → { ok, text, provider, tookMs }
  • POST /dsh-voice/polish — { text } → { ok, text }
  • GET /dsh-voice/status — Returns daemon status, active chains, SenseVoice and realtime config.
  • GET /dsh-voice/realtime — WebSocket upgrade for low-latency audio streaming (OpenAI Realtime API / Sherpa-ONNX). Accepts binary audio chunks, returns JSON text deltas.

📄 License

MIT © GooDAnDReaDY

  • 👻 Live Ghost Interim Preview (0.8.32): Visual real-time preview of spoken words before final chunk transcription.
  • ⏱️ SVG Silence Ring Timer (0.8.32): Circular animated countdown indicator in the composer dock during the pending message delay.
  • 🗣️ Hands-free Voice Actions (0.8.32): Spoken commands ("send", "cancel", "clear", "new line") trigger UI actions directly.
  • 🧩 Structured Prompt Voice Selection (0.8.32): Spoken options automatically select buttons in interactive assistant choice prompts.
  • 📖 Developer Lexicon & IT Jargon Correction (0.8.32): Phonetic normalization for developer slang (GitHub, Docker, Kubernetes, pnpm) with customizable settings dictionary.
  • 🌐 Realtime Audio Streaming (0.8.11): Low-latency WebSocket bridge (/dsh-voice/realtime) for OpenAI Realtime API or local Sherpa-ONNX streaming. API keys stay securely on the host.
  • 🌊 Liquid Wave & Dynamic Orb Visualizer (0.8.12): Smooth animated audio visualization in the recording pill with real-time mic volume reactivity. Switch between organic multi-layer liquid waves, pulsating radiant orb, classic bars, or off.