DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Voice — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins
V

@nn12138/dsh-voice

Voice

DeepSeek Harness 的语音输入插件:麦克风 → 本地/浏览器语音识别 → 作为普通聊天消息提交文本(仅输入,不依赖预设)

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add @nn12138/dsh-voice@0.2.6
README兼容性版本

兼容性与来源证明

Voice 以 @nn12138/dsh-voice 发布,当前版本为 0.2.6。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
web
发布来源
npm
Registry 更新时间
2026/8/21

版本

0.2.6stable
2026/8/21
0.2.5stable
2026/8/21

相关插件

正在加载相关插件…

最新版
0.2.6
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
161.3 kB
文件数
44
Surface
web
许可证
MIT
发布源
npm
GitHub
★ 7
周下载
0
最近提交
2026/8/21
查看源码 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题

相关插件

继续浏览 vision-media 分类下经过校验的插件。

Tool Describe Image@linxin666/dsh-tool-describe-image面向模型的 describe_image 工具,用于 dsh Web GUI:通过在兼容 OpenAI 的端点调用视觉语言模型,为文本模型提供图像理解能力,以描述一张图像(本地路径、http(s) URL 或附件引用)。可热插拔 —Modlens@liustack/modlens面向仅支持文本的 LLM 的插件视觉能力,由免费的 Antigravity CLI 提供支持Vision Toolkit@anionex/dsh-vision-toolkit面向 Harness 原生集成的 DeepSeek 与 agent-vision-toolkit:图像问答、OCR、定位、UI 还原、像素差异、Artifacts 和 Web UI。Deepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio,以及 27 个可编辑视频模板,以原生 DeepSeek Harness 对话视图呈现。

README

dsh-voice 🎤

English | 中文

A voice input plugin for DeepSeek Harness: click 🎤 in the web UI (or press a hotkey), speak, and the recognized text is submitted as a normal chat message. Input only — it never touches the agent preset/persona, so it behaves like "another input method" in every mode.

Features

  • 🎤 Voice input: microphone button (platform design-system UI) + configurable global hotkey (Ctrl+Space by default)
  • ⚡ Live recognition: streaming partial transcripts are echoed while you speak; VAD finalization commits on stop (0.6s tail padding keeps sentence endings)
  • 🧠 Adaptive dual engine: host-native ASR (sherpa-onnx-node zipformer2 + silero VAD, offline/private) with automatic fallback to browser Web Speech (zero extra dependencies)
  • 🔌 Preset-agnostic: does not touch the persona/system prompt; works with code/standard/minimal/custom presets
  • 📦 Optional models: zero-config out of the box; run dsh-voice-models when you want native offline recognition

Installation

# Plugin
dsh plugin --profile web add @nn12138/dsh-voice

# Optional: offline native recognition (the plugin does not auto-install this runtime)
dsh plugin --profile web add sherpa-onnx-node
dsh-voice-models            # one-shot model download (~100MB) → ./dsh-voice-models

# Optional: configuration (edit ~/.dsh/profiles/web/cordis.patch.yml)
- id: voice
  config:
    modelDir: './dsh-voice-models'   # native ASR model directory
    hotkey: 'ctrl+space'             # global hotkey
    vadThreshold: 0.3                # lower = less clipping at sentence boundaries
    tailPadSeconds: 0.6              # tail-padding duration
    engine: auto                     # auto (default) | native | browser

The row-level config is received by the host half. engine and hotkey are synced to the browser half over the /voice.config loopback RPC, so there is no separate client config to write. auto probes host native capability: with a model it uses native; without one it falls back to Web Speech, so zero-config users keep working. Restart dsh web after changing the config.

dsh web   # 🎤 button appears on the left of the composer, or press Ctrl+Space

See USAGE.md and INSTALL.md (Chinese) for details.

How it works

Browser captures mic audio (auto-resampled to 16 kHz)
  → PCM base64 chunks (256 ms) → /voice RPC channel (loopback)
  → host: silero VAD + zipformer2 streaming decode
  → partials returned per chunk (live echo) / finals committed
    (VAD segmentation + 0.6s tail padding)
  → conversation service submits the text (same path as typing)

Engine selection: the host resolves the effective engine (config + model-load result) and the client consumes it via /voice.ping — native unavailable falls back to browser Web Speech. /voice.config carries the row-level engine/hotkey from host to client.

Development

pnpm install --ignore-workspace        # standalone deps (no DSH monorepo needed); no install-time scripts
pnpm --ignore-workspace test           # unit tests (including real-model smoke tests)
pnpm --ignore-workspace typecheck      # type check
pnpm --ignore-workspace build          # build (tsc host half + tsdown client half)
pnpm --ignore-workspace verify:package # pack-level checks: scripts-free install, complete runtime files

The build only ever runs on the publisher's side (prepack — at npm pack/npm publish time) and in CI before release; consumers installing @nn12138/dsh-voice from the registry never execute lifecycle scripts, so --ignore-scripts installs are complete and usable (see issue #2).

Real-model smoke tests look for the local voxelf assets and skip when absent; override with: DSH_VOICE_MODEL_DIR (model directory) / DSH_VOICE_TEST_WAV (test wav) / DSH_VOICE_DOWNLOADED_MODELS (downloaded model directory).

Layout: src/index.ts (host half) / src/client/ (browser half) / src/core/ (recognition core) / tools/ (wire-protocol smoke tools + model downloader).

License

MIT