CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
377 plugins@lamplitisles/kepos-speechKepos Speech
Kepos Speech plugin for Alibaba and ByteDance Chinese speech synthesis plus short-audio recognition in DeepSeek Harness Web
dsh-voice-scribeVoice Scribe
Voice input plugin for DSH: tap or hold Alt to talk, then release or tap again to convert speech to text. Supports a hotword replacement list (hot.txt), custom polishing prompts, and a recording level indicator. Uses local offline recognition by default (SenseVoice; zero configuration, zero key, and
dsh-image-viewerImage Viewer
Zoom, pan, download, gallery, and region-note image viewer for DeepSeek Harness
dsh-live2d-avatarLive2d Avatar
A switchable Live2D avatar stage and desktop companion for DeepSeek Harness.
dsh-grok-adaptationGrok Adaptation
Normalize undersized Grok Responses images with Sharp while leaving supported large images unchanged.
dsh-vision-pro-bridgeVision Pro Bridge
Give text-only DeepSeek-V4-Pro real vision with zero new dependencies and DeepSeek-only routing: images are described by deepseek-v4-flash-vision-exp (your existing DEEPSEEK_API_KEY), then the text is handed to V4-Pro.
dsh-vision-toggleVision Toggle
DSH model vision toggle: adds a “Model Vision” row to the settings page for switching the input vision modality of manually declared models under llm-pi-ai custom routes via the official settings.mutate channel, with changes taking effect immediately.
dsh-asr-voiceAsr Voice
Speak and it becomes text · Speak-to-prompt for DeepSeek Harness: cloud ASR speech recognition + prompt optimization + fill into draft/auto-send, cross-platform macOS / Windows.
@dshp-inx/vision-bridgeVision Bridge
DeepSeek Harness (DSH) vision bridge: let text-only models delegate image understanding to multimodal vision models via vision_describe tool, with persistent settings UI and primary/fallback auto-retry. 纯文本模型通过 vision_describe 工具把图片转交给多模态视觉模型识别的桥接插件。
dsh-pdf-to-wordPdf To Word
DeepSeek Harness plugin: PDF→Word (.docx) conversion with layout fidelity (fonts/tables/images/borders), OCR scan mode, and optional multimodal LLM verification. Registers the pdf_to_word model tool.
dsh-vision-patchVision Patch
Per-model and per-route image-input (vision) checkboxes on the Models page's custom-provider cards, writing through the llm-pi-ai settings namespace.
dsh-plugin-visionPlugin Vision
Give the DSH agent a pair of eyes: call online VLMs (multi-provider, OpenAI-compatible) to analyze local images, URLs, and attachments uploaded in conversations; the selected models’ specialized capabilities are injected into the system prompt in real time
dsh-imgdrawImgdraw
Text-to-image for DeepSeek Harness: a `draw_image` model tool, an input-bar 生图 button with a prompt popup (async generation, 4-grid results, download / keep / delete), an /imgdraw image route, and persisted history. Backends: DashScope wan2.7-image (free default) and SiliconFlow Qwen-Image.
dsh-vision-autoswitchVision Autoswitch
DeepSeek Harness plugin: auto-route image-bearing requests to deepseek-v4-flash-vision-exp, then fall back to the original model.
dsh-vision-assistVision Assist
DeepSeek Harness vision assistant plugin: equips models without vision capabilities with a switchable multimodal recognition model. Images in the input field are automatically saved to disk and rewritten as text prompts; the main model can view images by calling the vision_recognize tool. The recogn
dsh-vision-bridgeVision Bridgealt · alaxrpg
DSH Vision Bridge Plugin: Reuse DSH Provider or connect directly via OpenAI compatibility, with visual configuration and image pasting support
dsh-plugin-show-imagePlugin Show Image
Render local image files inline in the DSH conversation via a global show_image tool.
media-previewMedia Preview
DSH 插件:在聊天记录中自动将本地音视频/图片路径渲染为可播放的预览组件。When an assistant message or tool result contains a local media path, the path is replaced inline with a playable <audio>/<video>/<img> element backed by same-origin /api/media-preview/* route.
dsh-plugin-voicePlugin Voice
DeepSeek Harness plugin: voice + notification outputs—agents proactively contact users through cloud TTS (Volcano seed-tts / Xiaomi MiMo V2.5, with automatic fallback to SAPI on failure), desktop notifications, and alert sounds. Combines the native DSH integration of dsh-plugin-notify with the cloud
@maiziman/dsh-model-capabilitiesModel Capabilities
Automatic reasoning and image capability detection for custom DeepSeek Harness models
dsh-llm-multimodalLlm Multimodal
DSH plugin: Provides image/video generation tools in DSH, based on OpenAI-compatible APIs. Models are automatically discovered from existing llm-pi-ai settings.
@xiaokaizhou/dsh-media-previewMedia Previewalt · @xiaokaizhou
DSH 插件:在聊天记录中自动将本地音视频/图片路径渲染为可播放的预览组件。When an assistant message or tool result contains a local media path, the path is replaced inline with a playable <audio>/<video>/<img> element backed by same-origin /api/media-preview/* route.
dsh-image-amnesiaImage Amnesia
Keep native vision, drop historical images before they hit relay providers. Global DeepSeek Harness bundle for every agent.
dsh-voice-controlVoice Control
Voice control for DSH web: speech-to-text into the composer (with auto-send) and spoken playback of assistant replies via the Web Speech API
dsh-photosPhotos
Photo upload for the DeepSeek Harness composer — native picker feeding the stock paste pipeline
@jetecho/dsh-csv-and-image-previewCsv And Image Preview
Preview images / SVG and CSV tables in the DeepSeek Harness chat, rendered as real browser <img> / <table> elements. Preview-first workflow: show the user the asset, wait for approval, then apply the real change.
@goodandready/dsh-im-hub-mediaIm Hub Media
Multi-platform IM gateway for DeepSeek Harness (fork of dsh-im-hub with media): Telegram voice/photo/document/video + reply handling, STT (Deepgram primary, HF Whisper fallback), outbound MEDIA: markers. Feishu/WeCom kept as-is.
dsh-svw-waveformSvw Waveform
Native SVW waveform rendering plugin for DeepSeek Harness
dsh-multimodal-runtimeMultimodal Runtime
Mogu Multimodal Runtime - DeepSeek Harness 统一多模态能力运行时 (Comfy Local Provider V1)
dsh-speech-inputSpeech Input
A microphone button for DeepSeek Harness that writes browser speech recognition into the composer draft.