CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
376 pluginsdsh-ui-specUi Spec
DeepSeek Harness plugin that turns UI screenshots into implementation-grade web specs using OCR, deterministic geometry, scene graphs, assets, and render comparison.
dsh-voice-callVoice Call
dsh-voice-call — give the agent a voice it owns: the agent decides when to speak (offer_call), the human holds the answer key (接听/拒接/稍后). Local-first TTS via CrispASR + Qwen3-TTS CustomVoice, plain audio files under ~/.dsh/voice/. Fork of dsh-voice with a
dsh-model-capabilityModel Capability
Explicit model input-modality selection for the DeepSeek Harness Web UI
dsh-sound-labSound Lab
DSH 声音工坊(Sound Lab):Dialogue event sound effects, AI character voice generation, and sound library upload management, all through visual point-and-click controls; includes the 明日方舟安洁莉娜 desktop pet(60fps sprite animation, edge docking, and floating window configuration). Hot-pluggable dsh plugin.
dsh-plugin-deepseek-visionPlugin Deepseek Vision
DeepSeek Harness native vision Bundle: paste or drag in images and call OpenAI-compatible vision models through the hosted deepseek-vision-mcp.
dsh-plugin-image-toolsPlugin Image Tools
DSH image plugin: ask_user_choice image/image-text mixed options (Web GUI renders image selection cards, with zoom support) + show_images embeds images in replies (mixed image and text) + click to enlarge all images in the chat area, with scroll wheel/button zoom and drag panning in the enlarged vie
dsh-vision-linkVision Link
Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.
dsh-voice-kitVoice Kit
Voice input (Web Speech API) and read-aloud (speechSynthesis) for the DeepSeek Harness web GUI — DSH 语音输入 + 回复朗读套件
dsh-deepseek-visionDeepseek Vision
Out-of-tree dsh provider plugin: a DeepSeek gateway route that claims image input and transparently describes pasted images through a configured vision-language model (e.g. Qwen-VL) before the text-only DeepSeek wire sees them.
dsh-sightSight
Plug-in vision for text-only DeepSeek Harness (dsh) models: a `vision` tool with built-in cheap/free VLM presets, multi-image batch analysis, paste-to-hint image admission, and a web settings page with hot-reload.
dsh-open-eyesOpen Eyes
Open Eyes for DeepSeek Harness: delegate images to a configurable multimodal model through OpenAI Responses, Chat Completions, or Anthropic Messages.
@svenyu/dsvuDsvu
dsvu (DeepSeek-Powered Video Understanding), a low-cost video understanding tool: Bilibili links/BV/local videos → information layer (ASR + scenes + object trajectories + YOLO) → summaries + Q&A. Uses same-source instant viewing and answering with deepseek-v4-flash-vision-exp + dense grid frame samp
@dsh-extension/dsh-vision-bridgeVision Bridge
On-demand vision for text-only DSH sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
@maxwell-feng/dsh-windows-ocrWindows Ocr
dsh plugin: recognize attached images with the built-in Windows OCR engine (Windows.Media.Ocr) and send only the recognized text to the model. Text models never receive image bytes; vision passthrough is opt-in.
dsh-gemini-multimodalGemini Multimodal
DeepSeek Harness plugin: multimodal tools (image/audio/video/document understanding, transcription, image generation) via Gemini API or the local Antigravity CLI.
dsh-mac-visionMac Vision
Native macOS OCR and Vision tools for DeepSeek Harness
dsh-vision-fallbackVision Fallback
DSH Silent Vision Enhancement: Select the main model as usual; images are automatically sent to a fixed vision model and returned to the main model as hidden context.
dsh-tool-visual-primitivesTool Visual Primitives
DSH visual primitives tool: Based on DeepSeek’s paper “Thinking with Visual Primitives,” routes images to an external vision model and returns text analysis with visual primitives. In a text-only loop, conversational models can “see” images without native vision capabilities.
dsh-local-visionLocal Vision
DeepSeek Harness plugin: a model-callable local_vision tool that describes images with a local vision model through Ollama. No cloud, the image never leaves the machine. Two tiers: fast (text/colors/summary) and detailed (full description).
@zoytown/dsh-avatarAvatar
DeepSeek Harness (dsh) plugin for wallpaper theming — upload images from the Settings page, pick one, and the whole dsh web UI renders over it with an adjustable readability mask and blur.
dsh-tool-generate-imageTool Generate Image
Model-facing generate_image tool for DeepSeek Harness (dsh): generates brand-new images with the Google Gemini image-generation service via the Antigravity CLI, saves them, and returns the paths. Hot-pluggable — install with `dsh plugin --profile web add
@woyeshishen/dsh-vision-pluginVision Plugin
Provides external vision model capabilities for DeepSeek Harness: the text-only primary model uses the describe_image tool to call an external vision model to analyze images and receive a text-only description (multimodal completion). A static Cordis plugin that loads automatically when DSH starts.
dsh-vision-pluginVision Pluginalt · Xin-Zhang-IceMan
dsh-vision-plugin: give DeepSeek Harness text-only models a pair of eyes — pasted images are transcribed by a vision model before they reach a text-only main model, plus the vision_analyze tool and a bilingual settings page.
dsh-subagent-visionSubagent Vision
dsh bundle: subagent_vision — delegate image reading to a vision-capable model from a text-only session, plus paste-to-path so pasted images reach the subagent as file paths.
dsh-vision-guardVision Guard
Transparent image guard + vision analysis for DeepSeek Harness: text-only models read pasted images without the 400 session deadlock.
mimo-visionMimo Vision
DeepSeek Harness (DSH) native plugin: the describe_image tool, a vision bridge (image -> mimo-v2.5 -> text description) over the ctx.fs / ctx.credentials seams
dsh-plugin-qwen-imagePlugin Qwen Image
DeepSeek Harness out-of-tree plugin: give a text-only coding model eyes by routing images to a Qwen-VL (DashScope) route through ctx.llm and returning text.
@motong/dsh-voiceVoice
A community plugin that adds voice capabilities to DeepSeek Harness (DSH / DeepSeek Hermes): voice input in the input field (with configurable shortcut) and spoken responses (Microsoft Edge neural voices, with voice switching and preview), with no API key required.
dsh-speech-pluginSpeech Plugin
Speech plugin for DeepSeek Harness: per-message speak button, composer voice input, and auto-announce toggle, over cloud TTS/ASR with Web Speech API fallback
dsh-multimodal-bridgeMultimodal Bridge
DeepSeek Harness plugin bundle: qwen_vision (Qwen-VL image understanding) and qwen_generate (Qwen-Image text-to-image and image editing) tools for text-only models