CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
376 pluginsdsh-image-genImage Gen
Bring ChatGPT-like image generation to DeepSeek Harness — Gemini, OpenAI, Seedream, DashScope, local ComfyUI & more.
@eric.wen/dsh-sightSight
DeepSeek Harness plugin: direct multimodal image transfer declarations + per-session image clearing, reasoning-effort auto-fill, and progressive Figma MCP bridging (design-to-code + AI-driven design).
dsh-codex-toolsCodex Tools
Codex-backed web search, image generation, and image understanding tools for the DeepSeek Harness.
dsh-open-fileOpen File
Workspace-bound arbitrary file upload, reading, OCR, and rendering for DeepSeek Harness.
dsh-mindseyeMindseye
MindsEye: model-driven vision tools, structured evidence, and exact cache for DeepSeek Harness
@maxwell-feng/dsh-tesseract-ocrTesseract Ocr
dsh plugin: recognize attached images locally with Tesseract OCR and send only the recognized text to the model. Text models never receive image bytes; vision passthrough is opt-in.
@omdp/dsh-vision-bridgeVision Bridge
DSH Vision Bridge plugin: Automatically distinguishes multimodal and text models. Multimodal models view images directly; text models use a configurable multimodal endpoint (baseUrl + apiKey + model) to view them on their behalf. Supports pasted images, the read_image tool, and converting images int
dsh-vision-webVision Web
Eyes for text-only DeepSeek on DeepSeek Harness: image understanding via Doubao Web by default (zero cost, no API key), Antigravity IDE quota (flash/pro), Gemini API, or Cockpit proxy — model-invokable vision tool, wrapper adapters, evidence memory, and a
dsh-screenshot-feedback-hook-mcpScreenshot Feedback Hook Mcp
Screenshot feedback for DeepSeek Harness — let your coding agent SEE what it builds
dsh-video-understandVideo Understand
Low-cost video understanding tool: Bilibili links/BV/local videos → information layer (ASR + scenes + object trajectories + YOLO) → summaries + Q&A. Question-driven dynamic routing across levels (L0/L1/L2), semantic-layer reuse, and budget caps. Self-contained engine with no external dependencies.
dsh-dictateDictate
Context-aware voice input for DeepSeek Harness with Web Speech, local SenseVoice transcription, model polish, editable Composer drafts, and user-controlled sending
dsh-auto-visionAuto Vision
DeepSeek Harness Vision Bridge: automatically discovers your configured multimodal models, equips text-only primary models with a vision tool, and returns recognition results as plain text. Zero configuration, one-command installation.
@xlight-oss/visionary-dshVisionary Dsh
DeepSeek Visionary native plugin for DeepSeek Harness: deepseek_vision / status / login / logout native tools plus the text-model image bridge, all backed by the visionary-server CLI (DeepSeek web vision model, no API key).
dsh-plugin-deepeyePlugin Deepeye
DeepSeek Harness native plugin: vision capabilities for text-only LLMs (describe, OCR, VQA, layout analysis, clipboard)
dsh-better-inputBetter Input
Better input experience for DeepSeek Harness: voice input, AI polishing, PDF and image input (voice input first)
@goodandready/dsh-fal-image-genFal Image Gen
Image generation for DeepSeek Harness: a generate_image tool with pluggable providers — the FAL queue API or any OpenAI-compatible images API. The picture is shown inline in the conversation; the model receives either a link (works with any chat model) or
@alain-prot0s5/dsh-screenshotScreenshot
Screenshot-to-input for DeepSeek Harness: composer camera button + global hotkey (Alt+A) + listener bound to the app lifecycle, configurable in settings. 截图自动粘贴到 DSH 输入框:相机按钮 + 全局快捷键 + 生命周期绑定 + 设置页配置。
computer-userComputer User
Codex-style computer use for DeepSeek Harness (DSH): read the screen and drive the mouse & keyboard. Pairs with picturereader (image_scan/image_ocr) to close the look-act-verify loop. Windows.
picturereaderPicturereader
Unified image understanding plugin for DeepSeek Harness (DSH). Visual twin adapter for native thumbnails + auto-analysis on any text-only model (incl. pi-ai providers); privacy/smart/strict routing; local tools (scan/OCR×4 engines (windows/macos/paddle/ra
dsh-ocr-localOcr Local
Local OCR for DeepSeek Harness: paste/attach an image, get its text via PP-OCRv5 + ONNX Runtime, fully offline. TUI (cc-tui) and Web. / DeepSeek Harness 本地 OCR 插件:图片转文字,PP-OCRv5 + ONNX Runtime,完全离线,支持 TUI 与 Web。
dsh-dseyesDseyes
Give DeepSeek Harness free eyes — natively. Paste or drop an image into the Web GUI composer: it appears as a thumbnail attachment (like a normal AI chat), and before the text-only DeepSeek model is called, the host automatically reads the image with the
dsh-image-pluginsImage Plugins
Multimodal plugin for DeepSeek Harness: understand images and generate images through configurable OpenAI-compatible or DashScope endpoints.
dsh-vision-mixVision Mix
A DeepSeek Harness Mix plugin for vision routing plus GPT Image generation and editing.
dsh-vision-free-eyesVision Free Eyes
DeepSeek Harness plugin: a model-facing `vision` tool that describes and OCRs image files by calling the free Zhipu GLM vision API directly (no external CLI required).
dsh-vision-recognizerVision Recognizer
Adaptive image routing for DeepSeek Harness: pass images directly to native multimodal models and transcribe them only for text-only models.
dsh-youreyesYoureyes
Eyes for text-only DeepSeek on DeepSeek Harness: image understanding via Antigravity IDE quota (default, flash/pro), OpenAI-compatible VLM endpoints, Gemini API, or local Ollama — model-invokable vision tool, wrapper adapters for deepseek/opencode-go, evi
dsh-vision-proxyVision Proxy
DeepSeek brain + automatic image transcription for DeepSeek Harness: a deepseek-vision provider route that transcribes attached images to text via any OpenAI-compatible VLM before delegating to the text-only DeepSeek adapter. Paid fast path with a key (Da
dsh-vision-proxy-routeVision Proxy Route
DeepSeek Harness plugin: a configurable provider route that transcribes pasted images via free Zhipu GLM vision models before delegating to a text-only adapter.
dsh-design-qaDesign Qa
Let any text-only model in DeepSeek Harness read images. Image recognition is an on-demand tool—images do not enter the main model context, and you pay nothing if they are not viewed; includes an evaluation set with 23 defects for self-testing when switching models.
dsh-voice-webspeechVoice Webspeech
DSH Web voice input plugin: no server, no keys, and no model downloads; directly uses the browser's built-in Web Speech API (Edge = Microsoft Azure Speech, Chrome = Google Speech).