CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
377 plugins@hackerfish/dsh-voice-inputVoice Input
Voice input: microphone button in the dialog, browser-built-in Web Speech API (zh-CN) recognition, results inserted into the input field draft
@zenk/visionVision
DSH local vision capabilities: macOS Vision OCR + ollama qwen3-vl semantic description + uploaded image bridge (converts image blocks to text, available on text-only channels)
dsh-deepseek-model-routerDeepseek Model Router
Auto-switch DeepSeek models for DeepSeek Harness: routes every model request to the vision model when images are present, to the pro model for complex tasks, and to the fast model otherwise; includes a switch_model tool for manual per-session overrides.
dsh-voice-pluginVoice Plugin
Voice input plugin for DeepSeek Harness
dsh-funasr-voiceFunasr Voice
DSH Web local offline voice input plugin: the browser captures microphone audio → host launches local FunASR (SenseVoiceSmall) for recognition → text is entered into the input field.
dsh-imggenImggen
DeepSeek Harness plugin: text-to-image output with in-chat image cards, download button, history gallery, and provider selection tabs
dsh-attachment-downscaleAttachment Downscale
Automatic attachment fallback: DSH image limits (one side ≤2000px / ≤3.5MB / ≤40 million pixels) no longer reject oversized images; sharp automatically resizes or re-encodes them before storage, so original phone photos can be uploaded without exceeding limits.
dsh-plugin-88api-imagePlugin 88api Image
88API Image Studio for DSH: four Image2 and Nano Banana models for text-to-image, multi-reference editing, 2K/4K output, and sequential batches.
dock-mediaDock Media
Media player for the dock file explorer: plays audio (MP3, WAV, OGG, FLAC, M4A, ...) and video (MP4, WebM, MOV, ...) files, with a music player for pure audio and fullscreen playback for video.
dsh-text2img-compressText2img Compress
把长文本渲染成图片发送来压缩 LLM 输入 token(DeepSeek 视觉模型、每图 384 token 封顶)— text-as-image token compression for DeepSeek Harness
dsh-computer-use-visionComputer Use Vision
Windows computer-use capability for DeepSeek Harness: screenshot → vision model → simulated mouse/keyboard input, with self-evolving knowledge base.
dsh-ds-vision-auto-routeDs Vision Auto Route
Route image-bearing turns to a configurable image-capable model for DeepSeek Harness
dsh-omnifileOmnifile
DSH file adapter plugin: drag, paste, or click to select multiple local files; anydoc parses documents, multimodal models recognize images, and the content is integrated for the main model.
@dsh-extension/dsh-generation-imageGeneration Image
On-demand image generation for DeepSeek Harness (DSH): a generate_image tool that calls your own OpenAI-compatible image URL + API key and delivers the image into the session
@dsh-external/dsh-image-tilerImage Tiler
DSH agent tool: slice a large image into labeled ~800x800 tiles plus an overview thumbnail, preserving detail for vision models.
@local/dsh-image-describeImage Describe
DeepSeek Harness host plugin: enables text-only models that do not support image input to "see" images (describe_image tool + image marker replacement)
dsh-pro-visionPro Vision
Let DeepSeek-V4-Pro (text-only) use V4-Flash-Vision-Exp for attached images. Mac/Windows/Linux.
dsh-official-visionOfficial Vision
Direct DeepSeek official vision API bridge for DeepSeek Harness: registers the deepseek-v4-flash-vision-exp multimodal model as a provider route, with base64 inline and Files API image input. 直连 DeepSeek 官方视觉 API 的 DSH 插件
dsh-vision-api-localorwebVision Api Localorweb
DSH plugin: Connects to image recognition model APIs (local large model image recognition tool + settings interface). Configure an OpenAI-compatible image recognition endpoint (LM Studio / vLLM / Ollama, etc.); leave the endpoint blank to disable the image recognition model.
deepseek-visual-pluginDeepseek Visual Plugin
DeepSeek Harness visual understanding plugin: Converts images in user messages and tool results into textual descriptions for text-only task models (such as DeepSeek).
dsh-culture-moviesCulture Movies
Movie genre
dsh-voice-input-cnVoice Input Cn
Voice input plugin for DeepSeek Harness web (China-ready): Alibaba Cloud DashScope ASR via a local bridge. Mic button in the composer, streaming recognition, cursor-aware insertion, silence auto-stop.
dsh-imageditImagedit
Local image editing toolkit for DeepSeek Harness: one image_edit tool for deterministic edits — cutout (quick flood-fill or rembg AI), trim, flip, rotate, brightness/contrast/saturation, blur, sharpen, rounded corners, border, canvas resize+center, sprite sheets, and PNG/JPEG/WebP export. Also ships
aura-visionAura Vision
Aura Vision — free vision OCR plugin for DeepSeek Harness web profile: Zhipu GLM-4V-Flash (free tier), adaptive tile recognition for long documents, history with favorites and Markdown/Excel/Word/PNG export.
dsh-attachment-formatsAttachment Formats
Codex-style attachment format expansion for the DeepSeek Harness Web GUI: PDF text-layer extraction (pymupdf4llm / pdfjs), Office text extraction, long-document spill + index cards, scanned-PDF OCR (tesseract.js), and browser-decodable images to PNG.
@mengruo/dsh-vision-toolkitVision Toolkit
DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, OCR, grounding, UI restoration, pixel diff, Artifacts, and Web UI.
dsh-reelsmakerReelsmaker
DeepSeek Harness plugin: turn a topic into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.
dsh-voice-announcerVoice Announcer
End-of-conversation voice announcement: session name + turn count + result (edge-tts streaming / SAPI)
@mokuyoaxis/dsh-irisIris
Give DeepSeek Harness eyes and hands: multi-provider media generation, vision routing, and an integrated Iris workbench.
taishan-visionTaishan Vision
Taishan Vision - Enables DeepSeek Harness text-only models to understand images: GLM visual recognition + inference by the selected model (static plugin package, permanently effective after restart; modified by JingQing v1.1.8)