CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
377 plugins@ycp424c/dsh-luna-vision-bridgeLuna Vision Bridge
DSH LLM adapter that transcribes native image attachments with Codex Luna before delegating to DeepSeek
dsh-image-serveImage Serve
Let AI assistants directly "see" local images in sessions: /ws/<absolute path> turns files anywhere on disk into same-origin HTTP URLs renderable in session Markdown. Show any local file in DSH web sessions — no staging directory, any drive.
dshtools-sensevoice-inputTools Sensevoice Input
SenseVoiceSmall-powered local speech-to-text for the DSH Desktop input box: click the mic beside the composer, speak, and the recognized draft (with language/emotion tags) is written into the input.
dsh-computer-use-windowsComputer Use Windows
Windows Computer Use for DeepSeek Harness: window-bound screenshots, robust OCR, verified clicks, pure-OCR mode, pluggable vision models.
dsh-unlimited-ocr-skillUnlimited Ocr Skill
Unlimited-OCR long-document parsing with a native DeepSeek Harness tool and GUI configuration.
dsh-media-serveMedia Serve
DeepSeek Harness host plugin: serves a configured workspace/media folder over the web server under /media, so any conversation can display local images by referencing http://<host>:<port>/media/<relative-path>.
agnes-mediaAgnes Media
Registers generate_image and generate_video tools for Agnes AI media models (agnes-image-2.1-flash, agnes-video-2.5-flash) in DeepSeek Harness. Supports multi-image reference via reference_image_urls.
@local/dsh-tool-imagegenTool Imagegen
DSH conversational inline image generation plugin: OpenAI-compatible API; images are displayed directly in the dialog as generated-image blocks; settings are located under Settings → Plugins → Configurable.
dsh-vox-inputVox Input
Voice (speech-to-text) input for the DSH Web composer via the browser Web Speech API — tap the mic, speak, the transcript fills the input box. Zero server, zero API keys, nothing leaves the machine. · DSH Web 语音输入:点麦克风说话,识别文字回填输入框(Chrome/Edge,无需 API key)。
dsh-ppt-masterPpt Master
PPT Master skill for DeepSeek Harness: AI-driven presentation workflow for generating editable PPTX decks, SVG snapshots, native template filling, and PPTX enhancement.
dsh-v4flash-tilerV4flash Tiler
DSH plugin: auto-tiles oversized chat images into labelled grid pieces (row/col metadata, overlap-aware, multi-image group isolation) so the DeepSeek vision model keeps fine detail instead of the ~800px downsample.
dsh-vocoVoco
DSH plugin dsh-voco
dsh-anydoc-markdownAnydoc Markdown
Converts documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, PDF) to clean Markdown via the Rust firecrawl-anydoc crate and describes every embedded image with a vision model. Host plugin exposing a `convert_document` tool; the conversion + batched-vision engine is a bundled Python wra
dsh-llm-visionLlm Vision
Model-facing describe_image + extract_text tools for the DeepSeek Harness web GUI: gives a text-only model reliable image understanding and OCR through an OpenAI-compatible vision endpoint, with critical-inspection prompts, auto-preprocessing, retries, and a persistent answer cache.
dsh-live-voiceLive Voice
DeepSeek Harness live voice preview with exact-session consent, a local synthetic demo, and one bounded manual turn
dsh-tool-imagegenTool Imagegenalt · Pappet
Text-to-image generation tool for DeepSeek Harness via OpenRouter's unified Image API, with capability-gated parameters and workspace file output
@difimim/dsh-voice-inputVoice Input
DeepSeek Harness voice input plugin: adds a microphone button to the input field and uses the browser Web Speech API to transcribe speech into text in real time.
dsh-screenshot-reviewScreenshot Review
Screenshot review dsh skill: the model takes screenshots, reviews them, modifies the code, and iteratively improves the frontend.
dsh-video-genVideo Gen
Bring text-to-video and image-to-video generation to DeepSeek Harness — DashScope Wanx, Volcengine/Doubao Seedance, Google Veo, OpenAI Sora & compatible relays.
dsh-inline-media-viewerInline Media Viewer
Persistent inline image, video, and audio previews for DeepSeek Harness Web conversations, with workspace-confined local reads and a configurable ComfyUI proxy.
@mengli114/dsh-image-tilerImage Tiler
DSH agent tool: slice a large image into labeled ~800x800 tiles plus an overview thumbnail, preserving detail for vision models.
dsh-drop-previewDrop Preview
DSH WebUI drag-and-drop file preview: drop files anywhere, full-page preview images and rendered Markdown (inline images resolved from dropped files or auto-searched in the session workspace), zoom/rotate images on click, persistent composer file box, smart desktop screenshot (full screen / free reg
visionVision
Vision DSH bundle: give the agent screen/window vision — see tool captures via the bundled Python cvision and returns the image natively
@alger-ai/dsh-image-previewImage Preview
Image preview in the dsh session: read_image tool results render as a small thumbnail by default and open full-size in the built-in lightbox on click
dsh-labnanaLabnana
Labnana image generation for DeepSeek Harness: text-to-image / image-to-image / precise editing with credits estimation, subscription balance and web settings UI.
dsh-read-image-jpeg-fallbackRead Image Jpeg Fallback
DSH profile plugin: converts read_image PNG/WebP attachments to opaque sRGB JPEG so LM Studio's openai-completions endpoint accepts them
dsh-video-playerVideo Player
Floating, draggable, resizable video scene player for DeepSeek Harness. Plays scene-per-MP4 clips from a Stash-style scene server, plus best-effort YouTube / Twitch / Jellyfin / custom links.
phone-eyePhone Eye
Let your AI agent see and operate a real Android phone — vision + UI-tree fusion over adb, for any MCP client
dsh-voice-inputVoice Inputalt · lhenlihai-hub
Voice dictation and transcript cleanup for DeepSeek Harness, using the current session model.
dsh-grok-imageGrok Image
DeepSeek Harness plugin: Grok Imagine image generation as a model tool (subscription)