CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
377 pluginsxby-asr-1Xby Asr 1
General-purpose speech recognition supporting multiple languages and lesser-spoken languages.
xby-asr-5Xby Asr 5
Five commonly used speech recognition languages: Mandarin Chinese, English, Japanese, Korean, and Cantonese, with automatic language detection.
xby-asr-fXby Asr F
Speech recognition supporting Mandarin Chinese and over 20 dialects and accents.
xby-asr-zhXby Asr Zh
Chinese speech recognition
xby-birdXby Bird
Detect and identify birds in images.
xby-captchaXby Captcha
CAPTCHA recognition toolkit supporting the recognition of text, slider, rotation, selection, and other verification methods. Note: Always comply with the terms of use and applicable laws and regulations of the target website or system, and use it only where permitted.
xby-cellphone-detectionXby Cellphone Detection
Input an image to detect mobile phones and output the bounding boxes, confidence scores, and labels for all targets in the image.
xby-classifyXby Classify
Classify images into 1000 ImageNet classes and return the Top-5 classes and confidence scores.
xby-daily-object-detectionXby Daily Object Detection
Input an image, detect people, pets, cars, fire, and cardboard boxes, and output the bounding boxes, confidence scores, and labels for all detected objects in the image.
xby-detect-vehicleXby Detect Vehicle
Input an image to detect vehicle types (car/truck/bus/motorbike/tricycle/carplate) and output the bounding boxes, confidence scores, and labels for all detected objects.
xby-detectXby Detect
Includes everyday object detection, insect recognition, plant recognition, animal recognition, electric bicycle detection, mobile phone detection, gesture detection, flame detection, cigarette detection, human head and body detection, wildlife detection, bird recognition, pet emotion recognition, di
xby-dishXby Dish
Dish recognition, output possible dish names and probabilities.
xby-ebike-detectionXby Ebike Detection
Input an image, detect the electric bicycles in it, and output the bounding boxes, confidence scores, and labels for all targets in the image.
xby-extract-imageXby Extract Image
The MCP server provides functionality to extract images from local files and URLs and convert them to base64 format, suitable for LLM analysis.
xby-fire-detectionXby Fire Detection
Detect flames in various general scenarios; best suited for security camera and traffic camera views.
xby-general-recognitionXby General Recognition
Recognizes labels in images containing a primary object and outputs the object's category label. Currently covers more than 50,000 object categories.
xby-gesture-detectionXby Gesture Detection
Input an image, detect the gestures in it, and output the bounding boxes, confidence scores, and labels for all detected targets.
xby-head-person-detectionXby Head Person Detection
Input an image, detect human heads and bodies, and output the bounding boxes, confidence scores, and labels for all detected targets.
xby-helmet-headXby Helmet Head
Input an image to detect human bodies, heads, and safety helmets, and output the bounding boxes, confidence scores, and labels for all detected objects.
xby-image-detectXby Image Detect
Detects 80 classes of COCO objects (people, vehicles, animals, everyday items, etc.) in images, outputting bounding boxes, confidence scores, and class labels.
xby-insect-recognitionXby Insect Recognition
Identify the name of an insect or other arthropod (or its order, family, genus, or species).
xby-logo-analyzeXby Logo Analyze
An intelligent logo extraction and processing MCP server that supports automatically identifying and extracting logo icons from website URLs, and provides image processing and vector conversion features.
dsh-omi-voiceOmi Voice
Read DeepSeek Harness conversations aloud: Doubao audio quality · click to read/pause/resume · Doubao Key stays only in the Omi engine (BYOK)
@lijian-ui/dsh-vision-toggleVision Toggle
Per-model vision (image input) toggle for DeepSeek Harness (dsh): list every configured model and flip a switch to enable/disable image support without hand-editing settings.yaml. 为 DeepSeek Harness 提供按模型的「支持图片」开关:无需手改 settings.yaml。
@dopilot/dsh-plugin-image-genPlugin Image Gen
OpenAI-compatible image generation tool for DeepSeek Harness
dsh-pro-vision-stdPro Vision Std
dsh-std Community v0.15 ModelProvider: DeepSeek-V4-Pro with images captioned by V4-Flash-Vision-Exp. Requires @dsh-std/adapter-dsh on DeepSeek Harness.
dsh-periscopePeriscope
DSH Desktop plugin: keep text-only models (deepseek-v4-flash / deepseek-v4-pro) as the session default and automatically route image-bearing requests to a configured vision model; also ships the generic file-attachment feature (drag/paste any file as a durable attachment the agent reads by path) as
dsh-handwritten-ocrHandwritten Ocr
Local OCR for DSH: handwritten Chinese and math formulas to Markdown with LaTeX. GPU (DirectML) / CPU / NPU backends, settings-driven, one-click install.
dsh-image-pickerImage Picker
DeepSeek Harness Web GUI input box 📎 image-picker button: Add reference images through the system file picker, bypassing drag-and-drop environment issues and reusing the official attachment pipeline (thumbnail generation, limit validation, and upload with the message).
dsh-ocr-pluginOcr Plugin
DeepSeek Harness OCR plugin — based on 小笨羊 OCR capabilities, providing image text recognition for text-only models