CATEGORY
Vision & Media DSH plugins
Verified manifests and exact versions in this category.
All plugins
377 pluginsdsh-filesFiles
DeepSeek Harness dual-face plugin: a composer file-upload button (paperclip + folder + drag-drop, @ dual-source candidates) with session-isolated storage, TTL sweep and sha256 dedup; native image support that hands JPEG/PNG/WebP/GIF to any image-capable model via the core attachment pipeline; plus a
dsh-maclensMaclens
Bridge Apple's on-device Vision framework (macOS) into DeepSeek Harness: OCR, image classification, face detection, and document layout as local dsh tools. No network, no API key, no daemon.
dsh-tu4-inline-imagesTu4 Inline Images
Inline images in conversations DSH plugin — renders local image paths in the input box, user messages, and LLM replies directly as images (loopback routing + strong token + multiple-root allowlist + three size presets). Inline local images in DSH web conversations.
dsh-model-auto-hot-switchModel Auto Hot Switch
Automatic per-task model hot-switching for DeepSeek Harness (dsh): image-aware tasks route to the vision model automatically, every other task keeps your default model. Zero extra tokens, no context disturbance.
xby-obbXby Obb
检测图像中的旋转目标,输出旋转边界框、角度、置信度和类别标签。支持15个目标类别:plane, ship,storage tank,baseball diamond,tennis court,basketball court,ground track field, harbor, bridge,large vehicle,small vehicle, helicopter, roundabout,soccer ball field,swimming pool。
xby-ocr-bank-cardXby Ocr Bank Card
Identify bank card numbers, issuing banks, and card types; validate card number validity using the Luhn algorithm.
xby-ocr-captchaXby Ocr Captcha
Submit a common CAPTCHA image and return its text.
xby-ocr-driver-licenseXby Ocr Driver License
Recognizes the main page of a driver's license (license number, name, gender, nationality, address, date of birth, permitted vehicle classes, date of first issuance, validity period) and the secondary page (file number).
xby-ocr-handwritingXby Ocr Handwriting
Input an image containing handwritten text to automatically detect text lines and recognize their content. Suitable for handwritten notes, signatures, handwritten forms, and more.
xby-ocr-id-cardXby Ocr Id Card
Recognizes the front of an ID card (name, gender, ethnicity, date of birth, address, and ID number) and the back (issuing authority and validity period), automatically determines the front or back, and validates the ID number.
xby-ocr-passXby Ocr Pass
Recognizes the permit number, name, gender, date of birth, validity period, place of issue, and other information on Exit-Entry Permits for Travelling to and from Hong Kong and Macao and Taiwan Travel Permits, with support for MRZ parsing.
xby-ocr-passportXby Ocr Passport
Recognizes passport number, Chinese name, English name, gender, nationality, date of birth, date of issue, expiry date, place of issue, and other information; supports MRZ parsing.
xby-ocr-proXby Ocr Pro
High-precision text recognition. Input an image containing text to automatically detect and recognize its content. Suitable for various scenarios, including documents, billboards, and screenshots.
xby-ocr-vehicle-licenseXby Ocr Vehicle License
Recognizes the license plate number, vehicle type, owner, address, make and model, engine number, vehicle identification number, and other information on motor vehicle registration certificates. Supports automatic orientation detection and filtering of primary and secondary pages.
xby-ocrXby Ocr
Text recognition balancing speed and accuracy. Automatically detects and recognizes content in images containing text. Suitable for various documents, billboards, screenshots, and other scenarios.
xby-pedestrianXby Pedestrian
Input an image, detect pedestrians in the image, and output the bounding boxes, confidence scores, and labels for all detected objects.
xby-pet-detectXby Pet Detect
Recognize pet facial expressions and output 4 emotion categories: Angry / Happy / Relaxed / Sad.
xby-picXby Pic
Includes general text recognition, handwriting recognition, license plate recognition, ID card recognition, everyday object detection, insect recognition, plant recognition, passport recognition, Hong Kong, Macao, and Taiwan travel permit recognition, bank card recognition, business license recognit
xby-plant-recognitionXby Plant Recognition
Identify the plant name (or its family, genus, species, or subspecies).
xby-plate-recognitionXby Plate Recognition
Recognize license plate numbers, license plate colors, single- and double-layer license plates, and bounding boxes.
xby-poseXby Pose
Detect people in images and output bounding boxes and keypoint coordinates. Each person has 17 keypoints, with each point representing a different body part, in order: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, lef
xby-read-pdfXby Read Pdf
An MCP server that enables AI assistants to read and analyze PDF files, providing PDF metadata extraction, page-range reading, and keyword search.
xby-recogXby Recog
Includes general text recognition, handwriting recognition, license plate recognition, ID card recognition, passport recognition, Mainland Travel Permits for Hong Kong, Macao, and Taiwan, bank card recognition, business license recognition, driver's license recognition, and vehicle registration cert
xby-reflective-vestXby Reflective Vest
Input an image, detect whether personnel are wearing reflective clothing, and output bounding boxes, confidence scores, and labels (safe/unsafe) for all targets in the image.
xby-segXby Seg
Instance segmentation goes a step beyond object detection: it not only identifies individual objects in an image, but also segments them from the rest of the image. Perform instance segmentation on the 80 COCO object categories in an image, outputting bounding boxes, masks, confidence scores, and cl
xby-smoking-detectionXby Smoking Detection
Input an image to detect cigarette targets and output the bounding boxes, confidence scores, and labels for all detected targets in the image.
xby-speech-synthesisXby Speech Synthesis
An MCP server integrating Microsoft Edge's high-quality text-to-speech capabilities, supporting multilingual speech generation, audio merging, and cloud storage.
xby-wikimedia-search-imagesXby Wikimedia Search Images
This MCP server enables AI assistants to search for images on Wikimedia Commons, providing detailed metadata and optional thumbnail sets to help AI models perform visual comparisons.
xby-wild-animal-detectionXby Wild Animal Detection
Input an image and output the bounding boxes, confidence scores, and labels of all detected wild animals in the image.
xby-animal-recognitionXby Animal Recognition
Identifies animal category labels in images without any additional input.