DeepSeek Harness Plugin Hub

Publish and manage complete Harness Profiles. Discover Plugins for your next setup.

Explore

PluginsPresetsDocsNews

Community

Publish a pluginContactReport an issue

Resources

Plugin Hub on GitHubDeepSeek HarnessSystem statusPrivacy notice
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

Independent and unofficial. Not affiliated with, authorized by, or endorsed by DeepSeek.

Plugin Vision Toolkit — DSH Plugin for DeepSeek Harness
DeepSeek Harness Plugin Hub
ProfilesPluginsCategoriesNewsDocsSign inManage Profiles
ProfilesPluginsCategoriesNewsDocsSign in
← Plugins

dsh-plugin-vision-toolkit

Plugin Vision Toolkit

Vision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images

The plugin will be installed here. Keep web if you are unsure.

npx -y @deepseek-ai/dsh plugin --profile web add dsh-plugin-vision-toolkit@0.1.1
READMECompatibilityVersions

Compatibility and provenance

Plugin Vision Toolkit is published as dsh-plugin-vision-toolkit and currently resolves to version 0.1.1. The Hub verifies its manifest and preserves the exact installation source for reproducible installs.

DSH compatibility
*
Runtime surfaces
any
Release source
npm
Registry updated
9/20/2026

Versions

0.1.1stable
8/22/2026
0.1.0stable
8/13/2026

Related plugins

Loading related plugins…

Latest
0.1.1
DSH
*
HMR
Process restart
Tree shaking
Safe tree shaking not declared
Unpacked size
29.1 kB
Files
28
Surface
any
License
MIT
Source
npm
GitHub
★ 0
Weekly downloads
65
View source ↗Project homepage ↗
README badge

Click the badge to copy Markdown for your README.

Do you maintain this Plugin?Claim benefit · Priority security scan

Verify the GitHub repository declared in package.json to manage this listing. After you claim it, Hub will prioritize a security scan of the current version and publish the result when it passes.

Claim this Plugin →
Report an issue

Related plugins

More verified plugins in vision-media.

Tool Describe Image@linxin666/dsh-tool-describe-imageModel-facing describe_image tool for the dsh web GUI: gives a text-only model image understanding by asking a vision-language model at an OpenAI-compatible endpoint to describe one image (local path, http(s) URL, or attachment reference). Hot-pluggable — Modlens@liustack/modlensPlug-in vision for text-only LLMs, powered by the free Antigravity CLIDeepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio and 27 editable video templates as a native DeepSeek Harness conversation view.Imagegen@dickpy/dsh-imagegenAI image generation plugin for the dsh web GUI: text-to-image and image-to-image through configurable provider channels (gpt-image-2 / grok-imagine-image / nanobanana series / seedream-5.0-pro / dall-e-3, with native xAI Grok Imagine, Google Nano Banana a

README

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.

What it does

Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.

Tools

  • glance -- describe, ask about, or OCR an image
  • ground -- locate a specific element (returns bounding box)
  • detect -- find all instances of an element kind
  • crop -- cut a region from an image

Install

dsh plugin --profile your-profile add dsh-plugin-vision-toolkit

Configuration

Set environment variables:

export VISION_API_KEY=sk-xxx           # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://...     # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2       # Vision model name

Usage examples

# Describe an image
glance screenshot.png

# Ask a question
glance screenshot.png -q "What error is shown?"

# OCR
glance screenshot.png --ocr

# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870

# Find all buttons
detect screenshot.png "buttons"

# Crop a region
crop screenshot.png 450,820,620,870 button.png

How it works

The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:

  1. Reads the image file
  2. Encodes it as base64
  3. Sends it to the vision API with a prompt
  4. Returns the text response

The agent never sees raw pixels -- it gets text descriptions it can reason about.

Supported vision providers

  • DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
  • OpenAI GPT-4V / GPT-4o
  • Any OpenAI-compatible multimodal endpoint

License

MIT -- YYTbit