DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Plugin Vision Toolkit — DeepSeek Harness 插件(DSH Plugin)
← Plugins

dsh-plugin-vision-toolkit

Plugin Vision Toolkit

DeepSeek Harness 的视觉工具包——glance、ground、detect、crop CLI 工具,帮助仅支持文本的智能体理解图像

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add dsh-plugin-vision-toolkit@0.1.1
README兼容性版本

兼容性与来源证明

Plugin Vision Toolkit 以 dsh-plugin-vision-toolkit 发布,当前版本为 0.1.1。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
any
发布来源
npm
Registry 更新时间
2026/9/20

版本

0.1.1stable
2026/8/22
0.1.0stable
2026/8/13

相关插件

正在加载相关插件…

最新版
0.1.1
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
29.1 kB
文件数
28
Surface
any
许可证
MIT
发布源
npm
GitHub
★ 0
周下载
65
查看源码 ↗项目主页 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录

相关插件

继续浏览 vision-media 分类下经过校验的插件。

Tool Describe Image@linxin666/dsh-tool-describe-image面向模型的 describe_image 工具,用于 dsh Web GUI:通过在兼容 OpenAI 的端点调用视觉语言模型,为文本模型提供图像理解能力,以描述一张图像(本地路径、http(s) URL 或附件引用)。可热插拔 —Modlens@liustack/modlens面向仅支持文本的 LLM 的插件视觉能力,由免费的 Antigravity CLI 提供支持Deepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio,以及 27 个可编辑视频模板,以原生 DeepSeek Harness 对话视图呈现。Imagegen@dickpy/dsh-imagegendsh Web GUI 的 AI 图像生成插件:通过可配置的提供商渠道实现文生图和图生图(gpt-image-2 / grok-imagine-image / nanobanana series / seedream-5.0-pro / dall-e-3,支持原生 xAI Grok Imagine、Google Nano Banana a

README

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.

What it does

Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.

Tools

  • glance -- describe, ask about, or OCR an image
  • ground -- locate a specific element (returns bounding box)
  • detect -- find all instances of an element kind
  • crop -- cut a region from an image

Install

dsh plugin --profile your-profile add dsh-plugin-vision-toolkit

Configuration

Set environment variables:

export VISION_API_KEY=sk-xxx           # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://...     # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2       # Vision model name

Usage examples

# Describe an image
glance screenshot.png

# Ask a question
glance screenshot.png -q "What error is shown?"

# OCR
glance screenshot.png --ocr

# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870

# Find all buttons
detect screenshot.png "buttons"

# Crop a region
crop screenshot.png 450,820,620,870 button.png

How it works

The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:

  1. Reads the image file
  2. Encodes it as base64
  3. Sends it to the vision API with a prompt
  4. Returns the text response

The agent never sees raw pixels -- it gets text descriptions it can reason about.

Supported vision providers

  • DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
  • OpenAI GPT-4V / GPT-4o
  • Any OpenAI-compatible multimodal endpoint

License

MIT -- YYTbit