DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Deepsee — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins

@chang416/deepsee

Deepsee

DeepSee:DeepSeek Harness 视觉能力、多模型路由,以及交付前的 Gemini 视觉自检

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add @chang416/deepsee@4.0.2
README兼容性版本

兼容性与来源证明

Deepsee 以 @chang416/deepsee 发布,当前版本为 4.0.2。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
web
发布来源
npm
Registry 更新时间
2026/9/20

版本

4.0.2stable
2026/8/15
4.0.1stable
2026/8/15
4.0.0stable
2026/8/14

相关插件

正在加载相关插件…

最新版
4.0.2
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
424.7 kB
文件数
34
Surface
web
许可证
MIT
发布源
npm
GitHub
★ 2
周下载
54
最近提交
2026/9/4
查看源码 ↗项目主页 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题

相关插件

继续浏览 vision-media 分类下经过校验的插件。

Tool Describe Image@linxin666/dsh-tool-describe-image面向模型的 describe_image 工具,用于 dsh Web GUI:通过在兼容 OpenAI 的端点调用视觉语言模型,为文本模型提供图像理解能力,以描述一张图像(本地路径、http(s) URL 或附件引用)。可热插拔 —Modlens@liustack/modlens面向仅支持文本的 LLM 的插件视觉能力,由免费的 Antigravity CLI 提供支持Deepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio,以及 27 个可编辑视频模板,以原生 DeepSeek Harness 对话视图呈现。Codexdsh-codexChatGPT OAuth、Codex 模型、搜索、read_image URL 支持,以及适用于 DeepSeek Harness 的 gpt-image-2 生成

README

DeepSee — vision and model routing for DeepSeek Harness

DeepSee

Vision and smart model routing for DeepSeek Harness.

简体中文 · Troubleshooting · Configuration · Output contract · Security

DeepSee turns DeepSeek Harness into a multimodal, multi-model coding workspace. Gemini sees. DeepSeek codes. Choose Flash or Pro directly, or let Auto and Customize split work between them. Paste screenshots into a text-only DeepSeek session, route tasks by cost and difficulty, and visually check the result before delivery.

npx -y @deepseek-ai/dsh plugin --profile web add @chang416/deepsee@latest

Then confirm what landed with dsh plugin --profile web list. pnpm v11 quarantines releases published in the last few days and can install an older version instead while still reporting success; troubleshooting fixes that in one line.

Open DeepSee Settings to add free Gemini keys (one per line), choose a default preview URL, and customize which work belongs to Flash or Pro.

DeepSee routes a task, reads the rendered UI, and visually checks the result

Watch the 30-second DeepSee product film

Watch the 30-second product film

How DeepSee combines Gemini vision with DeepSeek model routing

Highlights

DeepSee is both sight and a DeepSeek-native team. The model selector keeps direct V4 Flash and V4 Pro choices and adds two orchestration modes:

  • DeepSee Auto ships with a free-first routing policy. Flash takes discovery, documentation, tests, small edits, and bounded bug fixes; Pro takes architecture, security, risky refactors, integration, and final review. Independent subtasks run in parallel and the coordinator merges the result.
  • DeepSee Customize lets each user choose Flash or Pro for every work category. Selecting it for the first time opens DeepSee Settings inside the Harness interface, where the routing map can be changed without editing JSON.
  • It looks before it delivers. For UI work, DeepSeek can start the local preview and call Gemini at meaningful milestones and again before delivery. Gemini returns a strict PASS or a screen location plus the defect to fix; DeepSeek iterates within a configurable free-quota limit instead of asking the user to discover visual mistakes afterward.
  • DeepSeek writes the code. Gemini is used only as the visual reader. The coding lanes stay on DeepSeek V4 through whichever live route the user already has: the official provider, OpenCode's deepseek-v4-flash-free, or OpenCode Go's deepseek-v4-flash and deepseek-v4-pro.
  • Multiple free Gemini keys. Paste one key per line. DeepSee deduplicates them and automatically advances to the next key on authentication, quota, or rate-limit exhaustion; saved keys are write-only in the interface and never returned to the browser.

One install adds the native read_image bridge, DeepSee Auto and Customize, multi-key Gemini rotation, OpenCode/OpenCode Go-aware DeepSeek routing, and the deepsee_visual_check delivery gate. If dsh warns declares no dsh.bundle, see troubleshooting.

Pasting an image works two ways. ① Just paste. On a text-only model the pasted image lands as a private temp file and its path enters the composer — the same interaction OpenCode and Pi ship — and the read_image tool takes it from there. ② Pick a (deepsee vision) entry in the model selector (it remembers your choice, so once is enough), then paste: the thumbnail stays visible in your message, closer to the Codex app feel, and the image is converted to structured evidence at request time, answered by the same underlying route. The plugin auto-discovers every provider route carrying text-only DeepSeek or GLM models and adds a wrapped entry per route (a stock install gets DeepSeek-V4-Flash (deepsee vision) and DeepSeek-V4-Pro (deepsee vision); extra routes like opencode-go or zai get their own); the two families' own vision models are excluded automatically. Which paste route applies is the host's per-model call: only a model its metadata positively confirms text-only is taken over, anything unconfirmed is left alone, so vision models keep their native paste (details).

Paste an image and DeepSeek can use it. No model swap and no manual transcription.

  • Native and removable. DeepSee is one dsh plugin or one skill folder, with no local proxy daemon. Remove it and the host returns to its original behavior.
  • Zero-config start. Reuses what Claude Code, Codex, OpenCode, or Pi already have set up: the multimodal models on your machine go straight to work. Nothing at all? Antigravity CLI is a free no-key channel, and a free Gemini key brings a read down to 5-10 seconds.
  • Evidence, not imagination. Full transcription, reading-order layout regions, entity and relation lists. The model quotes specifics.
  • Install once, use everywhere. Verified on real machines in Claude Code, Codex, Pi, and OpenCode.

Installation

Step 1, hand it to your AI. Send it this line:

Install and configure the deepsee skill following https://github.com/chang416/deepsee/blob/main/INSTALL.md, then run the health check and tell me the result.

The install starts by checking what your machine already has. An existing login in Claude Code, Codex, OpenCode, or Pi can be enough: deepsee asks before reusing any of them, and the health check tells you where things stand.

Step 2, only if the health check comes back empty, set up a free engine. The recommended choice is a free Gemini API key (about three minutes at Google AI Studio, no credit card), which also makes every read 5-10 seconds. A free OpenAI-compatible key from another platform works too. To avoid any sign-up, install Antigravity CLI instead, then sign in:

curl -fsSL https://antigravity.google/cli/install.sh | bash
agy                                                           # sign in, then exit

The install also inventories vision reachable through your other local harness CLIs (Codex, OpenCode, Pi) and asks, per harness, whether deepsee may reuse it. Granted logins join the engine pool as equals, and every reused read is labeled with whose quota it spent.

Usage

Once installed, just chat. Paste an image or drop a path, ask anything, and the skill triggers on its own: the image goes to a vision engine and the answer comes back grounded in what it read.

Vision engines: five built-in providers, four reusable CLIs, one failover chain

DeepSee does not depend on any single vision service. Nine sources of vision in total: five built-in providers, any one of which is enough, plus four local agent CLIs whose logins can be reused. The built-ins:

ProviderWhat it needsSpeed per readGood for
gemini-apia free Gemini API key (3 minutes, no card)5-10sthe recommended default
openaiany OpenAI-compatible endpoint (key + baseUrl + model)5-10sqwen-vl, GLM, self-hosted gateways
anthropican Anthropic API key5-10smachines already holding one
antigravity-clithe free agy CLI, one browser sign-in, no key15-45szero-signup starts
claude-clia signed-in Claude Code20-45sriding your existing Claude subscription

Without a pinned provider, every configured engine forms one failover chain: the fast API providers try first, the agent CLIs back them up, the first good result wins, and meta.attempts records every attempt so a fallback is never silent.

openai is a universal socket, not just OpenAI

Any endpoint speaking the OpenAI chat-completions protocol with image input plugs straight in — that covers most of the vision-model world:

deepsee config set openai.baseUrl https://dashscope.aliyuncs.com/compatible-mode/v1   # qwen-vl
deepsee config set openai.apiKey  <key>
deepsee config set openai.model   qwen3-vl-plus

The same three keys work for GLM's open platform, SiliconFlow, OpenRouter, a self-hosted vLLM/Ollama, or any gateway of your own. If your favorite vision model has an OpenAI-compatible API, DeepSee can drive it.

Reusing what your machine already has

Two more sources of vision need zero new keys, each behind one explicit consent recorded in config:

  • The harness you are talking in right now. Running inside Claude Code with a subscription signed in? claude-cli reads images through it out of the box. The install flow asks the same question for whichever harness you install into.
  • Every other agent CLI on the machine. deepsee doctor discovers them, you grant per harness, and they join the same failover chain with no priority over your own keys. Every reused read is labeled in meta.warnings with whose quota it spent, so nothing is ever silently billed:
Reused CLIWhat it needsGrant withRides as
Codexa signed-in Codex CLI with a vision modelconfig set reuse.codex trueagent lane, 15-45s
OpenCodea vision model configured in OpenCodeconfig set reuse.opencode trueagent lane, 15-45s
Pimodel credentials held by Piconfig set reuse.pi truean API key upgrades to the 5-10s inline lane, OAuth drives Pi itself
Groka signed-in Grok CLI (SuperGrok)config set reuse.grok trueagent lane, 15-45s

Picking and routing

Two knobs: deepsee config set provider <name> states a preference (the chain still backs it up), -p <name> pins exactly one with no fallback. Machines behind a proxy set HTTPS_PROXY or deepsee config set proxy <url> and the API providers route through it. Details: the CLI manual for defaults and flags, Configuration for every key, and Security for who fetches what on remote URLs.

Documentation

DocRead it when
Install guideInstalling the skill step by step (written for an agent)
CLI manualThe CLI the skill drives: flags, config, doctor
TroubleshootingA command failed and the message needs decoding
ConfigurationSetting a key, switching providers, fixing config
Output contractParsing the JSON or building on it
Harness setupWiring it into Codex, Claude Code, Pi, or OpenCode
SecurityFile permissions, image content as untrusted input
CHANGELOGFinding what changed in a version

Contributing

Focused pull requests are welcome. Keep each PR scoped, explain the user-visible behavior, add or update tests, and run pnpm lint, pnpm typecheck, pnpm test, and pnpm build before opening it.

  • Open an issue. Bugs, suggestions, confusing errors, and unclear docs all help.
  • Read CONTRIBUTING.md. Security reports follow SECURITY.md, not public issues.

Disclaimer

Provided as-is under the MIT License below. The author makes no warranty and gives no endorsement for any particular use, commercial use included. Your use of upstream engines (Antigravity CLI, the Gemini, OpenAI, and Anthropic APIs, and any OpenAI-compatible endpoint) is governed by their own terms and quotas, which you are responsible for.

Acknowledgements

DeepSee is designed, developed, and maintained by chang416. Early exploration referenced a small amount of the MIT-licensed ModLens project; its required copyright notice remains in LICENSE. DeepSee's product architecture, Auto/Customize orchestration, settings experience, Gemini key rotation, OpenCode-aware routing, and visual self-check loop are developed for DeepSee.

License

MIT