DeepSeek Harness Plugin Hub

Publish and manage complete Harness Profiles. Discover Plugins for your next setup.

Explore

PluginsPresetsDocsNews

Community

Publish a pluginContactReport an issue

Resources

Plugin Hub on GitHubDeepSeek HarnessSystem statusPrivacy notice
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

Independent and unofficial. Not affiliated with, authorized by, or endorsed by DeepSeek.

Llm Vision — DSH Plugin for DeepSeek Harness
DeepSeek Harness Plugin Hub
ProfilesPluginsCategoriesNewsDocsSign inManage Profiles
ProfilesPluginsCategoriesNewsDocsSign in
← Plugins
L

dsh-llm-vision

Llm Vision

Model-facing describe_image + extract_text tools for the DeepSeek Harness web GUI: gives a text-only model reliable image understanding and OCR through an OpenAI-compatible vision endpoint, with critical-inspection prompts, auto-preprocessing, retries, and a persistent answer cache.

The plugin will be installed here. Keep web if you are unsure.

npx -y @deepseek-ai/dsh plugin --profile web add github:1710782766/dsh-llm-vision#a4a5272448c1b6df7d90b7b7afc6170e26c2ef43
READMECompatibilityVersions

Compatibility and provenance

Llm Vision is published as dsh-llm-vision and currently resolves to version 0.3.2. The Hub verifies its manifest and preserves the exact installation source for reproducible installs.

DSH compatibility
*
Runtime surfaces
web
Release source
github
Registry updated
8/31/2026

Versions

0.3.2stable
8/31/2026
0.3.0stable
8/24/2026

Related plugins

Loading related plugins…

Latest
0.3.2
DSH
*
HMR
Process restart
Tree shaking
Safe tree shaking not declared
Unpacked size
Unavailable
Files
Unavailable
Surface
web
License
Apache-2.0
Source
github
GitHub
★ 1
Weekly downloads
0
Last push
8/31/2026
View source ↗
README badge

Click the badge to copy Markdown for your README.

Do you maintain this Plugin?Claim benefit · Priority security scan

Verify the GitHub repository declared in package.json to manage this listing. After you claim it, Hub will prioritize a security scan of the current version and publish the result when it passes.

Claim this Plugin →
Report an issue

Related plugins

More verified plugins in vision-media.

Tool Describe Image@linxin666/dsh-tool-describe-imageModel-facing describe_image tool for the dsh web GUI: gives a text-only model image understanding by asking a vision-language model at an OpenAI-compatible endpoint to describe one image (local path, http(s) URL, or attachment reference). Hot-pluggable — Modlens@liustack/modlensPlug-in vision for text-only LLMs, powered by the free Antigravity CLIDeepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio and 27 editable video templates as a native DeepSeek Harness conversation view.Image Gendsh-image-genBring ChatGPT-like image generation to DeepSeek Harness — Gemini, OpenAI, Seedream, DashScope, local ComfyUI & more.

README

dsh-llm-vision

English | 中文

Give your DeepSeek Harness a pair of eyes — reliable image understanding and OCR for text-only models, configured entirely in the GUI.

Paste an image and the model describes or reads it; big screenshots are auto-compressed, transient failures retry, identical images hit a persistent cache. Built-in free presets (Zhipu / Gemini / DashScope) get you running without touching a config file.

Quick start

dsh plugin --profile web add dsh-llm-vision@0.3.2
  1. Install with the command above (or see Install).
  2. Restart the GUI once — plugins load at boot, so the card is not visible until then. Configuration changes after install never need a restart.
  3. Open Settings → Plugins → llm-vision and pick a Provider preset: zhipu, gemini, or dashscope fill the endpoint fields for you — free routes, no payment details (see the free presets table).
  4. Paste your API key into the card's API key field and Save — it is stored in your owner-only settings document and never shown again.
  5. Use it — paste / drag / drop an image into the composer and send; the model now sees it. Or call the llm_vision_check tool for a full pipeline diagnosis.

Why

Text-only models (DeepSeek V4, GLM text series, …) cannot see images. This plugin registers model-facing tools backed by any OpenAI-compatible vision endpoint:

ToolPurpose
describe_imageImage understanding with two perspectives: normal (natural description) and critical (objective inspection that actively reports text misalignment, overlap, occlusion, wrapping anomalies, missing elements, and separates fact from guess). The critical lens is the antidote to vision models rationalizing rendering bugs — use it for page/UI problem reports and screenshot-vs-design comparisons. Accepts a single image or a batch of up to 8 images read together in one call.
extract_textOCR & document parsing through a dedicated OCR model — ID cards, invoices, receipts; structured output (JSON/CSV) on request; verbatim extraction that never guesses missing text.
llm_vision_checkDiagnostics: verifies the configuration, that an API key resolves, and that the endpoint answers an authenticated probe — optionally with a real end-to-end vision call (testCall). The key itself never appears in the report.

Plus the DSH-native experience:

  • Paste / drag / drop images into the composer and send — the browser half rewrites the image-bearing send into attach references the text model can resolve, and upgrades the references into inline thumbnails in the transcript.
  • Live settings card (Settings → Plugins → llm-vision): endpoint, models, prompts, bounds, retries, preprocessing, and cache — saves apply to the very next call.
  • Three input kinds per call: local absolute path, http(s) URL (redirects refused), or attachment reference.
  • The image never enters the session log — only the returned text crosses into the conversation.

Install

dsh plugin --profile web add dsh-llm-vision@0.3.2

Then restart the GUI once — plugins load at boot, so the plugin and its settings card become visible only after the restart (configuration changes after that never need one).

The version is pinned on purpose: pnpm 11 holds back packages published in the last 24 hours, so a bare add dsh-llm-vision (latest) would silently install the previous release on launch day. This line is bumped with every release. --profile web is the GUI profile of this deployment — use your own profile name if it differs.

Requires dsh ≥ 0.1.2-alpha.1 — the settings-card host API and the browser-half store moved in that release; older harness builds cannot serve the card.

From a source checkout the same command accepts a tarball or local path (pnpm pack names the tarball after the current version — use that name):

pnpm install && pnpm build && pnpm pack   # → dsh-llm-vision-<version>.tgz
dsh plugin --profile web add ./dsh-llm-vision-<version>.tgz
# or: dsh plugin --profile web add /path/to/dsh-llm-vision   (build first — lib/ is gitignored)

The tarball ships prebuilt lib/ (both the node half and lib/client.js), so no build step runs on the installing machine.

Configure

Everything is configured in the GUI — the Settings → Plugins → llm-vision card. No patch file, no environment exports required:

  1. Open Settings → Plugins and find the llm-vision card.
  2. Pick a Provider preset for a zero-config route, or set baseURL / model / ocrModel yourself (custom).
  3. Paste the API key into the card's API key field — the simple path: it is stored in the harness's owner-only settings document (~/.dsh/settings.yaml, 0600) and never shown again. Advanced: leave the field empty and let apiKeyEnv resolve through the credential seam instead (presets prefill e.g. DASHSCOPE_API_KEY; the default is VISION_API_KEY) — for users who prefer environment variables.
  4. Save — the change reaches the very next tool call, no restart.

Before configuring, the first call fails with a clear hint (llm-vision: baseURL must be an absolute http(s) URL) — that is the expected unconfigured state, not a broken install.

The values live in the harness settings document (~/.dsh/settings.yaml, 0600, shared across profiles) and are written by the GUI. A profile patch layer may still provide deployment defaults for the card (shown as "Inherit"), but the card's saved values always win — the GUI is the only configuration surface a user needs. A deployment without a settings provider falls back to the built-in defaults.

Free presets (zero-cost routes)

The Provider preset selector fills baseURL / model / ocrModel / apiKeyEnv for you — free routes, no payment details. Free policies change, so re-check the provider docs if a call stops working:

PresetEndpointGetting a free key
zhipuZhipu BigModel — permanently free GLM-4V-Flash; the best default in mainland Chinaopen.bigmodel.cn — register, create an API key; free tier, no card
geminiGoogle Gemini — free key from Google AI Studio (aistudio.google.com)AI Studio → "Get API key", no card; not reachable from mainland China without a proxy
dashscopeAlibaba DashScope (the default models) with free quotaAlibaba Cloud Bailian console (bailian.console.aliyun.com) — free quota; reachable from mainland China

Picking a preset prefills the endpoint fields (still editable before saving); explicit field values always win at call time. The free presets reuse the vision model for OCR (extract_text drives it with the OCR prompt) — free tiers are rate-limited, so they suit interactive use better than batch runs.

KeyDefaultMeaning
providercustomEndpoint preset: custom (all fields explicit), dashscope, zhipu (free GLM-4V-Flash), or gemini (free key). Explicit fields win.
baseURL— (required for custom)OpenAI-compatible root URL; /chat/completions or /responses appended per apiStyle.
modelpreset, else qwen3-vl-plusVision model for describe_image; optional thinking suffix :off/:low/:medium/:high.
ocrModelpreset, else qwen3.5-ocrOCR model for extract_text; same suffix support.
apiKey—Inline key, stored in the settings document (secret: never shown by the GUI).
apiKeyEnvVISION_API_KEYCredential-reference (env var name) resolved through the credential seam; empty disables.
criticalPromptbuilt-indescribe_image critical-perspective prompt when the model passes none.
normalPromptbuilt-indescribe_image normal-perspective prompt when the model passes none.
ocrPromptbuilt-inextract_text prompt when the model passes none.
apiStylechat-completionschat-completions or responses.
maxBytes10485760Image byte bound (local files and downloads). Hi-res PNG wallpapers (10–30 MB) exceed the default; raise it — preprocessing compresses after loading.
maxOutputTokens1024Output-token cap sent to the endpoint.
timeoutMs60000Per-attempt timeout.

Reliability engineering

  • Auto-preprocessing — images over 1568px are scaled, oversize files re-encoded (JPEG q85, transparent formats kept as PNG) via the macOS built-in sips; every failure silently falls back to the original image. HEIC/HEIF inputs are always re-encoded to JPEG (endpoints support HEIC unevenly), failing loudly only when sips is absent. Fixes the classic "big screenshot times out" failure.
  • Retries — transient errors retry up to maxRetries with exponential backoff (≤ 4s) under a shrinking per-attempt budget (total ≤ 2× timeout). Exhausted retries append (已重试 N 次). Caller cancellation aborts immediately without retry.
  • Persistent cache — identical image + model + prompt + preprocessing settings hit a content-addressed cache (SHA-256 over the image bytes) at ~/.cache/dsh-llm-vision/; only the text answer is stored, never image bytes; TTL 30 days, 500 entries, atomic writes, 0600/0700 permissions. Note: OCR results of sensitive documents are stored in plain text there — set cacheEnabled to false when that matters.

Security model

  • The vision request and any image download refuse HTTP redirects (redirect: 'error') — bearer credentials and image bytes never leave the configured endpoint.
  • Request bodies carry the base64 image but never the key; parsed credentials are never logged.
  • Only http(s) URLs and local paths are accepted; all other schemes are rejected.
  • Attach uploads are validated (strict base64, magic bytes, byte bound) before the attachment store persists them; only the reference JSON (text) enters the session.
  • Response bodies are capped (maxOutputTokens × 8 + 64 KiB) before parsing; error excerpts are bounded to 200 chars.
  • Calling the tools sends the image bytes to the configured endpoint — only hand the model images you are comfortable leaving your machine.

Testing status

A fully offline test suite (vitest, mock HTTP server, tmp-dir cache), a strict typecheck, and CI on every push. Verified end-to-end in the real DSH web GUI against a live OpenAI-compatible vision endpoint: describe_image reads a real image (DashScope qwen3-vl-plus), extract_text OCR returns real transcription (qwen3.5-ocr), the attach upload/readback routes work through the live web server, and the settings card renders and saves in the plugin-configuration page — see Configure.

Known limitations

  • Attachment/upload channel: PNG / JPEG / GIF / WebP only (the official attachment store's type set). HEIC/HEIF images are read directly by the tools from local paths and URLs — preprocessing re-encodes them to JPEG on macOS — but pasting a HEIC file into the GUI is rejected with a hint; convert it or pass the path instead. On Windows/Linux (no sips) a HEIC/HEIF read fails with a clear message.
  • Preprocessing relies on macOS sips (zero dependencies); on Windows/Linux the plugin degrades silently and sends the original bytes — never an error, but oversized images are then likelier to time out. The bound gates loading (maxBytes), so a too-small bound is a clean rejection, never a crash.

Development

pnpm typecheck   # tsc -b + vitest program
pnpm test        # vitest run (fully offline)
pnpm build       # tsc -b && tsdown → lib/ + lib/client.js
pnpm watch       # tsdown --watch

License & attribution

Apache-2.0. Built on: deepseek-harness packages/vision/tool-describe-image (whitelonng/dsh-plugin-describe-image, MIT), the dsh-web-ui plugin family (Apache-2.0), and the llm_vision design (MIT). See NOTICE and AGENTS.md.

maxRetries2Retries for transient failures (timeout / network / 429 / 5xx); 0 disables.
maxEdge1568Max image edge (px) before auto-scaling; 0 disables preprocessing.
compressEnabledtrueAuto scale/re-encode oversize images (macOS sips; skipped elsewhere).
cacheEnabledtruePersistent content-addressed answer cache (cross-session).
cacheDir$XDG_CACHE_HOME/dsh-llm-visionCache directory.
cacheTtlDays30Cache entry lifetime (days).
cacheMaxEntries500Cache capacity; oldest evicted.
renderImagePreviewtrueUpgrade attach references into inline thumbnails (display only).
interceptImageSendtrueRewrite image-bearing sends into attach references at submit; turn off to hand raw image blocks to other vision plugins.