DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Vision Bridge — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins

@goodandready/dsh-vision-bridge

Vision Bridge

DeepSeek Harness 的旗舰多模态视觉中心:约 40 个工具,支持 PDF 拖放、LaTeX 公式、复杂表格、二维码、UI 流程图和多模型共识。

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add github:GooDAnDReaDY/dsh-vision-bridge#3065a4fee2f5cb2dec6aff158b8f0896a9ea9571
README兼容性版本

兼容性与来源证明

Vision Bridge 以 @goodandready/dsh-vision-bridge 发布,当前版本为 0.6.4。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
web
发布来源
github
Registry 更新时间
2026/9/20

版本

0.6.4stable
2026/9/19
0.6.3stable
2026/9/18
0.6.2stable
2026/9/18
查看其余 83 个版本收起版本
0.6.1stable
2026/9/17
0.6.0stable
2026/9/17
0.5.37stable
2026/9/13
0.5.36stable
2026/9/12
0.5.35stable
2026/9/11
0.5.34stable
2026/9/11
0.5.33stable
2026/9/11
0.5.32stable
2026/9/10
0.5.31stable
2026/9/10
0.5.30stable
2026/9/9
0.5.29stable
2026/9/7
0.5.28stable
2026/9/7
0.5.27stable
2026/9/5
0.5.26stable
2026/9/4
0.5.25stable
2026/9/4
0.5.24stable
2026/9/4
0.5.23stable
2026/9/2
0.5.22stable
2026/9/2

相关插件

正在加载相关插件…

最新版
0.6.4
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
未提供
文件数
未提供
Surface
web
许可证
MIT
发布源
github
GitHub
★ 1
周下载
1,019
最近提交
2026/9/19
查看源码 ↗项目主页 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题
0.5.21stable
2026/9/2
0.5.20stable
2026/9/2
0.5.19stable
2026/9/2
0.5.18stable
2026/9/1
0.5.17stable
2026/9/1
0.5.16stable
2026/9/1
0.5.15stable
2026/9/1
0.5.14stable
2026/9/1
0.5.13stable
2026/9/1
0.5.10stable
2026/9/1
0.5.12stable
2026/9/1
0.5.11stable
2026/9/1
0.5.9stable
2026/9/1
0.5.8stable
2026/9/1
0.5.7stable
2026/9/1
0.5.6stable
2026/9/1
0.5.5stable
2026/9/1
0.5.4stable
2026/9/1
0.5.3stable
2026/9/1
0.5.2stable
2026/9/1
0.5.1stable
2026/9/1
0.5.0stable
2026/9/1
0.4.13stable
2026/9/1
0.4.12stable
2026/9/1
0.4.11stable
2026/9/1
0.4.10stable
2026/9/1
0.4.9stable
2026/9/1
0.4.8stable
2026/9/1
0.4.7stable
2026/9/1
0.4.6stable
2026/9/1
0.4.5stable
2026/8/30
0.4.4stable
2026/8/29
0.4.3stable
2026/8/27
0.4.2stable
2026/8/27
0.4.1stable
2026/8/26
0.4.0stable
2026/8/26
0.3.9stable
2026/8/26
0.3.8stable
2026/8/25
0.3.7stable
2026/8/25
0.3.6stable
2026/8/25
0.3.5stable
2026/8/25
0.3.4stable
2026/8/25
0.3.3stable
2026/8/25
0.3.2stable
2026/8/25
0.3.1stable
2026/8/25
0.3.0stable
2026/8/25
0.2.14stable
2026/8/25
0.2.13stable
2026/8/25
0.2.12stable
2026/8/25
0.2.11stable
2026/8/25
0.2.10stable
2026/8/25
0.2.9stable
2026/8/25
0.2.8stable
2026/8/25
0.2.7stable
2026/8/25
0.2.6stable
2026/8/25
0.2.5stable
2026/8/25
0.2.4stable
2026/8/25
0.2.3stable
2026/8/25
0.2.2stable
2026/8/24
0.2.1stable
2026/8/24
0.2.0stable
2026/8/24
0.1.4stable
2026/8/19
0.1.3stable
2026/8/19
0.1.1stable
2026/8/18
0.1.0stable
2026/8/18

相关插件

继续浏览 vision-media 分类下经过校验的插件。

Tool Describe Image@linxin666/dsh-tool-describe-image面向模型的 describe_image 工具,用于 dsh Web GUI:通过在兼容 OpenAI 的端点调用视觉语言模型,为文本模型提供图像理解能力,以描述一张图像(本地路径、http(s) URL 或附件引用)。可热插拔 —Modlens@liustack/modlens面向仅支持文本的 LLM 的插件视觉能力,由免费的 Antigravity CLI 提供支持Deepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio,以及 27 个可编辑视频模板,以原生 DeepSeek Harness 对话视图呈现。Imagegen@dickpy/dsh-imagegendsh Web GUI 的 AI 图像生成插件:通过可配置的提供商渠道实现文生图和图生图(gpt-image-2 / grok-imagine-image / nanobanana series / seedream-5.0-pro / dall-e-3,支持原生 xAI Grok Imagine、Google Nano Banana a

README

📦 @goodandready/dsh-vision-bridge

Universal Vision Bridge for DeepSeek Harness: Seamless Multimodal Attachments with Text-Only Chat Models

🇬🇧 English • 🇷🇺 Русский • 🇨🇳 中文说明

⭐ If you like this plugin, please star it on GitHub — it shows me that the plugin is useful to you and motivates me to keep developing it.

🐛 If you find a bug or would like to request a feature, open a GitHub issue in any language — I will review your proposal and implement useful suggestions in a future plugin version.

⚡ Overview & The Problem

When interacting with text-only chat models (any provider/model that accepts text only) in DeepSeek Harness, users cannot natively attach and send images:

  1. In DSH 0.1.2-alpha.2+, the core session controller performs a strict server-side modality check (ctx.llm.resolveModelInfo). If the active conversation model lacks image in inputModalities, the prompt is immediately rejected with a session/attachment-invalid error ("Model does not support image input").
  2. Standard text-only adapters throw errors when encountering raw multimodal image blocks in their message payload.

How dsh-vision-bridge Solves This

dsh-vision-bridge acts as an intelligent intermediary inside the Cordis runtime:

  • Server-Side Modality Bridge (v0.5.3+): Decorates ctx.llm.resolveModelInfo and ctx.llm.listModels so the session controller accepts image attachments on all models when bridging is active.
  • Automatic Image Rewrite (agent/pre-step & llm/stream): Automatically intercepts image blocks, routes them to a configured vision model (a catalog provider/model, or a local OpenAI-compatible endpoint), receives a descriptive synthesis, and rewrites the image block into text context [The user attached an image. Description: ...] before handing it to the text-only chat model.
  • Native Passthrough: Automatically detects models that natively support vision and allows images to pass directly without unnecessary rewriting.
  • Rich Visual Tool Suite: Exposes ~40 specialized tools for on-demand OCR (incl. local Tesseract), visual question answering, bounding-box grounding, document/table/formula extraction, QR/barcode reading, UI-flow reconstruction, multi-model consensus and more. vision_consensus is opt-in via the consensusEnabled setting.

🏗️ Architecture

graph LR
    User["User attaches Image in Web UI"] --> Gateway["DSH Session Controller"]
    Gateway --> BridgeCheck{"Modality Bridge (v0.5.3)"}
    BridgeCheck -->|"Augments inputModalities"| SessionAllowed["Prompt Accepted"]
    SessionAllowed --> Hook["agent/pre-step Hook"]
    
    Hook --> CheckNative{"Does chat model support vision natively?"}
    CheckNative -->|"Yes (Native Passthrough)"| NativeLLM["Send Raw Image to Chat LLM"]
    CheckNative -->|"No (Text-Only)"| VisionRouter["Vision Bridge Channels"]
    
    VisionRouter --> VisionModel["Dedicated Vision Model\n(DSH / OpenAI / Ollama / Webhook)"]
    VisionModel --> Description["Generated Text Description + OCR"]
    Description --> Rewrite["Substitute Image with Text Marker"]
    Rewrite --> ChatModel["Send Enriched Text to Chat LLM"]
    ChatModel --> Answer["Assistant Response in Chat"]

✨ Key Features

1. Processing Modes

  • hybrid (default): Automatically describes attached images in chat turns while keeping all ~40 explicit vision tools available for follow-up reasoning.
  • llm: Pure auto-rewrite mode — images are transparently converted to text context; tools remain callable.
  • tools: Auto-rewrite disabled — the chat model is expected to explicitly invoke describe_image or OCR tools when required.

2. Multi-Channel Endpoint Routing & Fallback

Chain multiple vision backends with automatic failover, parallel racing, and circuit breaker:

  • dsh-catalog: Auto-detect or select any vision-capable model already registered in DSH.
  • openai-compatible: Standard OpenAI-compatible vision endpoints (vLLM, SGLang, OpenRouter, etc.).
  • ollama: Auto-discovery and local inference through any vision-capable model the local Ollama instance exposes.
  • webhook / custom: External HTTP or JSON-RPC vision endpoints.

3. High-Performance LRU Description Cache

Caches vision responses by hash(bytes + prompt + model + mode) to eliminate redundant vision API calls and save token quota on repeated questions about the same image.

4. Comprehensive Visual Tool Inventory (~40 Tools)

Tool CategoryToolsDescription
Coredescribe_image, read_image, inspect_imageGeneral image analysis by attachment ID, file path, or URL.
Geometry & Detectionvision_ground, vision_crop, vision_detect, vision_compare, vision_presentBounding box coordinates (0–1000 scale), object inventory, multi-image comparison.
OCR & Textvision_ocr, vision_ocr_local, vision_long_ocr, vision_trace, vision_colors, vision_extract_foregroundTranscription, local Tesseract OCR (offline), long screenshot stitching, SVG tracing, color palettes.
Structured & UIvision_describe_structured, vision_vqa, vision_ui_layout, vision_translate_imageJSON breakdown ({summary, ocr, layout, entities}), short VQA, UI section analysis.
Pixel & Diagnosticsvision_pixel_diff, vision_quality_checkSemantic visual diff, quality scoring (blur/lighting).
Documents & Intelligencevision_extract_formula, vision_extract_table, vision_scan_barcode, vision_extract_structured, vision_audit_accessibilityFormula (LaTeX), table (Markdown/HTML), QR/barcode scan, JSON-schema extraction, WCAG accessibility audit.
Scenarios, Consensus & Memoryvision_ui_flow, vision_consensus, vision_memory_searchUser-journey graph (Mermaid), multi-model consensus, semantic search over remembered images.
Attachments (v0.5.33)vision_attach_pages, vision_attach_frames, vision_attach_imagesPublish PDF pages, video frames and local/remote images as conversation attachments so native-vision chat models read the pixels themselves.

📦 Installation

dsh plugin --profile web add @goodandready/dsh-vision-bridge

After installation, restart the DSH Web UI. The configuration card is available under Settings → Plugins → vision-bridge.


⚙️ Configuration (settings.yaml)

dsh-vision-bridge:
  # Operation mode: 'hybrid' | 'llm' | 'tools'
  mode: hybrid
  
  # Auto-detect vision model or specify provider/model explicitly
  visionProvider: ""
  visionModel: ""
  
  # Allow vision models to receive images natively without rewriting ('prefer' | 'never' | 'always')
  nativePassthrough: prefer
  hideRedundantTools: true   # hide bridge tools when the chat model already sees images
  attachMaxItems: 8          # pages/frames/files one attach call may publish
  
  # Enable LRU description cache
  cacheEnabled: true
  cacheMaxEntries: 200
  
  # Request timeout in milliseconds
  timeoutMs: 120000
  
  # Multi-channel routing configuration
  channels: []
  channelFallback: sequential # 'sequential' | 'parallel-race'

Parameter Reference

ParameterTypeDefaultDescription
modestring"hybrid"Processing mode (hybrid, llm, tools).
visionProviderstring""ID of the vision provider (empty = auto-detect).
visionModelstring""ID of the vision model (empty = auto-detect).
nativePassthroughstring"prefer"Behavior for native vision models (prefer, never, always).
hideRedundantToolsbooleantrueWhen the chat model supports images natively, hide the compensation tools from that agent and keep only the extra instruments.
attachMaxItemsnumber8Maximum images one vision_attach_* call publishes (PDF pages, video frames, files). Hard ceiling 32.
cacheEnabledbooleantrueEnables LRU caching for descriptions.
cacheMaxEntriesnumber200Maximum number of cached items in memory.
timeoutMsnumber120000Execution timeout in milliseconds.
channelFallbackstring"sequential"Channel routing (sequential, parallel-race); ordering via channelOrderMode.

15. 🛡️ Enterprise Security & Settings Governance (v0.5.27)

  • Masked Secrets: GET /channels automatically masks sensitive provider API keys (sk-p...7890 or ********) and reports hasApiKey: true to prevent secret leakage in browser DevTools/XHR. POST /channels preserves existing keys when a mask is submitted.
  • CSRF & Origin Guard: All mutating and cost-incurring endpoints (/config, /channels, /upload-pdf, /test, /bench, /batch, DELETE /journal, DELETE /cache, and /doctor?probe=1) validate sec-fetch-site !== 'cross-site' via isTrustedSettingsRequest, rejecting cross-origin attacks with 403 Forbidden. The default GET /doctor report is static (no channel probes).
  • SSRF fetch policy: model-supplied image URLs are fetched only through safeFetch — non-http(s) schemes, localhost names and private/loopback/link-local hosts (incl. IPv4-mapped IPv6 and NAT64) are refused on every redirect hop; response bodies are capped. Use allowedUrlHosts to explicitly re-allow an internal endpoint.
  • apiKeyRef: channels may reference a credential-service entry or environment variable by NAME instead of storing a plaintext apiKey in settings; keys resolve at call time. GET /channels keeps masking values; key preservation on save matches channels by identity, not position.
  • Privacy & honesty: maskPII, stripEXIF, auditLog and consensusEnabled settings are wired end-to-end; the previously decorative face-blur / NSFW / tiling toggles and no-op presets were removed.
  • Native settingsScope Integration: Settings UI binds directly to ctx.settingsScope (namespace: 'dsh-vision-bridge'), supporting reactive snapshot listeners and kernel state synchronization.

📝 Changed in v0.5.30

Security and honesty release. Additive summary of what changed for users:

  • SSRF fetch policy: model-supplied image URLs (describe_image urls, inspect_image, and the headless-chrome tools) are now fetched only through a policy layer — non-http(s) schemes, localhost names and private/loopback/link-local hosts (incl. IPv4-mapped IPv6 and NAT64) are refused on every redirect hop; bodies are capped. New setting allowedUrlHosts (exact-hostname allowlist) deliberately re-allows an internal endpoint.
  • apiKeyRef: channels can reference a credential-service entry or environment variable BY NAME — plaintext API keys are no longer required in settings.yaml. Existing inline apiKey values keep working; masked-key preservation on save now matches channels by identity, not position.
  • Route guards: POST /bench, POST /batch, DELETE /batch/:id, DELETE /journal, DELETE /cache now require same-origin like the other mutating routes. GET /doctor is static by default; channel probes run only with ?probe=1 (same-origin required).
  • Settings honesty: maskPII, stripEXIF, auditLog and consensusEnabled are now wired end-to-end. The previously decorative blurFaces, nsfwFilter, tileLargeImages/tileThreshold toggles and the no-op Local/Cloud/LM Studio presets were removed. Changed in v0.5.30: if you relied on them, note they never had an effect.
  • English source language: all user-facing strings are English; the bundled Russian dictionary was removed — the translation plugin supplies Russian at runtime.
  • Core split: the pure kernel (config schema + helpers) moved to lib/vision-core.js; lib/index.js re-exports it — no API changes. describe_image lost a v0.5.13 regression that returned an empty description; vision_annotate works again; pHash caching no longer mixes up similar images.

📝 Changed in v0.5.31

Maintenance release — no user-facing behavior changes.

  • Internal structure: tool registrations moved to lib/tools/* domain modules (core / grounding / ocr / document / analysis / media); lib/index.js re-exports everything as before. Smaller files, explicit dependencies between the host and the tool domains.
  • Settings card hardening: locale registration is fault-tolerant, the redundant sidebar fallback was removed, plugin service access goes through a safe ctx.get wrapper, and /config accepts the expanded field set (cacheMaxEntries, channelFallback).

📝 Changed in v0.5.32

Stability and architecture release.

  • Internal structure: all ~44 tool registrations moved to lib/tools/* domain modules (core / grounding / ocr / document / analysis / media) with explicit dependencies; lib/index.js keeps the host wiring only.
  • Stability fixes: finished batch records are released after a 10-minute poll window (memory growth fixed); /upload-pdf rejects payloads above the new maxPdfBytes setting (20 MiB default) instead of buffering arbitrary bodies; vision_memory_search scores each attachment against its own description (previously all attachments matched identically); dead host code removed; journal labels are consistent between the legacy and channels paths.
  • Settings: new maxPdfBytes setting (upload hard cap).

📝 Changed in v0.5.33

Vision-first release: the bridge now serves chat models that see images natively, not only text-only ones.

  • Attach domain (vision_attach_pages, vision_attach_frames, vision_attach_images): PDF pages, sampled video frames and local/directory/URL images are published as conversation attachments, so a native-vision chat model looks at the pixels itself instead of paying for a second vision call. Every image is compressed by the existing imageMaxWidth/imageMaxHeight/imageQuality settings and bounded by maxImageBytes; URL sources go through the same SSRF policy as the rest, and local paths are restricted by allowedImageDirs.
  • Model-aware tool exposure: on a route whose chat model already accepts images, the compensation tools of the bridge are hidden from that agent (the model does not need them) and only the extra instruments remain; on a text-only route the attach tools are hidden instead, because that model cannot see an attachment. The behaviour is controlled by the new hideRedundantTools setting (on by default).
  • New settings in the plugin card: an Attachments group exposes attachMaxItems (whole number 1–32, default 8 — how many images one attach call publishes) and hideRedundantTools. Both are validated on save; a value outside the range is rejected with a clear message. The image dimension fields are now also written to the live settings snapshot, not only to the route.
  • Batch API: DELETE /batch/:id releases a finished batch immediately, under the same same-origin guard as start/cancel. The batch record's TTL timer no longer keeps a short-lived process alive, which cut the test suite from 10 minutes to ~3.5 seconds.
  • Tools mode: images attached in chat are now indexed before the sanitisation gate, so tools that take an attachment id (and the read_image alias) work in tools mode as well; previously the ids were unavailable there.
  • Fixes: one unreadable source no longer aborts vision_attach_images — it is reported as Skipped N: <name>: <reason> while the readable sources still attach, and a fetch-policy refusal stays a hard error; the PDF text layer actually appears now (pdftotext was never detected, so the layer silently never shipped); a page range or frame count cut by the cap is reported through truncated and the note.
  • Internal: CI installs poppler without sudo, runs once per commit, serializes per ref and installs ffmpeg for the frame tests; the repository gained the pull-request template and ignores .

📝 Changed in v0.6.0

Major stability, reliability, and lifecycle release.

  • In-App Auto-Updater: added dedicated one-click update endpoint (/api/dsh-vision-bridge/update) and Settings card controls (UpdaterBlock) in settings.plugin.item. Features real-time npm version check, progress state, and restart guidance, protected by loopback and CSRF origin verification gates.
  • Error Resilience & Catch Elimination: audited and replaced all 76 unannotated empty catch blocks across core runtime and tools. Added bestEffort(label, fn, fallback) helper for non-fatal side effects, explicit warning propagation in image processing (smartOptimizeImage returns preprocessed: boolean and warnings: string[]), and transparent fallback reporting.
  • Identity & Scope Integrity: strictly aligned scoped package identity @goodandready/dsh-vision-bridge across manifest, Cordis patch, client module loader, and internal runtime metadata.
  • Hardened Settings Security: upgraded settings mutation route to a strict fail-closed validator (isTrustedSettingsRequest) enforcing loopback origin, CSRF header checks, bearer tokens, and rejecting suspicious remote headers.
  • Native Theme Compliance: client diagnostics and card styles fully migrated to DSH CSS design tokens (--dsw-alias-*), eliminating all hardcoded color literals.
  • Sanitized Public Release: integrated plumbing-based release script (publish.sh) and .gitattributes export filters ensuring zero private development artifacts in public distribution.

📝 Changed in v0.6.2

Reliability hardening, error visibility, and degradation warnings release.

  • Real bestEffort Logging: replaced all remaining dummy /* bestEffort ... */ void err catches across core runtime, tools, and routes with active bestEffort(label, fn) logging to console.debug.
  • Degradation Warnings in Tools: all tool modules now declare warnings: { type: 'array', items: { type: 'string' } } in output schemas and surface warnings: string[] on partial JSON parsing errors or engine fallbacks.
  • Image Preprocessing Failure Capture: preprocessing exceptions (deskew, enhance, stripEXIF, compression) are recorded and surfaced to tool callers instead of failing silently.
  • Expanded Verification: added unit test suite test/issue-314-besteffort-warnings.test.js, bringing total test coverage to 329 passing tests across 104 suites.

📄 License

MIT © GooDAnDReaDY

.worktrees/