DeepSeek Harness Plugin Hub

Publish and manage complete Harness Profiles. Discover Plugins for your next setup.

Explore

PluginsPresetsDocsNews

Community

Publish a pluginContactReport an issue

Resources

Plugin Hub on GitHubDeepSeek HarnessSystem statusPrivacy notice
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

Independent and unofficial. Not affiliated with, authorized by, or endorsed by DeepSeek.

Video Lens — DSH Plugin for DeepSeek Harness
DeepSeek Harness Plugin Hub
ProfilesPluginsCategoriesNewsDocsSign inManage Profiles
ProfilesPluginsCategoriesNewsDocsSign in
← Plugins
V

dsh-video-lens

Video Lens

Give text-only DeepSeek Harness agents video understanding: scene-aware frame sampling + VLM + optional ASR transcript fused into timeline evidence. / 给纯文本模型的视频理解插件(场景感知抽帧 + VLM + 可选语音转录)

The plugin will be installed here. Keep web if you are unsure.

npx -y @deepseek-ai/dsh plugin --profile web add github:dundunhan/dsh-video-lens#cc8e9a0cfae73c904d463ccbc7610967ef6547f1
READMECompatibilityVersions

Compatibility and provenance

Video Lens is published as dsh-video-lens and currently resolves to version 0.3.2. The Hub verifies its manifest and preserves the exact installation source for reproducible installs.

DSH compatibility
*
Runtime surfaces
any
Release source
github
Registry updated
9/18/2026

Versions

0.3.2stable
9/18/2026

Related plugins

Loading related plugins…

Latest
0.3.2
DSH
*
HMR
Process restart
Tree shaking
Safe tree shaking not declared
Unpacked size
Unavailable
Files
Unavailable
Surface
any
License
MIT
Source
github
GitHub
★ 113
Weekly downloads
0
Last push
9/19/2026
View source ↗
README badge

Click the badge to copy Markdown for your README.

Do you maintain this Plugin?Claim benefit · Priority security scan

Verify the GitHub repository declared in package.json to manage this listing. After you claim it, Hub will prioritize a security scan of the current version and publish the result when it passes.

Claim this Plugin →
Report an issue

Related plugins

More verified plugins in vision-media.

Tool Describe Image@linxin666/dsh-tool-describe-imageModel-facing describe_image tool for the dsh web GUI: gives a text-only model image understanding by asking a vision-language model at an OpenAI-compatible endpoint to describe one image (local path, http(s) URL, or attachment reference). Hot-pluggable — Modlens@liustack/modlensPlug-in vision for text-only LLMs, powered by the free Antigravity CLIDeepseek Ivideodeepseek-ivideoiPolloWork HyperFrames Video Studio and 27 editable video templates as a native DeepSeek Harness conversation view.Imagegen@dickpy/dsh-imagegenAI image generation plugin for the dsh web GUI: text-to-image and image-to-image through configurable provider channels (gpt-image-2 / grok-imagine-image / nanobanana series / seedream-5.0-pro / dall-e-3, with native xAI Grok Imagine, Google Nano Banana a

README

dsh-video-lens

Video understanding for DeepSeek Harness — give text-only agents eyes and ears on video.

A DeepSeek Harness (DSH) plugin that lets text-only LLM agents understand local video files. It provides two tools:

ToolWhat it does
video_probeCheap, instant metadata via ffprobe: container, duration, resolution, fps, codecs, audio tracks, subtitles.
video_analyzeContent understanding: scene-change-aware frame sampling (ffmpeg scdet), optional ASR transcript (speech with timestamps), fused with any OpenAI-compatible vision model into structured evidence JSON.
video_askTime-anchored Q&A: parses explicit time references ("at 3:20", "第2分钟") or locates relevant speech via transcript keyword matching, re-samples frames from the matched windows, and answers with grounded evidence (answer + confidence + supporting timestamps).

v0.3.2. The plugin never locks you into a provider: vision and ASR are both OpenAI-compatible endpoints configured via baseUrl + model + key env var.

How it works

video file ──► video_probe ──► ffprobe ──► compact metadata JSON
           └─► video_analyze ──► scdet scene detection ──► shot boundaries
                                 ├─► ffmpeg frame sampling (one representative frame per shot, capped)
                                 ├─► ffmpeg audio extract ──► ASR transcript (timestamped)   [optional]
                                 └─► OpenAI-compatible vision API ──► evidence JSON
  • Scene changes are detected with ffmpeg's scdet filter (ffmpeg ≥ 6.0). Videos without detectable cuts fall back to uniform midpoint sampling.
  • ASR is strictly additive: if asrApiKeyEnv is unset or the provider fails, the visual analysis still completes and transcript is null.
  • All media work is delegated to ffmpeg/ffprobe on PATH — no native decoding in the agent.

Install

Prerequisites: Node.js ≥ 20, ffmpeg ≥ 6.0 (recommended) with ffprobe on PATH (brew install ffmpeg / apt install ffmpeg).

Option A — npm (recommended)

# in your DSH profile directory (the one containing package.json)
pnpm add dsh-video-lens

Option B — from source (development)

Clone the repo, then mount it into your DSH profile via a local link:

git clone https://github.com/dundunhan/dsh-video-lens.git

Either way, register the bundle in your profile's package.json — this exact block is the full profile configuration:

{
  "dependencies": {
    "dsh-video-lens": "^0.3"
  },
  "dsh": {
    "profile": {
      "bundles": [
        "@deepseek-ai/dsh-base",
        "@deepseek-ai/dsh-web-app",
        "dsh-video-lens"
      ]
    }
  }
}

Then export the keys and restart the profile:

export VIDEO_LENS_API_KEY=sk-...        # vision
export VIDEO_LENS_ASR_KEY=sk-...        # optional, ASR

Do not install the host runtime yourself. @deepseek-ai/dsh-tools is declared as an optional peer: the plugin always uses the dsh-tools that already ships with your DSH installation / DSH Desktop. Adding it to your profile as a dependency — or pinning one exact -rc version, which is what 0.3.1 did — installs a second, older runtime next to the host's, makes the Loader entry fail to import, and takes the whole plugin tree (and the app) down with it.

Configuration

All options are DSH config values:

KeyDefaultMeaning
visionBaseUrlhttps://api.siliconflow.cn/v1Vision endpoint (OpenAI-compatible)
visionModelQwen/Qwen3-VL-8B-InstructVision model name
visionApiKeyEnvVIDEO_LENS_API_KEYEnv var holding the vision key
asrBaseUrlhttps://api.siliconflow.cn/v1ASR endpoint (OpenAI-compatible /audio/transcriptions)
asrModelFunAudioLLM/SenseVoiceSmallASR model name
asrApiKeyEnvVIDEO_LENS_ASR_KEYEnv var holding the ASR key
maxFrames12Frame budget cap (1–max); actual count is duration-adaptive (~1 frame per 30s, denser for short videos)
frameMaxWidth768Max frame width; keeps payloads small
frameQuality4JPEG quality (ffmpeg -q:v)
sceneThreshold10scdet threshold (0–100); higher = fewer cuts
askPaddingSec2video_ask window padding around matched transcript segments
vlmMaxTokens1500Vision model max output tokens
vlmTimeoutMs90000Vision call timeout
asrTimeoutMs120000ASR call timeout

Usage

Ask the agent:

"What's in /tmp/demo.mp4?"

The agent calls video_probe first, then video_analyze. Evidence includes:

{
  "metadata": { "container": "mov,mp4,m4a,3gp,3g2,mj2", "durationSec": 268.4, "...": "..." },
  "shots": [{ "timeSec": 12.3, "score": 45.2 }],
  "framesSampled": [{ "timestampSec": 5.5, "jpegBytes": 12345 }],
  "transcript": {
    "text": "…",
    "segments": [{ "start": 0.0, "end": 2.4, "text": "…" }],
    "language": "zh"
  },
  "visionModel": "Qwen/Qwen3-VL-8B-Instruct",
  "analysis": { "overall_summary": "…", "timeline": [{"timestamp_sec": 5.5, "description": "…"}], "on_screen_text": "…", "visual_style": "…", "notable_moments": "…" }
}

Permissions & security

Read this before using or redistributing. DSH plugins run in the host process as trusted code and there is no official plugin review — self-review is on the author. See SECURITY.md.

What this plugin does

  • Reads: any local file path the agent passes to its tools (via ffprobe/ffmpeg).
  • Executes: ffprobe and ffmpeg from PATH (never a shell — argv arrays only).
  • Network: one outbound call per video_analyze to the configured visionBaseUrl (frames + vision key), and optionally one to asrBaseUrl (audio + ASR key).
  • Does not: execute shells, eval code, phone home, auto-update, or read files on its own.

Operator responsibilities

  • Keys are only as safe as the endpoints they are sent to — configure only endpoints you trust.
  • The real access boundary is the DSH host sandbox; the plugin's readability check is a UX guard, not a security boundary.
  • Payload sizes are bounded: maxFrames × ~100–300 KB (768px JPEG) per analysis call.

Compatibility

  • Tested with DSH profile bundles @deepseek-ai/dsh-base + @deepseek-ai/dsh-web-app.
  • Host runtime is not pinned: @deepseek-ai/dsh-tools is an optional peer resolved from the host installation, so the plugin follows the core it is loaded by (verified against core 0.1.0-rc.7 and 0.1.5-rc.2, the upstream version DSH Desktop 2.0.5 pins).
  • Node ≥ 20 (uses AbortSignal.any / built-in fetch / FormData).
  • ffmpeg ≥ 6.0 for scdet; older versions degrade to uniform sampling.
  • macOS verified. Windows: the 0.3.1 boot failure reported on Windows was not platform-specific — it was the pinned old dsh-tools runtime (see Troubleshooting); the code paths themselves are OS-neutral (ffmpeg/ffprobe are spawned via argv, no shell).

Troubleshooting

dsh-plugin-desktop: plugin tree failed to load: failed to apply loader entry include (cordis:include): AggregateError: loader entries failed to apply — the client no longer starts.

0.3.1 hit this on DSH Desktop 2.0.5. The full error underneath is an import failure of the plugin (or of the host tools entry):

failed to import loader entry video-lens (dsh-video-lens): The requested module '@deepseek-ai/dsh-llm' does not provide an export named 'CallId'
  [cause]: profiles/<name>/node_modules/@deepseek-ai/dsh-tools/lib/index.js:4

Cause: 0.3.1 pinned @deepseek-ai/dsh-tools@0.1.0-rc.7, so the profile got a second, older dsh-tools while the host ran a newer core (0.1.5-rc.2). Any failing Loader entry fails the whole tree, so the app cannot boot until the plugin is removed.

Recovery (0.3.1 installed and the app will not start):

  1. Use the client's Recovery page to return to the last healthy profile, or remove the plugin from the profile: dsh plugin --profile <name> remove dsh-video-lens (Desktop: run that in its terminal).
  2. Install dsh-video-lens@^0.3.2, where the host runtime is an optional peer and nothing is installed into the profile.

Uninstall

  1. Remove dsh-video-lens from dsh.profile.bundles in your profile package.json.
  2. Remove the dependency: pnpm remove dsh-video-lens (npm install) — or delete the link: entry if you installed from source — then reinstall the profile.

Roadmap

  • v1.0: frame caching by file hash, evaluation table in README (5 video types × metrics), publish to npm (in progress).
  • Beyond: native video-input models as an optional fast path when the configured VLM supports them.

License

MIT — see LICENSE.