dsh-local-vision
Give a text-only DeepSeek Harness agent local "eyes". This plugin registers a
model-callable local_vision tool that describes an image using a local
vision model through Ollama. The image never leaves your
machine — no cloud, no uploads, no API keys.
Requirements
-
Ollama running (default http://localhost:11434).
-
At least one vision-capable model pulled, e.g.:
ollama pull qwen2.5vl:3b # small, fast — text / colors / summary
ollama pull gemma4:12b-it-qat # larger — thorough description
Install
# from npm (recommended)
npm install -g @deepseek-ai/dsh
dsh plugin --profile web add dsh-local-vision
# or from this repository (use `file:` so @deepseek-ai/dsh-tools resolves)
dsh plugin --profile web add file:/path/to/deepseek-harness-plugins/dsh-local-vision
Restart the DSH session so the new bundle is composed. The package then appears
in Settings → Plugins → Plugin list and the tool local_vision is available.
Usage
local_vision(image: "/abs/path/screenshot.png", mode: "fast")
local_vision(image: "/abs/path/mockup.png", mode: "detailed", prompt: "Describe the navigation and buttons")
local_vision(image: "screen", mode: "detailed") # capture the current display (macOS screencapture)
Modes
mode | intent | default model |
|---|
fast | transcribe visible text, name dominant colors, one-line summary | qwen2.5vl:3b |
detailed | full factual description: layout, UI elements, all text, colors, state | gemma4:12b-it-qat |
Parameters
| Param | Type | Notes |
|---|
image | string (required) | Absolute path to a PNG/JPEG/WebP/GIF, or the literal screen. |
mode | "fast" | "detailed" | Defaults to fast. |
prompt | string | Optional focused question. Defaults per mode. |
model | string | Optional override of the Ollama model id for this call. |
Configuration
Optional, at $DSH_HOME/local-vision.json (default ~/.dsh/local-vision.json):
{
"host": "http://localhost:11434",
"models": { "fast": "qwen2.5vl:3b", "detailed": "gemma4:12b-it-qat" },
"temperature": 0.1,
"timeoutMs": 300000
}
All keys are optional; defaults are shown above. The plugin sends the image to
Ollama's OpenAI-compatible /v1/chat/completions endpoint as a base64 data URI.
Notes & limitations
- The first call after the model is cold loads it into memory (a 7 GB model takes
a while); subsequent calls are fast.
image: "screen" requires macOS and Screen Recording permission for the terminal.
- The tool calls Ollama directly (not through the DSH sandbox executor).
License
MIT