DeepSeek Harness Plugin Hub

Publish and manage complete Harness Profiles. Discover Plugins for your next setup.

Explore

PluginsPresetsDocsNews

Community

Publish a pluginContactReport an issue

Resources

Plugin Hub on GitHubDeepSeek HarnessSystem statusPrivacy notice
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

Independent and unofficial. Not affiliated with, authorized by, or endorsed by DeepSeek.

Semantic Search — DSH Plugin for DeepSeek Harness
DeepSeek Harness Plugin Hub
ProfilesPluginsCategoriesNewsDocsSign inManage Profiles
ProfilesPluginsCategoriesNewsDocsSign in
← Plugins
S

dsh-semantic-search

Semantic Search

A DeepSeek Harness (dsh) plugin for local semantic code search. Builds a fragment-level index of a workspace (symbol-aware chunking for many languages), computes local embeddings (lightweight lexical TF-IDF feature vectors by default, or a configurable OpenAI-compatible embedding endpoint), and answ

The plugin will be installed here. Keep web if you are unsure.

npx -y @deepseek-ai/dsh plugin --profile web add github:JohnXu22786/semantic-search#601c3f581506ee202591b2202e18ccd8abad380a
READMECompatibilityVersions

Description

A DeepSeek Harness (dsh) plugin for local semantic code search. Builds a fragment-level index of a workspace (symbol-aware chunking for many languages), computes local embeddings (lightweight lexical TF-IDF feature vectors by default, or a configurable OpenAI-compatible embedding endpoint), and answers natural-language queries with hybrid retrieval (vector cosine + BM25, fused with reciprocal-rank fusion). Ships dsh tools sema_search / sema_reindex / sema_stats plus a stand-alone CLI.

Compatibility and provenance

Semantic Search is published as dsh-semantic-search and currently resolves to version 0.1.0. The Hub verifies its manifest and preserves the exact installation source for reproducible installs.

DSH compatibility
*
Runtime surfaces
any
Release source
github
Registry updated
8/20/2026

Versions

0.1.0stable
8/20/2026

Related plugins

Loading related plugins…

Latest
0.1.0
DSH
*
HMR
Process restart
Tree shaking
Safe tree shaking not declared
Unpacked size
Unavailable
Files
Unavailable
Surface
any
License
MIT
Source
github
GitHub
★ 2
Weekly downloads
0
Last push
9/12/2026
View source ↗
README badge

Click the badge to copy Markdown for your README.

Do you maintain this Plugin?Claim benefit · Priority security scan

Verify the GitHub repository declared in package.json to manage this listing. After you claim it, Hub will prioritize a security scan of the current version and publish the result when it passes.

Claim this Plugin →
Report an issue

Related plugins

More verified plugins in developer-tools.

Web App@deepseek-ai/dsh-web-appThe dsh browser-surface bundle: the web patch layer over dsh-base plus the runtime glue plugin (frontend dist serving, web-surface prompt, bash runtime variables, URL line)Sdk Minimal@deepseek-ai/dsh-sdk-minimalThe standalone minimal SDK profile bundle: JSON-RPC, one DeepSeek adapter, persistent shell, and JSONL sessionsSdk App@deepseek-ai/dsh-sdk-appThe dsh SDK profile bundle: stdio JSON-RPC serving and process lifecycle over dsh-baseSubagent Codex@deepseek-ai/dsh-subagent-codexOne-shot Codex subagent provider over the official app-server protocol

README

dsh-semantic-search

Local semantic code search for DeepSeek Harness (dsh) — and any Node script.

中文文档:README.zh.md

sema builds a fragment-level index of a workspace: source is tokenized by a language-aware tokenizer (camel/snake/kebab splitting, CJK n-grams), chunked with symbol-aware boundaries (functions/classes stay intact), and embedded into fixed dimension vectors — locally and dependency-free by default (feature-hashed lexical TF-IDF), or via any OpenAI-compatible embedding endpoint. Queries run a hybrid retrieval (vector cosine + BM25, fused with reciprocal-rank fusion) over the index, so meaning-based search works even when exact terms don't match.

Ships as a dsh plugin bundle — the sema_search, sema_reindex and sema_stats tools on the harness tool registry — plus a standalone sema CLI.


Highlights

  • Works offline by default — the built-in lexical provider needs no network, no model download, no API key. Use it as a fast BM25-plus code search.
  • Symbol-aware chunking — boundaries from a conservative per-language table (16 languages) keep functions/classes intact; a missed boundary degrades to a plain chunk rather than breaking.
  • CJK-aware tokenizer — n-gram tokenization (default bigrams) aligns Chinese queries and documents without a segmentation library; full-width punctuation is folded, not treated as a hard break.
  • Hybrid retrieval with RRF — vector cosine + BM25 channels are fused by reciprocal-rank fusion, so a document found by one channel still ranks.
  • Graceful degradation — if a configured remote provider is unreachable, the index falls back to the local lexical provider (configurable via allowFallback).
  • Incremental refresh + file watching — sema_reindex diffs by size+mtime, and an optional watcher keeps the index live.
  • Persistence — the index is saved to <root>/.sema atomically (JSON metadata
    • binary vectors), with staleness detection when the provider/dimension changes.
  • Deterministic — the same workspace and options produce the same index and the same ranked answers.

Supported languages

TypeScript, JavaScript, Python, Go, Rust, Java, Kotlin, Scala, C, C++, C#, Objective-C, Ruby, PHP, Swift, Bash, plus common data/markup formats (JSON, YAML, TOML, Markdown, HTML, XML...).


Installation

As a dsh bundle

The package declares "dsh": { "bundle": { "patch": "./cordis.patch.yml" } }. The patch inserts one plugin row that mounts the bundle and registers sema_search / sema_reindex / sema_stats on ctx.tools.

# from npm (name reserved; publish pending access setup)
npm install -g dsh-semantic-search

# straight from this repository
dsh plugin --profile demo add github:JohnXu22786/semantic-search

# or from a local checkout
dsh plugin --profile demo add /path/to/semantic-search

As a standalone CLI

npm install -g dsh-semantic-search   # or: npm run build && node bin/sema.mjs
sema --help

CLI usage

sema index                build the full index from the workspace
sema reindex [--full]     incremental refresh (or full rebuild with --full)
sema search <query...>    hybrid vector + BM25 search, prints top hits
sema stats [--json]       index health, provider, and sizing numbers

Global options:

--root <dir>          workspace root (default: current directory)
--data-dir <dir>      index storage directory (default: <root>/.sema)
--provider <kind>     embedding provider: lexical | openai (default: lexical)
--dim <n>             embedding dimension (lexical default: 4096; openai 0 = auto)
--base-url <url>      OpenAI-compatible embeddings endpoint base URL
--model <name>        embeddings model name (openai only)
--api-key <key>       API key (openai only; env: SEMA_EMBEDDING_API_KEY)
--top <n>             hits to print for search (default: 20)
--json                machine-readable output where supported
--help                show this help

CLI exit codes

  • 0 — success (including a search with zero hits and a --version/--help call).
  • 1 — a runtime failure (config error, build/index/search error).
  • 2 — a usage error: unknown command, unknown flag, or a missing query.

Configuration

The plugin is configurable through the bundle row's config (see cordis.patch.yml for an example), the CLI flags above, or defaults in code:

OptionDefaultMeaning
rootcwdworkspace root to index
dataDir.semaindex storage directory
provider.kindlexicallexical (offline) or openai
provider.dimension0 (lexical: 4096)embedding dimension; 0 = auto-infer from the endpoint
provider.baseUrlhttps://api.openai.com/v1OpenAI-compatible endpoint root
provider.modeltext-embedding-3-smallembedding model name
provider.apiKeyEnvSEMA_EMBEDDING_API_KEYenv var holding the API key
allowFallbacktruefall back to the lexical provider when a remote one fails
include / ignoredefaultsglob sets of files to index / skip
maxLinesPerChunk80hard chunk size upper bound
nGram2CJK n-gram size (1 disables n-gramming)
topK20hits returned by default
rrfK60RRF fusion constant
vectorK300candidates per channel before fusion
autosavetruepersist the index after builds
autoIndextruebuild lazily on first search
watchtruewatch the workspace for changes

Development

npm ci
npm test          # build + run the node:test suite (76 tests)
npm run typecheck
npm run build     # tsc -> lib/

License

MIT — see LICENSE. © 2026 dsh-semantic-search contributors.