dsh-project-memory
如果这个插件帮你省下 1 小时 Debug 时间,请点个 Star。
English | 简体中文
A persistent project development memory for DeepSeek Harness (dsh) agents. Built specifically for project development, natively integrated with dsh's task system: task lists and files read during a session are automatically persisted as cross-session task records, with tasks ↔ files linked — workflows can be switched and resumed, no need to re-scope the whole project, solving context loss. Documents (PDF/Markdown/txt) and code symbols are stored separately per workspace; documents are automatically cross-linked to the code symbols they mention. Experience notes (problem → solution) are automatically deduplicated, preventing repeated mistakes. All data is stored per project on disk, survives session compaction and handover; recalls include path:line citations for source verification. Only one dependency, no vector DB, no native builds.
The plugin keeps a compact project memory on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project. Tasks and experience persist across session compactions and handovers.
alt text
The workflow panel is collapsible, automatically adapts to dsh and theme plugin styles, and offers four card style options to switch between.
alt text
Features
- TaskBridge: cross-session development tasks — the plugin watches each session's live todo list (
todo_write events) and file reads (tool/call): progress snapshots (steps) and touched files sync into durable per-project task entities. An unbound session that writes a todo auto-creates a task. Associated files are kept in recency-weighted order (written/edited first; a read never outranks a written file) so a resumed session sees at a glance where to look. New sessions continue by list_tasks → select_task (bind / rename / unarchive); query_memory gains type: 'task' and appends a task-count hint to type: 'all' results. The user-side /tasks command shows the task stack, step progress, involved files, and the current session binding. A task named by the model via select_task(title=…) keeps that title; one auto-created by the first todo_write is titled from its first list entry (≤48 chars), falling back to the first human message (the part after the last colon), then Untitled Task. Sessions spawned as subagents are excluded from auto-creation (origin: 'subagent' / delegationDepth > 0); merging delegated work back into a task is deliberately unbuilt — see §11 in Design tradeoffs. Capacity is project-size adaptive (fileCount/20, clamped 5–100). Storage: .dsh-project-memory/tasks.json + binding.json. Auto-sync requires a dsh build with session events + todo_write (verified on 0.1.2-alpha.x, re-verified against the 0.1.5-rc.1 host surface); on older hosts the task tools still work as a plain record list.
- Task Panel (v0.4.2+): Floating task panel in dsh web — built on the real dsh web 0.1.5-rc.1 client plugin contract (cordis inject + apply, registered into host
shell.overlay slot). Draggable cards show steps/files (click to copy path); collapse to a draggable mini-bar; hide completely (summon with /task / /tasks). Render errors have error boundaries — panel crash no longer takes down the host.
- Task Panel Behavior —
- Default hidden: panel does not show on dsh web startup
- Explicit summon: type
/tasks or /task (list form) to open; model calls show_task_panel tool to open
- Session switch: only syncs data in background, does not auto-open panel
Performance
Synthetic Benchmark (Node 24.19, WSL2 on 20 vCPU, Linux file system)
| Scenario | Scale | Measured |
|---|
| Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
| Cold load | 5,000 files | 40 ms |
| Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
| query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
| query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
| Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
| Cold load | 10,000 files | 90 ms |
| Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on WSL2 / 20 vCPU / Linux file system, measured 2026-09-14. Reproduce with npm run bench:synthetic -- 5000 (harness: scripts/bench-synthetic.mjs). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + precomputed searchText; the first query after a write rebuilds IDF (106 ms at 40k entries, 57 ms at 20k, 12 ms at 4k), subsequent queries hit the cache.
Real Project Storage
| Project | Files | Entries | Store Size | Per Entry |
|---|
| Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
| Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
Reproduce it on your own project
Rather than asking you to trust the numbers above, the measurement itself ships with the repository and with the published npm package (scripts/ is part of the tarball). It needs no dsh instance, no network and no model calls, and it never touches your project's own store — results go to a temp directory and are removed when it finishes:
npm run bench -- /path/to/your/project
# or, with options:
node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size and bytes per entry. Example — our internal Vue project (289 files / 2,141 entries, Node 24, 20 CPU, Linux):
cold index 253 ms (read+hash 9 ms · extract 229 ms · commit 13 ms) ← 2nd, warm-cache run
store 1.10 MB · 538 bytes/entry · cold load 4.6 ms
hot query p50 0.80 ms · p95 1.35 ms (2,141 entries)
re-index 1 file p50 0.33 ms
Two caveats we would rather state than hide: read+hash depends on the OS page cache — on that corpus the first run spent 787 ms and the second 253 ms, so say which run you quote — and real projects score slower than the synthetic table above — on a 3,000-file slice of a large TypeScript repository (15,594 entries) hot queries were p50 7.5 ms, because real declaration text is longer than generated stubs. Pass --queries your-queries.json to run the same labeled-set method (hit@5 / hit@10 / MRR) against your own project.
How it works
The design follows four principles:
- Volatility — context is ephemeral; it is lost when a session is compacted.
- Persistence — the memory is stored on disk and survives compaction and new sessions.
- Compactness — the code layer stores one declaration line per symbol, so code-heavy projects stay near 0.5% of the source (8.8 MB of source → 49 KB of index in the example project), and recall replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected
summary, a bounded terms set covering the whole chunk for retrieval, and a precomputed searchText. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): terms ≈ 27.5% of source and the on-disk store ≈ 166% of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
- Verifiability — recalls carry a
path:line citation where applicable, so the agent can confirm details against the source.
Building the memory does not require an upfront scan: files are memorized as the model reads them, so the memory grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the memory stays fresh with minimal ongoing overhead.
The store is per-project and follows the codebase: changed files are re-extracted by content hash, deleted files are removed. Experience notes are retrieval-only, so accumulation does not affect context.
Installation
The plugin relies exclusively on stable public APIs (defineTool, llm.stream, Schema) declared via peerDependencies, ensuring compatibility with future rc/alpha releases without changes.
cd dsh-project-memory && dsh plugin --profile web add . -w
The -w (workspace-root) flag is required: the profile directory is a pnpm workspace root, and pnpm rejects add there without it. From any other directory, the path form works the same: dsh plugin --profile web add /path/to/dsh-project-memory -w.
The plugin is also published on npm as a scoped package:
dsh plugin --profile web add @yolk_vat-y/dsh-project-memory -w
A prebuilt tarball is published with each release, installable without a build step:
dsh plugin --profile web add /path/to/dsh-project-memory.tgz
Each indexed project has its own store at <root>/.dsh-project-memory/. Add it to .gitignore if it should not be committed.
Usage
The tools below are invoked by the agent, not typed by the user. In the chat, just ask naturally — e.g. "index this project" or "what does the auth module do?" — or simply keep working, and the agent calls the matching tool automatically. By default (lazyIndexing) files are indexed the moment the model reads them, so memory fills in while you work. watch_repo keeps explicitly-watched roots fresh in the background; index_repo forces a full backfill of a project (unchanged files are skipped).
| Tool | Purpose |
|---|
index_doc file_path | Index one document (PDF/MD/txt): chunk → deterministic summary + whole-chunk terms → store with path:line. Unchanged files are skipped. |
index_repo root | Index a whole project: docs get deterministic summaries + whole-chunk terms, code files get a zero-token symbol table. Incremental, cleans up deleted files, cross-links docs to symbols. A root that does not exist — including a Windows-style path resolved on Linux/macOS — is rejected before anything is written. |
watch_repo root | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts; a non-existent root, the filesystem root and the shared temp directory are all refused, and roots that disappear are dropped instead of being re-created. |
memory_stats root | Show what the store contains: totals (files / entries / experience notes), last index time, and the per-file list sorted by recency. |
query_memory query | BM25 search over docs + symbols + experience + insights (lessons / decisions / procedures), optionally query-expanded by the LLM. type selects a layer (all / doc / symbol / experience / insight / task). Returns ranked hits with relative scores, sources or insight ids, and doc→symbol references. |
list_tasks | List task records for the project (archived marked). Call first in a new session before continuing work. |
select_task | Bind the session to a task so its todo list and file reads sync into it. Exact taskId, or exact title (multiple matches return candidates; no match creates a new task). Pass title with taskId to rename. Auto-unarchives. |
archive_task | Archive a task (hide from default views, exclude from capacity, stop syncing). select_task restores it. |
show_task_panel | Show the task panel in the UI. Call when the user asks to see the task list or when you want to display the panel. |
/tasks (typed by the user, not the model) | Shows the task stack: title, step progress, involved files, and which task the current session is bound to. |
Design
.dsh-project-memory/
format.json layout marker (v2, sharded)
shards/ one self-describing JSON per indexed source file
({ relPath, record, entries }) — writes touch only dirty shards
experience.json problem → solution notes (retrieval-only)
watch.json watched roots
tasks.json TaskBridge task entities (cross-session)
binding.json current session ↔ task binding
insights.json v0.5 project-scope insights (lessons/decisions/procedures); v0.4 experience notes imported once, non-destructively
Stores created before v0.2.0 (single entries.json / index.json) migrate automatically and idempotently on first load. Within one dsh process, all tool calls share a single in-memory store per project, so hot-path indexing writes only the shard that changed.
- Incremental — content hash per file; only changed files are re-extracted.
- Cross-linking — after indexing, doc summaries are matched against symbol names; matches are attached to the doc entry as
references and surfaced by query_memory.
- Query expansion — when
llmQueryExpansion is on, query_memory asks ctx.llm to rewrite the query into several variants (synonyms, EN/CN, identifier guesses) and merges BM25 scores across variants; when off, queries never touch the LLM. Indexing itself is model-free: keywords are rule-derived (title-weighted top terms), and doc↔symbol links surface English symbol names from Chinese hits.
- Consistency — the fact layer follows the codebase (hash re-extract / remove-on-delete); the experience layer is retrieval-only with supersede and
forget. Store writes are serialized per memory directory; the lock is in-process, so avoid running multiple dsh instances against the same project store concurrently.
Architecture (Task Panel)
TaskPanel (Container)
├── task-data-store (server data, cross-tab sync via BroadcastChannel)
├── task-ui-store (local UI state, localStorage)
├── task-hooks (useTaskDrag, useTaskEdit)
└── TaskComponents (MiniBar, TaskCard — presentational only)
Design tradeoffs
These are deliberate scope choices.
1. Synchronous lock-free transactions over async locks
We do: All writes go through store.commit(fn) — a synchronous in-process transaction. The callback fn performs all validation and mutations; only on success is the result atomically written to disk. The JS event loop guarantees no interleaving. CAS (applyFileUpdate) makes concurrent writes idempotent.
We don't: Async mutexes, file locks, or multi-process coordination.
Why: DSH runs on Cordis, which is single-process by design. Adding locks would complicate the hot path (every remember/forget/index_doc call) for a scenario (multi-process DSH) that would require a breaking ecosystem change. Synchronous transactions keep the hot path at ~2 ms median with zero contention overhead in practice.
2. Watch: compute outside, commit inside
We do: Heavy work (mtime/hash/scan/parse/PDF extraction) runs outside the transaction; a single commit applies all changes atomically. On failure, the snapshot rolls back so the next poll retries automatically.
We don't: Hold a lock during parsing, or use fs.watch events.
Why: PDF extraction and large-file parsing take time — holding a lock would block remember/forget/query_memory. Polling with mtime+content-hash is platform-agnostic (works on network drives, Docker volumes, WSL) and avoids the "double fire / missed events" nightmare of fs.watch.
3. Corrupt files are quarantined, not auto-repaired
We do: On JSON parse failure, the bad file is renamed to *.corrupt, an error is logged, and that file's store starts fresh. The rest of the store remains intact.
We don't: Write-ahead logs, embedded databases (SQLite/LMDB), or automatic partial recovery.
Why: A corrupted shard means one source file has a bad index — quarantining it costs near zero. A WAL or embedded DB adds a heavy dependency, increases binary size, and introduces new failure modes (lock contention, corruption of the WAL itself). The tradeoff: lose one file's index vs. add 500 KB+ of native code.
4. No vector embeddings, no semantic search at query time
We do: BM25 with CJK phrase boost (3+ chars ×1.5 on title/keywords), synonym expansion (bidirectional table), field weighting (title ×5), and experience-layer phrase boost. All at query time, zero LLM calls.
We don't: Vector embeddings, dense retrieval, rerankers, or hybrid search.
Why: Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For the queries this plugin targets, lexical BM25 is already sufficient and measurable: on our benchmark suite (29 queries over a real Vue project) file-level hit@5 is 96.6%, and 28 of the 29 are exact symbol lookups that lexical search answers essentially always. Whole-chunk terms took document-term coverage from 27.3% to 100% while queries that already worked kept their ranking (MRR 0.958 vs 0.955). Those figures come from an internal Vue project with a hand-labeled 29-query set, so they are not reproducible outside it — but the method now ships as scripts/bench.mjs --queries <your-set.json>, so you can run the identical measurement on your own project. The marginal gain from semantic search doesn't justify the 10x complexity/cost increase.
5. Indexing is deterministic and model-free
We do: Derive keywords with a rule (title-weighted top terms) and build a whole-chunk terms set — both deterministic and reproducible. Doc↔symbol links surface English symbol names from Chinese queries, and CJK tokenization keeps cross-language hits working. With llmQueryExpansion: false, queries never touch the LLM.
We don't: Call a model at index time to translate or paraphrase a document, and we don't translate queries at search time.
Why: An index-time model call makes indexing slower, non-deterministic and unverifiable — the same document can index differently on two runs. Query-time translation adds latency and a hard failure mode (a bad translation means zero recall). Rules plus symbol linking cover the common cases, work offline, and keep indexing at zero model calls.
6. Model-facing memory: the agent writes, and no human has to be in the loop
We do: Treat the agent as a first-class writer. remember / save_lesson write any scope at any time (task / project / global) with no human step, and promotion is deterministic and runs inside the ordinary write path: cross-task token-overlap dedupe accumulates sourceTaskIds, then promoteAllTasksToProject / promoteProjectToGlobal move an entry up once its corroboration counts are met (≥2 tasks for project, ≥ globalPromoteTasks — 3 by default — for global). Nothing waits on the task panel: a user who never opens the UI still gets a memory that fills, dedupes and graduates.
We do (labeling): Keep inferred content distinguishable from recorded content. The v0.5 reflection path (opt-in, off by default) is the only writer that infers rather than records: it writes task-scoped drafts stamped draft: true / source: 'reflect', and recall plus silent injection skip draft entries while they remain drafts.
We don't: Require human approval for memory to become useful, or make the UI a step in the write path. draft is a provenance label plus a corroboration threshold, not an approval queue.
Why: The agent is the consumer and it is usually headless — memory that only graduates when a human clicks a card is memory that never graduates. Labeling keeps the useful half of the caution (inferred ≠ recorded, and unreviewed single-task inference stays out of the prompt) without taxing the normal path. A draft graduates on corroboration: a second task matching it through the model's own writes, or the model writing the same knowledge at project scope, which links the existing entry instead of duplicating it.
7. Full entries returned directly
We do: query_memory returns complete entries with path:line citations. Every hit can be verified against source.
We don't: Return a minimal index first, then require a second tool call for details.
Why: Returning full entries preserves verifiability — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300-char summary + citation, plus a search-only terms field that never enters the prompt); the token cost is lower than a second tool call + context switch.
8. Symbol extraction focused on what developers search for
We do: Regex-based symbol extraction (functions, classes, methods, interfaces, type aliases) with string/comment masking, multi-line signatures, and cross-file linking by symbol name. For TypeScript/JavaScript projects, an optional L2 enhancement layer uses the TS Compiler API to infer return types, resolve generics, and extract interfaces — all cached by content hash for instant reuse.
We don't: Tree-sitter AST parsing, import graphs, call graphs, or full-program type resolution across files.
Why: Our regex scanner handles 8 languages with zero dependencies, runs in <1 ms/file, and captures the declarations developers actually search for (names, signatures, generics). The optional TS layer adds semantic depth for TS/JS without native deps. Cross-file linking by name covers the most common "find related code" use case. Full-program analysis would add native binaries, 10x install size, and version fragility — for marginal gain on the remaining 5% of edge cases.
9. forget by query is aggressive; prefer ID deletion
We do: forget query deletes all experience notes with ≥0.5 token overlap.
We don't: Interactive confirmation, soft-delete/trash, or exact-match-only.
Why: Experience notes are low-stakes, high-volume, and retrieval-only. Aggressive deletion prevents stale noise from polluting search. For precision, delete by ID (shown in query_memory output).
10. TypeScript enhancement is optional, lazy, and cached
We do: L2 TS Compiler API enhancement runs async in a priority queue (P0 on fs/observed, P1 on watch, P2 on index_repo), results cached by content hash in type-cache/. Zero config — just npm i -D typescript@5 or typescript@6. Falls back to L1 regex if TS absent or disabled.
We don't: Mandatory TS, blocking enhancement, or full-program type checking.
Why: Mandatory TS would break installs for non-TS projects. Blocking enhancement would stall index_repo on large codebases. Full-program checking is 10x slower and memory-heavy. Our design: enhance what's read, cache it, never block the hot path.
11. Subagent sessions are out of scope for now
We do: Exclude sessions spawned as subagents (origin: 'subagent' / delegationDepth > 0) from auto-creating or binding a task. Their todo_write events do not create tasks, and they inherit no task binding.
We don't: Merge a delegated run's steps and files back into the task that spawned it. That is not designed yet: there is no parent-link model for delegated work, and the naive version mints one project task per subagent.
Why: Every subagent that writes a todo would otherwise create its own task entity, so one fan-out run would flood the task list with ephemeral entries nobody resumes. Excluding them keeps the task list equal to the work the user actually owns. The cost is that a delegation's progress is invisible in the task record; merging it properly (child steps folded into the parent, or a separate delegated-work view) is future work.
Configuration
| Key | Default | Meaning |
|---|
memoryDir | .dsh-project-memory | store directory inside each indexed root |
chunkChars | 3000 | max chars per document chunk |
maxChunksPerFile | 40 | max chunks per document |
maxFileSizeMb | 50 | skip documents (incl. PDF) and code files larger than this (MB) |
maxOutputChars | 8000 | cap for query_memory result text (chars) |
tasklist.enabled | true | enable TaskBridge auto-sync (task entities from the session todo list and file reads) |
tasklist.syncHostOnAdopt | true | when select_task//task switch binds a task, push its steps to host todo/write so dsh's task list mirrors the task |
maxPdfPages | 1000 | PDF page cap when pages are not otherwise limited |
llmQueryExpansion | false | expand queries via ctx.llm before BM25 (off by default to save tokens) |
expansionCount | 6 | max expansion variants |
lazyIndexing | true | index files the moment the model reads them (fs/observed) |
autoIndexOnFirstUse | false | full scan of the current working directory on plugin load (opt-in) |
watch | true | enable the background refresh |
watchInterval | 15 | poll interval (seconds) |
tsPath | (auto) | optional absolute path to a specific typescript install; if omitted, resolves from project cwd → plugin node_modules |
enableTypeScript | true | set false to disable L2 TS enhancement entirely (L1 regex only) |
insight.* |
Injection admission (why it stays quiet)
Automatic injection used to be a retrieval problem ("which entry is most related to this text?"), which is total — a ranking always returns something, so noise was structural. It is now an admission problem ("is this step about to cross a boundary I have been burned by?"), with silence as the default:
- Only
when triggers, and it is a low-dimensional typed signal: normalized ops, the files this step is about to write, and intent words from the human message after stripping quoted/path references. Corpus text — raw tool arguments, file contents, filenames — can never trigger anything.
guard only narrows. Extension/name globs (*.pptx, README*) are ignored outright: they can only lie, never narrow.
- Ratio plus an absolute floor. The hint channel needs the relative score and an IDF-weighted coverage floor and at least two shared terms —
relative:1.00 also happens on entries that share nothing with the step.
- Frequency is bounded. At most one item injection every
gateCooldownSteps, capped per session by count and characters. The resident task card is exempt (it is a snapshot that should update); the budget is a ceiling, not a target.
- Prefix-cache discipline. Injections are appended as a user message at the tail of the history, so the cached prefix is never rewritten. What they add is resident cache-read tokens, not cache misses; nothing is ever edited in place.
- It is auditable. Every real injection appends one line to
injection-audit.jsonl (reason, dropped candidates, session budget), and npm run eval:injection scores 8 labelled scenarios — currently precision 1.00 / recall 1.00 with a clean control group.
Toggling features
The two most relevant switches are lazyIndexing (index a file the moment the model reads it; default on) and autoIndexOnFirstUse (full scan of the current working directory on plugin load; default off). Lazily indexed project roots are automatically registered with the watcher, so changed files stay fresh without an explicit watch_repo.
Settings live in the plugin's config object. To change them, add an override entry to your profile's cordis.patch.yml — for the web profile that is ~/.dsh/profiles/web/cordis.patch.yml:
- id: project-memory
config:
lazyIndexing: true # on: index files as the model reads them (default)
autoIndexOnFirstUse: false # off: no upfront full scan (default)
llmQueryExpansion: false # off: do not spend tokens on LLM query expansion (default)
watch: true # on: background refresh for watched roots (default)
watchInterval: 15 # poll interval in seconds
enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
# budgetLog: once # debugging: log budget drops to stderr (default off = silent)
# reinjectItemsAfter: 20 # debugging: allow the same insight again after N steps (default 0 = once per session)
# tsPath: /custom/path/to/typescript # optional: force specific TS install
Only list the keys you want to change; the rest fall back to the plugin defaults. Verify the result with dsh --profile web --dump-config.
For a one-off run without editing the profile, pass the override as a CLI patch overlay:
dsh web --patch ./config.yml
where config.yml contains the same override block.
Development (for contributors)
These commands are for maintaining the plugin code — regular users do not need them. Installing the plugin only requires the command in Installation.
npm install
npm test # 313 tests (184 core + 16 TaskBridge + 11 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 6 readiness-eval + 6 ops + 6 injection-audit + 5 injection-budget + 6 injection-scenarios)
npm run eval:injection # scenario P/R: 14/14 hits, 0 false positives, control group clean
npm run selfcheck:triggers # which entries can still push, which declarations are dead
npm run bench -- /path/to/project # index/query performance on any project — no dsh needed
Release notes live in CHANGELOG.md and on GitHub Releases.
License
MIT