description: "The out-of-tree wechat_scrape tool: retrieves WeChat official-account articles, filters promotional blocks, and saves each as a Markdown file, for maintainers of this machine's dsh profiles."
kind: "package-reference"
@deepseek-ai/dsh-ab-wechat-scrape
Summary
dsh-ab-wechat-scrape gives this machine's dsh profiles one way to read WeChat official-account articles: a single wechat_scrape call retrieves one link or several, filters the promotional blocks WeChat mixes into the article body, and saves every article as its own Markdown file named after the article title. A link may be an article page or a list or collection page: a collection page's article links are parsed out with the cursors that address its remaining pages, a generic list page's own next-page link is followed while the deployment sets a depth, and every link is then scraped the same way an individually named article is. It resolves the short-link interstitial, paces and retries its own requests with a jittered interval, carries the session cookies WeChat sets during the call, holds the whole instance back after a refusal, and falls back to a browser render for a page whose body is built by client script or whose plain request was refused — waiting for that body to fill rather than trusting a fixed delay. It reports WeChat's block pages instead of returning their HTML, and it writes a file only for a page that is still an article with a usable body after filtering: a refusal, a withdrawn article, and a body the filtering emptied are reported as failures and produce no document. It also contributes the scope-aware system-prompt section that says when to reach for it.
Table of Contents
Use this package
Mount the row in a profile when the agent should read official-account articles. The row injects ctx.tools, ctx.systemPrompt, and ctx.fs: the tool reaches the model catalog through the ordinary schema assembly, the prompt through the ordinary section assembly, and every article file is written through the mounted filesystem capability, so a deployment's file policy sees the write rather than the plugin going around it. A profile that mounts no filesystem provider fails the row at load instead of writing unobserved.
./invariant
The package publishes none. An invariant companion is warranted only when independent observations of one owned relation can diverge; here the plugin owns no relation that a second observer reads, so there is nothing for an invariant to compare. The same-title collision rule is proven by the persistence suite instead.
Mount it in a profile
The package is out-of-tree and published to the npm registry, so the profile resolves it from npm — never a local link:, symlink, or junction. dsh plugin add writes a real version range into the profile package.json, and pnpm install pulls the published artifact into node_modules.
# from <DSH_HOME>/profiles/<profile>
dsh plugin add dsh-ab-wechat-scrape
pnpm install
The cordis.patch.yml insert names the same package and carries the deployment overrides; it is what makes the plugin behaviour adjustable through cordis.patch.yml without a rebuild:
# <DSH_HOME>/profiles/<profile>/cordis.patch.yml
- insert:
- id: wechat-scrape
name: 'dsh-ab-wechat-scrape'
config:
minIntervalMs: 3000
minIntervalJitterMs: 1500
blockCooldownMs: 120000
renderer: auto
rendererDriver: playwright
browserChannel: msedge
renderWaitMs: 15000
minBodyChars: 20
outputDir: wechat-articles
The dsh_install_dsh_ab_plugin.py helper performs the same dsh plugin add + pnpm install sequence (with the registry and peer-dependency guards) for every out-of-tree plugin, including this one, and is the supported install path.
Configuration
Every field carries a default, so a bare row mounts and a deployment overrides only what its network and session require.
| Field | Default | Meaning |
|---|
| userAgent | desktop Chrome UA | User agent every request and browser context sends |
| referer | https://mp.weixin.qq.com/ | Referer sent with each request; empty omits it |
| cookie | empty | Cookie header of a logged-in session; empty sends none |
| cookieJar | true | Whether cookies WeChat sets during a call are replayed for the rest of it |
| minIntervalMs | 2000 | Minimum milliseconds between two requests from this instance |
| minIntervalJitterMs | 1000 | Random milliseconds added to every wait, so the cadence is not a fixed period |
| blockCooldownMs | 60000 | Milliseconds the instance stands down for after WeChat refuses a page |
| timeoutMs | 20000 | Per-request and navigation timeout |
| maxRetries | 2 | Retries after the first attempt (0-5) |
| backoffBaseMs | 800 | Base delay of the exponential backoff |
| maxRetryAfterMs | 60000 | Longest wait a server Retry-After hint may impose |
| maxChars | 40000 | Maximum characters of the body the model reads; the saved file always holds all of it |
| maxImages | 50 | Maximum image URLs the recorded value may carry |
| maxArticles | 10 | Maximum articles one call may scrape (1-200), and the ceiling a list walk stops at |
| maxListPages | 0 | Collection pages one list walk may read; 0 reads until the collection ends, and only then is a generic list page's own next-page link followed |
| filterAds | true | Whether promotional blocks are removed from the body |
| promoMaxChars | 200 | Longest block the promotion text rule may drop, in characters |
| promoPatterns | [] | Extra promotion rules, as regular-expression sources, beside the built-in ones |
| minBodyChars | 10 | Shortest filtered body that is still saved; below it the call reports a failure instead of a file |
What each call records
The model supplies one link or a list of links, an optional body format (markdown, the default, or plain text), whether to list image URLs (default true), whether every link is a list page, and — instead of a link — an optional keyword query. A link that is an article is saved directly. A link that is a list page has its article links parsed out, and each of those is scraped the same way; the list page itself is not saved. A page that carries no article body is read as a list page automatically, so a collection link needs no flag. A query searches the WeChat ecosystem through Sogou Weixin and returns the ranked results as links (title, account, snippet, verified badge, and the Sogou redirect link); when searchScrapeTop is above zero, the top results are resolved to their real article URLs and scraped and saved exactly as named links are.
Every article that was retrieved is written to its own file under outputDir and reported with its canonical value: the requested link, the resolved link, the article facts, the body's length in characters, a bounded excerpt of it, the image URLs, and whether the excerpt was clipped or the page was browser-rendered. The body itself is not in the value — it is in the file, and file names it — so a call never pays an article's full length in context or in the session log. Every list page walked is reported with its title and the article links it produced. A link or article that failed is reported beside the ones that succeeded, so one refused article does not discard the rest of a batch.
What is never written
A file is the record of an article, so the tool writes one only for a page that turned out to be an article: it carries a body, and that body is still at least minBodyChars long after promotion filtering. Every other outcome is a reported failure with no document beside it.
- A refused page. WeChat's verification and rate-limit pages are reported, not saved. The refusal also holds the whole instance back for blockCooldownMs, because it says the instance is asking too fast.
- A withdrawn article. A page whose own text says the publisher deleted it, that WeChat removed it, or that the account is closed is reported as having no article. These markers are only decisive when the page has no body element, so a live article that quotes one of them is still saved.
- A body the filtering emptied. filterAds can remove everything a page carried. The rule is applied to the text that would actually be written, so an article that was nothing but a follow prompt and a QR code is reported rather than saved as an empty file.
- A body below the minimum. minBodyChars is the deployment's floor. Lower it, or set filterAds: false, when a deployment genuinely needs very short articles.
Where article files go
outputDir resolves against the calling session's workspace directory. wechat-articles under a session whose workspace is /work/project writes /work/project/wechat-articles/<title>.md; an absolute outputDir is used as given. The file carries YAML front matter with the title, account, author, publish time, resolved link, cover, the body's character count, and the rendered flag, then the whole body and its image list. A collection is written flat into the same directory, one file per article.
Writing under a confining filesystem
A deployment that mounts @deepseek-ai/dsh-fs-sandbox (the web profile does) fences every write by a per-call sandbox policy, and that policy names the workspace the calling session runs in. The plugin resolves ctx.sandboxPolicy for each call and stamps it onto the write — the same thing write and edit do — so a file lands inside the session's own workspace.
This is not an optimization. A confining backend that receives no policy falls back to the deployment's own workspaceRoot, which for the web profile is the server's launch directory (process.cwd()), not the session workspace. A session working anywhere else is then fenced out of its own output directory, and every article fails with file access denied under workspace-write mode while the built-in write tool keeps working. A row that mounts a confining filesystem with no ctx.sandboxPolicy fails at load instead of silently writing nowhere.
Because the policy's root is what a relative outputDir resolves against, outputDir must name a directory inside the session workspace. An absolute path elsewhere is refused, and the refusal says which root it was fenced by.
The file is the record and is never truncated: maxChars bounds only the excerpt the model reads, so changing it cannot cost the saved article its tail.
Understand the implementation
Implementation internals — click to expand
Design commitments
- The call is self-contained. The tool reads no session state beyond the workspace directory it writes into, and its whole result is the canonical value, so a replay renders the same text.
- One page is classified before it is used. An article body decides an article; an article title without a body is an article shell a render must fill; a page reporting a withdrawn article is neither an article nor a list; anything else that links to articles is a list page. The classification runs before a browser render, so a collection page never costs one.
- Every article link is scraped the way a named article is. A link a list page produced goes through the same fetch, interstitial, filter, and save path as a link the model supplied, so a collection cannot behave differently from a single call.
- Anti-crawl behavior is policy, not luck. Requests carry a browser navigation header set, an optional session cookie, and the cookies WeChat sets during the call, never start closer than minIntervalMs plus a fresh random jitter to the previous one, retry only rate-limit and server failures — waiting the server's own Retry-After hint, capped by the deployment, when it gives one — hold the whole instance back after a refusal, and turn WeChat's verification pages into an actionable error instead of an empty article.
- Client-rendered and refused pages are explicit fallbacks. The short-link interstitial is resolved from its own inline script before a browser is considered. A page that carries no article body, and a page WeChat refused, reach the renderer; the render waits, bounded, for the article body element to fill before it reads the page, and the browser instance is reused, closed on idle, and disposed with the plugin fiber.
- Promotion filtering is two rules over one body. A structural rule drops the elements WeChat marks as advertising; a text rule drops a block whose whole text is a follow or promotion prompt and which is short enough to be a prompt rather than an article section. A deployment widens the text rule with its own patterns and moves the length ceiling.
- A file is written only for an article. The body the filtering leaves is what the minimum applies to, so a refusal, a withdrawn article, and an emptied body all end as reported failures with no document written.
- The file name is derived, never trusted. The article title is normalized for the characters no file system accepts and for the invisible ones a reader cannot see, capped at fileNameMaxChars, checked against Windows' reserved device names, and — when the title yields nothing at all — replaced by the article's own link identity.
- Bounds are deployment policy. The body, image, article, and list-page ceilings, the retry count, the cool-down, the timeouts, the minimum body, the promotion rules, the name length, and the renderer choice are validated Config fields. A malformed promotion rule fails the row at load rather than partway through the first article.
Further Exploration
Model Experience
System-prompt section
What the model sees
Every request where the tool is visible carries tool:wechat_scrape at order 2970, after the built-in tool band and before the SDK section. It routes official-account article and collection links to wechat_scrape and names what one call saves and returns. The section is empty in any scope where the tool is not visible.
Token effect
A fixed per-request cost for the section text wherever the row is composed and the tool is visible.
KV Cache effect
Prefix-stable while the section text, its order, and the row's visibility are unchanged.
Tool schema
What the model sees
One tool schema, wechat_scrape: an optional url, an optional urls array, an optional query string, an optional list flag, an optional format enum (markdown, text), and an optional includeImages flag. A call must name either a link (url or urls) or a keyword query; a query takes the search path, a link takes the scrape path.
Token effect
Fixed schema cost on every request where the tool is visible.
KV Cache effect
Prefix-stable while the definition and visibility are unchanged.
Tool-call history and result
What the model sees
Each assistant tool call retains the requested links and options in its arguments. Success returns one text block per list page walked, naming the page, its link, how many articles it produced, and each article link; then one text block per saved article with the title heading, an account/author/publish/link line, the body, the image list, and the path the article was written to; then one block per failure with the exact failure text. Result text for an article that carried promotional blocks is the filtered body. A link that turned out not to be an article — a refusal, a withdrawn article, a body the filtering emptied — appears only as a failure block, with no file named.
Token effect
Arguments and result are data-dependent and resent until compaction; each excerpt is bounded by maxChars, each image list by maxImages, and at most maxArticles articles are returned, so the whole result is bounded by maxArticles × maxChars however long the articles are. A list block carries one line per link it produced. Reading a whole article costs one read of the file the result named, which is also what makes the record recoverable after compaction.
KV Cache effect
None beyond the ordinary request prefix.
Known Limitations and Deferred Work
- List walking is one level deep: an article found through a collection is scraped and never expanded again, so a collection of collections is not followed.
- Only
mp/appmsgalbum collection links page by begin_msgid/begin_itemidx. A generic list page continues through its own next-page anchor, and only while maxListPages is above zero; a list page that paginates by some other mechanism yields only its first page.
- A collection page carries no title of its own, so a walked collection is reported by its link rather than by a name.
- The article write requires a mounted
ctx.fs; a profile with no filesystem provider cannot load this row. The capability's local backend publishes atomically and creates the output directory, so the plugin does no directory work of its own.
- Under a confining filesystem the write is fenced by the calling session's sandbox policy, so an outputDir outside the session workspace is refused rather than created. A deployment that needs a shared absolute output directory has to grant it in the deployment's own
workspaceRoot or mode; the plugin cannot widen a fence on the caller's behalf.
- The per-call policy is resolved once per tool call from
ctx.sandboxPolicy, so a session whose sandbox mode changes between calls is fenced by whatever mode it holds at the call that writes the file.
- The body converter is a focused tag walker for the elements WeChat article bodies use; it is not a general HTML-to-Markdown engine, and nested tables lose their cell boundaries.
- Promotion filtering is rule-based, not learned. It cannot recognize a promotional block whose text matches none of the rules, and the text rule can drop a short genuine sentence that names a promotion verb; promoPatterns, promoMaxChars, and
filterAds: false are how a deployment adapts it.
- The minimum-body rule is a character count, not a judgement: a genuine article shorter than minBodyChars is reported instead of saved, and a long page of boilerplate that survived filtering is saved.
- A page that carries an article title but no body is treated as an article shell and rendered, not as a list page. A list page that carries such a title needs
list: true.
- A withdrawal marker is read from the page head only, and only on a page with no body element; an article whose body was emptied by filtering is caught by the minimum-body rule rather than by that marker.
- A block page is reported, not circumvented: when the browser render is refused as well, an expired or absent cookie still fails the call, and the remedy is a deployment cookie plus a lower request rate.
- The cookie jar lives for one plugin instance and is never persisted, so each process starts as a fresh visitor; the deployment's configured cookie is what carries an authenticated session across restarts.
Publish
The package is publishable to a public registry by default. Its manifest carries publishConfig.access: "public", license: MIT, keywords, engines, and a real ^ range for every @deepseek-ai/* dependency (resolved from the checkout's published versions), so a clean pnpm verify followed by pnpm release publishes it as written.
pnpm verify # typecheck + test + build + lint (also runs as prepublishOnly)
pnpm release # pnpm publish --access public --no-git-checks
The runtime host deps (@deepseek-ai/dsh-tools, @deepseek-ai/schemastery) and the Cordis peer resolve from the registry. The filesystem, sandbox, and Cordis-loader packages stay dev-only (the tests mount a real local backend and a real Loader) and are not bundled into the published artifact. playwright is an optional dependency — the renderer lazily imports it, so a deployment without it still loads the plugin until a browser render is actually required.
Never publish under the reserved @deepseek-ai/dsh-* scope: that scope is for official plugins, and reusing it would shadow them. This package is published unscoped as dsh-ab-wechat-scrape.
Dev Note
pnpm build # tsc -p tsconfig.json && tsdown
pnpm typecheck # tsc -p tsconfig.json --noEmit
pnpm test # node --test (build first)
pnpm lint # publint && attw --pack .
pnpm verify # typecheck + test + build + lint
pnpm clean # rimraf lib
The host imports lib/index.js and the tests assert the emitted modules, so a green suite means the artifact the profile row resolves is the artifact that behaves. node --test discovers *.test.mjs only, so every suite is named that way. pnpm lint runs publint (manifest hygiene) and attw (types against the emitted d.ts), which is why files ships lib/**/* and the .d.ts paths in types/exports point at the flat lib/ layout tsc actually emits.