Install
Inside DeepSeek Harness, with dsh-market
dsh plugin --profile web add dshmarket
Or from the command line
dsh plugin --profile web add dsh-fs-encoding
Installing runs third-party code with your own permissions — it can read your files, use your credentials and reach the network. Review the source first, and pin a commit (github:owner/repo#sha) when you can.
README
A file-encoding governance plugin for DeepSeek Harness (DSH): stop the AI from wrecking BOMs and multi-byte encodings when it reads and writes files
🌏 中文 | English
dsh · dsh-plugin · plugin · encoding · BOM · GBK · Big5 · Shift-JIS · UTF-16 · AI agent · 编码 · 文件编码 · 乱码
Introduction
DSH's built-in read / write / edit tools only understand UTF-8, which breaks down on the encodings common in Chinese, Japanese and Korean projects:
- GBK, Big5 and Shift-JIS files simply cannot be read — they fail with
invalid UTF-8 textand the AI is left staring at an error; - UTF-8 files with a BOM can be read, but the BOM is silently swallowed — one casual edit and those first three bytes are gone. For PHP, older compilers and some Windows software, a missing BOM can mean a parse error or garbled output;
- Worse, there is no warning at all: you find out the file is broken days later, with no idea which edit did it.
This plugin takes over those three tools and, while preserving every existing behaviour (sandbox fence, read-before-write protection, version checking, diff display), makes non-UTF-8 files and BOMs round-trip correctly: whatever encoding a file had, it still has after the edit.
Features
- Byte-exact preservation: a file's encoding is decided once, at first read, and inverted on every save. Edit a GBK file and it is still a GBK file — never quietly converted to UTF-8.
- BOM fidelity: a BOM is restored exactly when the file had one, and never invented for a file that did not — correct in both directions.
- Line-ending fidelity: CRLF / LF / CR are detected at read and restored on save, so Windows projects are not rewritten to LF.
- No silent corruption: if the target encoding cannot represent the new content (an emoji in a GBK file), the plugin refuses the write and explains why, leaving the file untouched — instead of filling it with
?and destroying it. - It tells you what to do when it cannot read: for a non-UTF-8 file, the plugin lists the most likely encodings with a sample decode of each, so the AI (or you) can pick one and re-read — like VS Code's "Reopen with Encoding".
- Nineteen encodings: UTF-8 (with BOM), UTF-16, UTF-32, plus GBK, Big5, Shift-JIS, EUC-KR and the full Windows-125x family (Western, Central European, Cyrillic, Greek, Turkish, Hebrew, Arabic, Baltic).
- Zero learning curve:
read/write/edittake the same arguments and return the same shapes as the built-ins — drop-in replacements, so existing prompts and habits keep working;readandwriteeach gain one optionalencodingargument (the one onwriteapplies to new files only, see below). - Create files in a chosen encoding: one
encodingargument onwriteproduces a GBK, Shift-JIS or other legacy-encoded file directly — no writing UTF-8 and converting afterwards. - Three further tools:
insert(insert at a line number, which the built-ins cannot do),undo_last_edit(revert the last edit, content and encoding together) andstr_replace_editor(a compatibility layer for its four commands, also encoding-governed). See "Three further tools" below. - Simple configuration: one YAML file with a few switches, all overridable by environment variables.
Usage
Works out of the box
Install it and you are done — no configuration required. There is nothing to fill in, no initialization step, and nothing is written into your project. The AI keeps using read / write / edit as before, and the plugin handles the encoding:
- Reading a UTF-8 file (with or without a BOM): identical to the built-ins, and the BOM is no longer swallowed.
- Reading a UTF-16 / UTF-32 file that carries a BOM: detected automatically.
- Editing a GBK, Shift-JIS or other legacy-encoded file: saved back in its own encoding, never silently converted to UTF-8.
- Writing a new file: UTF-8 by default, or any other encoding with one extra argument.
None of this needs anything from you, and the AI needs no change in habits. In everyday use its calls are exactly the same as with the built-ins.
Where the AI's calls differ
One case behaves differently: a non-UTF-8 file with no BOM (typically a legacy GBK / Big5 / Shift-JIS file). The plugin does not guess. It stops and has the AI re-read with an explicit encoding:
[E_NOT_TEXT] legacy.txt is not valid UTF-8. Most likely gbk. Re-read with
read({ file_path: "legacy.txt", encoding: "gbk" }) to decode it, or set
autoGuessEncoding: true in the plugin config to decode automatically.
Candidates: gbk("你好,世界"), big5("斕疑"), shift_jis("ト羲")
The AI re-reads with the call shown in the message and handles it on its own — you are not involved. After that re-read the encoding turns from a guess into a known fact that every later save follows.
read and write therefore each gain one optional argument:
read({ file_path: "legacy.txt", encoding: "gbk" })
write({ file_path: "run.bat", content: "echo 中文\r\n", encoding: "gbk" })
UTF-16 / UTF-32 files without a BOM work the same way — specify explicitly with
read({ file_path: "<path>", encoding: "utf16le" }). Such files are rare on Windows, and without a BOM the byte order cannot be reliably detected, so the plugin does not guess.
Why does it ask by default? GBK, Big5 and Shift-JIS byte ranges overlap on short inputs, so a wrong guess is invisible in the UI — and would then be written back under the wrong encoding, ruining the file. The plugin therefore prefers "fail rather than guess wrong". If you would rather have best-effort decoding, set
autoGuessEncoding: true.Even with
autoGuessEncodingon, one case still fails loudly with the candidate list: a very short file (a few bytes) where two independent detectors name different pages. Nothing then distinguishes them — measured, the top pick is wrong about 79% of the time in that case — so the plugin lets you choose from the candidates instead of gambling for you. When both detectors name the same page it is adopted directly, and files of any ordinary length (tens of bytes and up) effectively never hit this.A second case that fails loudly: a detector names a single-byte page such as ISO-8859-1 or Windows-1252, but decoding with it yields text that is almost entirely non-ASCII. Real Western text is mostly letters and spaces, so it does not look like that — whereas 2-byte CJK, Korean or Cyrillic text read as a single-byte page turns every character into two Latin ones, which is exactly that shape. The plugin refuses the verdict and lists candidates instead. Measured, this catches 24 files that would otherwise be silently mis-read, and refuses none that previously read correctly.
Creating a file in a chosen encoding
A new file is UTF-8 without a BOM by default. To generate a GBK file for a legacy system, add encoding:
write({ file_path: "run.bat", content: "echo 中文\r\n", encoding: "gbk" })
The accepted names are the same as read's, and aliases and case are insensitive (cp936, Shift-JIS both work). Names that carry a BOM (utf8bom, utf16le, utf16be, utf32le, utf32be) write it; every other name writes none. Content the encoding cannot represent is refused, never written as ?.
encodingapplies to new files only. Passing it for an existing file fails withE_ENCODING_NOT_APPLICABLEinstead of converting — preserving a file's encoding is this plugin's core promise, and a conversion rewrites every character of the file in a way the AI cannot see from the reply. Do not delete the file to force a conversion: deleting bypasses the read-before-write gate, so content the session never read disappears silently behind abefore: null, and if the session had read the file, every later write failsFS_STALE_VERSIONand the path cannot be recreated for the rest of the session. This plugin does not convert encodings; when you need a copy in another encoding, write it to a new path.
Three further tools
Beyond the three above, the plugin registers three more tools, all encoding-governed (a save lands in the file's own encoding).
insert — insert at a line number, which the built-ins cannot do. A literal-match edit can only insert where it can quote surrounding text, so "add a line at the top" means reading that line first and reproducing it exactly:
insert({ file_path: "config.ini", insert_line: 0, new_string: "[core]" })
insert_line names the line to insert after: 0 is the very top, and the file's line count appends. Line numbers match exactly what read shows.
undo_last_edit — revert a file's most recent edit, restoring the content and the encoding together:
undo_last_edit({ file_path: "config.ini" })
Use it when an edit produced the wrong result, or when the AI notices its own last change was mistaken. The points that matter:
- Only the most recent edit is kept. Edit a file twice and only the second is undoable; undo once and the history is spent (there is no redo).
- The encoding is restored too. If that edit converted a GBK file to UTF-8 (see
normalizeToUtf8), the undo turns it back into GBK — not merely the characters. - A changed file is refused. Before reverting, the plugin checks the file still matches what that edit wrote; if anything changed it since — you, another tool, another session — the undo is refused with an explanation rather than overwriting those changes.
- In memory only. Nothing is written to disk and no repository is polluted, so an undo does not survive a DSH restart — the tool says it has no history rather than pretending to succeed.
- Creating a file leaves no undo point: "undo a creation" means deleting the file, which is too destructive. Delete it directly instead.
str_replace_editor — a compatibility layer for its four commands, so prompts and habits written for that tool keep working:
| Command | What it does |
|---|---|
view |
Shows a file with line numbers (narrow with view_range); for a directory, lists two levels deep, skipping hidden entries, node_modules and __pycache__ |
create |
Creates a file; refuses one that already exists |
str_replace |
Replaces the unique match; refuses an ambiguous one and names the lines it found |
insert |
Same as the insert tool, same insert_line semantics |
undo_edit |
Same as undo_last_edit — reverts the last edit |
str_replaceadditionally accepts an optionalreplace_all: omit it for the standard behaviour (a unique match is required), passtrueto replace every occurrence. It is the only argument this plugin adds.The native
str_replaceandinsertunderstand UTF-8 only — a GBK file fails outright or gets converted. That is precisely why this plugin takes them over.
Installation
⚠️ Conflict: any plugin that registers
read/write/edit/insert/str_replace_editor/undo_last_edit— any one of those names — on the same scope layer is mutually exclusive with this one; registering a name twice in a layer throws.If those names are already held by another plugin on the same layer, the install is refused with the offending tool named, rather than leaving a half-registered tool set. To resolve it, either remove this plugin from the profile or disable the plugin that holds the name:
# in the profile's cordis.patch.yml - id: <the other plugin's id> disabled: trueNote: the built-in
read/write/editlive on an outer host/preset layer and are not a conflict — shadowing them on the agent's own layer is exactly what this plugin does, matching how the native tools shadow each other.
Option 1: install from the plugin market (dsh-market, recommended)
If you already have dsh-market (the DSH plugin market): open Settings → Plugin Market, search for dsh-fs-encoding, click Install on the card and confirm, then restart dsh web.
Market card: https://awesome-dsh-plugin.com/p/MrWeiCodes/dsh-fs-encoding/
Option 2: let the AI install (easiest)
Just give your DSH AI assistant the repository URL, e.g. "install the plugin https://github.com/MrWeiCodes/dsh-fs-encoding". The AI handles plugin loading, dependencies and the patch for you; then restart dsh web.
Option 3: install from npm (recommended)
dsh plugin --profile web add dsh-fs-encoding
Why this is the recommended path: the npm package already ships the compiled lib/, so the install runs no build scripts at all — unaffected by pnpm's build-script gate, and independent of your local build environment. Then restart dsh web.
Option 4: install from GitHub
dsh plugin --profile web add -w github:MrWeiCodes/dsh-fs-encoding
A GitHub install fetches the sources, so lib/ is compiled on the spot by the prepare script — which means you may need to allow build scripts in the profile's pnpm-workspace.yaml (pnpm 10 and later blocks dependency build scripts by default; paste the line it prints and re-run). If you would rather skip that, use Option 3 — the npm package already contains the compiled output and has no such step.
Known issue with local-directory installs: on Windows, if the plugin directory and the profile are on different drives (e.g. plugin on
G:\, profile onC:\), pnpm mis-resolves thefile:dependency toC:\Users\<username>\...and the install fails. Use Option 5 instead.
Option 5: manual installation
Fallback for environments without pnpm or for offline use:
- Clone this repository into your DSH profile's plugin directory and build it once there (the
preparescript generateslib/):# example: web profile $dst = "$HOME\.dsh\profiles\web\packages\dsh-fs-encoding" git clone https://github.com/MrWeiCodes/dsh-fs-encoding.git $dst cd $dst npm install # also triggers prepare → generates lib/ - Add to the
dependenciesof the profile'spackage.json:"dsh-fs-encoding": "file:./packages/dsh-fs-encoding" - Append the contents of
cordis.patch.ymlto your profile'scordis.patch.yml. - Reinstall dependencies and restart:
pnpm install(ornpm install), thendsh web.
Updating
- Installed via Option 1 (plugin market): update from the plugin market card.
- Installed via Option 2 (AI): just tell your AI assistant "update the dsh-fs-encoding plugin".
- Installed via Option 3 (npm):
Then restartdsh plugin --profile web add dsh-fs-encoding@latestdsh web. The npm path involves no build step either. - Installed via Option 4 (GitHub):
If the latest commit is not fetched (git dependencies are cached), remove and re-add:dsh plugin --profile web add -w github:MrWeiCodes/dsh-fs-encoding
Then restartdsh plugin --profile web remove dsh-fs-encoding dsh plugin --profile web add -w github:MrWeiCodes/dsh-fs-encodingdsh web. The commit hash changes with the new commit, so if pnpm asks you to allow build scripts again, redo the step described under Option 4. - Installed via Option 5 (manual):
Then restartcd "$HOME\.dsh\profiles\web\packages\dsh-fs-encoding" git pull # or re-copy the new version over the old files npm install # run if dependency declarations changed; also re-runs preparedsh web.
Your configuration ($DSH_HOME/plugins/dsh-fs-encoding/config.yaml) is preserved across updates.
Uninstallation
Installed via Option 1 (plugin market): uninstall from the plugin market
Open Settings → Plugin Market, click Uninstall on the dsh-fs-encoding card (two-step confirmation); the plugin is removed immediately.
Installed via Option 2 (AI): let the AI uninstall
Just tell your AI assistant "uninstall the dsh-fs-encoding plugin"; the AI removes the plugin and cleans up the patch and dependencies.
Installed via Option 3 (npm) / Option 4 (GitHub)
dsh plugin --profile web remove dsh-fs-encoding
Then restart dsh web. The built-in read / write / edit come back immediately — nothing else to do.
Installed via Option 5 (manual)
- Remove this plugin's
insertrow from your profile'scordis.patch.yml. - Remove
"dsh-fs-encoding": ...from thedependenciesof the profile'spackage.json. - Reinstall dependencies and restart:
pnpm install(ornpm install), thendsh web.
Leftover locations (cleanup checklist for the AI)
- Config directory:
$DSH_HOME/plugins/dsh-fs-encoding/—$DSH_HOMEdefaults to~/.dsh(C:\Users\<username>\.dshon Windows); the plugin'sconfig.yamllives here. Uninstallation does not delete it automatically; delete the whole directory manually for a complete cleanup. - Plugin directory (Option 5 installs):
$DSH_HOME/profiles/<profile>/packages/dsh-fs-encoding/. - Dependency & patch (Option 5 installs): the
"dsh-fs-encoding": ...dependency in the profile'spackage.json, and theinsertrow incordis.patch.yml. - No global registry, npm global packages, or system-level writes; the plugin writes no events of its own into session logs.
Configuration
The defaults work — you normally do not need to touch this. Only change it to adjust how encoding guessing behaves.
Config file location (created automatically on first start, with full comments inside):
$DSH_HOME/plugins/dsh-fs-encoding/config.yaml
All options
| Option | Default | What it does |
|---|---|---|
autoGuessEncoding |
false |
On an unreadable non-UTF-8 file: guess, or fail and let you choose |
normalizeToUtf8 |
false |
Whether saving converts legacy files (GBK, …) to UTF-8 |
supportedEncodings |
built-in list | The encodings considered by automatic guessing |
excludeEncodings |
empty | Remove a few encodings from the built-in list |
maxFileBytes |
10 MiB | Read size limit for a single file |
Common needs (copy and paste)
Let it guess instead of asking every time
autoGuessEncoding: true
Be done with encoding problems for good (a GBK file becomes UTF-8 after its first save)
normalizeToUtf8: true
One encoding keeps guessing wrong (e.g. Cyrillic stealing Western text)
excludeEncodings: [windows-1251]
The encoding list: what the two keys do
The plugin ships a built-in list used for automatic guessing. You can subtract from it, or replace it wholesale:
| What you want | Which key | When the plugin adds an encoding later |
|---|---|---|
| Just drop one or two | excludeEncodings |
✅ You get it automatically |
| Use a list of your own | supportedEncodings |
❌ You will not receive it |
Why excludeEncodings is the better default: writing supportedEncodings freezes the list into your config file — encodings added in a later release will never reach you, and the plugin will not change it back for you (it never modifies an existing config).
The list only affects guessing, not reading. Any encoding can be read by naming it explicitly, even when it is not on the list:
read({ file_path: "legacy.txt", encoding: "windows-1253" })
Environment variables
Useful for a quick test or a container deployment. An environment override always wins over the config file:
| Variable | Option |
|---|---|
DSH_FS_ENCODING_AUTO_GUESS |
autoGuessEncoding |
DSH_FS_ENCODING_NORMALIZE_TO_UTF8 |
normalizeToUtf8 |
DSH_FS_ENCODING_SUPPORTED_ENCODINGS |
supportedEncodings |
DSH_FS_ENCODING_MAX_FILE_BYTES |
maxFileBytes |
excludeEncodings has no environment variable — it is the only option that lives in the config file alone.
normalizeToUtf8does not convert UTF-16 / UTF-32 files: they are already Unicode, and re-encoding them would break programs that depend on them.
Supported encodings
| Kind | Encodings |
|---|---|
| Unicode | utf8, utf8bom, utf16le, utf16be, utf32le, utf32be |
| East Asian | gbk (incl. gb18030, gb2312, cp936), big5 (incl. cp950), shift_jis (incl. sjis, cp932), euc-kr (incl. cp949) |
| Windows ANSI | windows-1250 (Central European), windows-1251 (Cyrillic), windows-1252 (Western), windows-1253 (Greek), windows-1254 (Turkish), windows-1255 (Hebrew), windows-1256 (Arabic), windows-1257 (Baltic) — each also accepted as cp12xx |
| Other | iso-8859-1 (incl. latin1) |
On
windows-1258(Vietnamese): not supported. Vietnamese needs combining sequences —ếis one code point but two bytes — and the underlyingiconv-litesingle-byte table cannot split them, so 52 of 67 common Vietnamese characters encode to?. Offering it would let a file be read but almost never saved, which reads as a bug rather than a limitation, so the name is deliberately absent.
This table is what can be read and written, not what gets guessed. Automatic guessing uses only the short list in the config (7 encodings by default) — a small set keeps the wrong-guess rate down. Everything else in the table still works when named explicitly:
read({ file_path: "x.txt", encoding: "windows-1253" }). See Configuration for how to adjust the list.
Names are case- and style-insensitive: Shift-JIS, shift_jis and SJIS all mean the same encoding.
Common errors
| Error | Meaning and fix |
|---|---|
E_NOT_TEXT |
Not valid UTF-8 and no BOM. Re-read with read({ file_path: "...", encoding: "..." }) as suggested, or enable autoGuessEncoding. |
E_UNMAPPABLE |
The target encoding cannot represent the new content (an emoji in a GBK file). The file was not modified — switch to an encoding that can, or migrate with normalizeToUtf8 (that migration applies to an existing file only; for a new file just name a different encoding). |
E_BAD_ENCODING |
Unknown encoding name; use one from the table above. |
E_ENCODING_NOT_APPLICABLE |
encoding was passed to write for a file that already exists. The argument applies to new files only; the file keeps its own encoding and nothing was changed. This plugin does not convert encodings — do not delete the file to force one (see above); write a copy to a new path instead. |
E_DECODE_FAILED |
The bytes do not decode under the requested encoding — the encoding is probably wrong; try another candidate. |
FS_STALE_VERSION |
The file changed after it was read (or was deleted). Read it again before editing, to avoid clobbering someone else's change. |
FS_NOT_OBSERVED |
This session has no encoding record for the file, so it cannot be rewritten: an existing file that was not read this session (or whose record was reclaimed) is refused rather than re-encoded as UTF-8. Read it once before writing; if it is not a text file, this plugin cannot rewrite it. Creating a new file is unaffected. |
FS_EDIT_NOT_FOUND |
The old_string to replace was not found; check the content and indentation. |
FS_AMBIGUOUS_EDIT |
old_string matched more than once. Add surrounding context to make it unique, or set replace_all: true. |
FS_SANDBOX_DENIED |
Refused by DSH's sandbox (e.g. writing outside the workspace) — the existing safety policy at work. |
Compatibility & conflicts
- Zero-intrusion: the plugin only uses DSH's public interfaces to register tools; it does not modify native DSH code and does not replace the
ctx.fsfilesystem itself. Uninstalling restores the built-in tools immediately, leaving nothing behind. - Every existing guarantee is kept: sandbox fence, read-before-write protection, version checking and observation records all behave exactly as before — anything that should be blocked or questioned still is.
- Mutually exclusive with any plugin registering the same tool names: this plugin registers
read/write/edit/insert/str_replace_editor/undo_last_edit, and registering a name twice in a layer throws. The test is whether any of those names is already occupied on that agent's own layer, regardless of who occupies it — on a collision it refuses to install and names the occupied tool (see Installation above). Built-ins on a host/preset layer are not a collision. - Session state: encoding information is kept in memory and isolated per session — never written to disk, never polluting your repository. After a DSH restart, the encoding is detected afresh on the first read.
- Publishes a service: once installed, the plugin provides the
fsEncodingservice so other plugins can reuse the same decoding rules (see For plugin authors below).
For plugin authors
DSH's ctx.fs is UTF-8-only: it decodes with a strict TextDecoder, so a GBK file is simply FS_NOT_TEXT to it. This plugin's tool layer solves reading for the model, but other plugins that call ctx.fs directly for content still hit that limit — so they either report an error for a file with perfectly readable content, or reimplement the guess themselves, and two parts of one deployment end up disagreeing about what a file says.
This plugin therefore publishes a service that keeps "what text is in these bytes" in one place:
// Consumer side: read it with ctx.get — no `inject` needed, and it is
// `undefined` when the plugin is not installed.
const fsEncoding = ctx.get("fsEncoding");
// Careful: when another plugin already owns the name, `ctx.get` returns THAT
// plugin's object rather than `undefined`, so test for the capability, not just
// for absence. `isFsEncodingService` verifies the members you name (with none
// named, `tryDecode` + `decode`). If you hand-roll the guard, the same rule
// applies: testing only `tryDecode` lets a `tryDecode`-only object through, and
// the later `decode` call then throws.
if (!isFsEncodingService(fsEncoding)) {
// Not installed (or the name is taken): fall back to what you did before.
}
// You own the IO — the sandbox and path resolution are your context.
const target = await ctx.fs.resolve(path, { cwd });
// All three arguments are required: `maxBytes` is the only cap `readBytes` has,
// and omitting it does not fail — it reads the whole file into memory. Both the
// signal and the cap come from YOUR execution context (you already hold an
// `exec`, and the cap is yours to choose).
const bytes = await ctx.fs.readBytes(target, exec?.signal, maxBytes);
const outcome = await fsEncoding.tryDecode(bytes, { displayPath: path });
if (outcome.ok) {
outcome.result.text; // decoded text, BOM removed
outcome.result.encoding; // e.g. "gbk"
outcome.result.decided; // "bom" | "utf8" | "hint" | "guessed"
outcome.result.hasBOM;
outcome.result.lineEnding;
} else {
outcome.refusal.message; // display-ready, same wording as the `read` tool
outcome.refusal.candidates; // candidate encodings, for a picker
outcome.refusal.ranked; // whether the order is evidence-based
outcome.refusal.adoptable; // whether the head is credible enough to adopt
outcome.refusal.autoGuessEnabled; // "not attempted" vs "attempted and failed"
}
Three design points:
decidedis the field you must look at. Only"bom"/"utf8"/"hint"are determined;"guessed"is a probabilistic pick. When rendering for a human, say which one you got — presenting a guess as the file's real encoding is exactly the failure this plugin exists to prevent.- Refusal is a value, not an exception.
tryDecodethrows for no input at all — including an argument error such as a non-stringencoding, which comes back as a refusal; usedecodewhen you want the refusal thrown with thereadtool's exact wording. - The service does not read files and writes no records. IO is yours; a decode writes no session encoding record, so a preview cannot count as having "read" a file and thereby authorize a later write. (It can read a record — see the next section. Reading one is not the same act as making one.)
Agreeing with a tool: recordedEncoding
The above answers "what text is in these bytes". But a consumer that has to agree with a tool needs more than that: edit, insert and str_replace_editor take no encoding argument, so their encoding comes from this plugin's own record. Guess from the candidates instead and you can land on a different page than the tool, so what you describe is not the same document the tool actually touches (measured: with a record of big5 the tool decodes as big5, while the top candidate gbk decodes the same bytes into a different document).
// `sessionId` must be the session YOUR execution belongs to: the record the
// tools write with is keyed by it. Without one (`undefined` or an empty string)
// you read the "no agent" bucket, which is invisible from every real session —
// so this necessarily returns `undefined`. Never use that as the basis; show
// nothing rather than something on the wrong page.
const sessionId = exec?.agent?.session?.id;
const info = await ctx.fs.stat(target);
if (info === undefined) {
// The file does not exist: there is no basis to agree with.
} else {
const rec = fsEncoding.recordedEncoding(sessionId, target, info.version);
if (rec) {
// rec.encoding is the encoding the tool will use
const bytes = await ctx.fs.readBytes(target, exec?.signal, maxBytes);
// Use `tryDecode`: a refusal is a value, not an exception.
// Note `maxBytes` is the MEMORY-GUARD cap you gave `readBytes`, which is a
// different thing from the service's `maxFileBytes` (10 MiB by default,
// deployment policy) — so "you could read it" does not mean "the service
// will decode it": a 10–64 MiB file comes back as `E_TOO_LARGE`, which you
// must handle as a refusal rather than let it throw.
// (You may pass `maxBytes` to raise the per-decode cap when you really mean
// to override deployment policy — but do not substitute your memory guard
// for it, since a small value rejects files the service could have decoded.)
const out = await fsEncoding.tryDecode(bytes, {
encoding: rec.encoding,
displayPath: path,
});
if (out.ok) {
out.result.text; // text on the same basis as the tool
rec.decided; // the real provenance, see below
rec.hasBOM; // the BOM a save will restore
rec.lineEnding; // the line ending a save will restore
} else {
// out.refusal.code is E_TOO_LARGE when the service's maxFileBytes was hit
out.refusal.message; // presentable explanation
}
} else {
// No usable record for this session: fall back to decoding the bytes
// (and present the result as a guess).
}
}
- Use the returned
rec.decided; do not re-derive it. Decoding withrec.encodingmakes the service answer"hint"("the CALLER specified this") — that is YOUR decision, not the file's real provenance. A file the record says was guessed would then be presented as a determination, which is the very failure this plugin exists to prevent. - Ask for what you call:
isFsEncodingService(fsEncoding, "recordedEncoding"). With no names it verifiestryDecode+decode(the historical answer); name the members you are about to use and it answers exactly that question. Naming them matters: the first registration of the service name wins, soctx.getmay hand you an older instance from an earlier mount (this method arrived in 1.4.0; earlier versions do not have it), and without the name you get "guard passes, call throwsTypeError". Do not over-ask either — checking members you never call rejects an instance whose decode entry points work perfectly. The answer is a boolean, so you can tell "the instance is too old to have this" from "this session has no record for that file" (the latter isrecordedEncodingreturningundefined) and warn instead of degrading silently. A consumer that cannot import this package writes the equivalent test by hand:typeof fsEncoding?.recordedEncoding === "function". The?.is not optional: when the plugin is not installedctx.getreturnsundefined, and without it the guard itself throws aTypeError— the very thing the guard exists to prevent (isFsEncodingServiceanswersfalseforundefinedrather than throwing). sessionIdis required, and must be your own session. The tools' record is keyed by the calling session (exec.agent.session.id); passingundefinedor an empty string reads the agentless bucket, which is invisible from every real session, so it necessarily returnsundefined. When there is no record, do not substitute another session's or a guess — show nothing instead.- Let the method decide staleness; do not compare versions yourself. Pass the
versionyou just got fromstatand a record taken at a different version comes back asundefined. Omitting the argument is not "skip the check" — it is read exactly as the write path reads an absent version (invalidateIfStale(…, undefined)deletes a versioned record), so a versioned record is reported as unusable. When you could not observe a version (astatthat returned nothing), treat that as "cannot confirm" rather than presenting the record as fresh. Passnullonly when you deliberately want the record as it stands, fresh or not (to show history) — that is the one spelling that skips the check, so merely lacking a version cannot reach it by accident. The version token is internal bookkeeping, and one shared comparison is what keeps a preview and the tool that follows it from disagreeing about whether the record still holds. undefinedcollapses "never recorded" and "record no longer valid" — the action is the same for both (decode the bytes and present them as a guess), and telling them apart would mean exposing the version token, the internal detail this method exists to keep out of the contract.- Cross-session reads are allowed. The service is host-plane, so one process serves many sessions; passing another
sessionIdreads that session's record. This cannot affect a write — the encoding a write uses is decided by the calling session alone — so reading another session's record can inform a display but never change what gets written.undefinedis the agentless caller, which has its own bucket, invisible to and from every real session. - It never writes a record, and cannot make a write possible. A record both decides the bytes of a write and is the guard that permits it, so a method able to create one could authorize a write the session never read. This one only reads: no state change, no events, no arming the gate — call it freely.
Other methods: isUtf8(bytes) (cheap check for whether these bytes are valid UTF-8 — note this does NOT mean ctx.fs.readText will succeed, since that also rejects content with a NUL byte in the first 8 KiB), supportedEncodings() (the set this deployment will actually try — use it for a picker instead of your own list), knownEncodings() (every usable encoding name), autoGuessEnabled(), recordedEncoding(sessionId, target, currentVersion?) (see above), isFsEncodingService(value, ...required?) (whether ctx.get handed you this service, by capability; name the members you will call in required, or omit it for tryDecode + decode).
tryDecode / decode accept a maxBytes option bounding a single decode, defaulting to this plugin's maxFileBytes; an oversized input comes back as a refusal (E_TOO_LARGE) instead of being fully ranked. Note there is no "unlimited" spelling: omitting maxBytes means this deployment's maxFileBytes, while an unusable value (Infinity, NaN, a non-positive number) is refused with E_BAD_ENCODING rather than silently treated as the default — otherwise the error would advise raising an argument that was just ignored.
tryDecode throws for no input, and an argument error always comes back as a refusal: a non-string encoding or displayPath, bytes that is not a Uint8Array, and an unusable maxBytes are all refused.
A mistyped opts is refused too, rather than treated as "not supplied". Only omitting the argument entirely (or passing undefined) means "use the defaults"; null, a string, a number and the like are all refused:
await fsEncoding.tryDecode(bytes); // defaults
await fsEncoding.tryDecode(bytes, {}); // defaults
await fsEncoding.tryDecode(bytes, { encoding: "gbk" }); // name an encoding
await fsEncoding.tryDecode(bytes, "gbk"); // refused: the braces were forgotten
await fsEncoding.tryDecode(bytes, null); // refused: null is not "not supplied"
This is deliberate. In JavaScript, reading a property off a non-object does not throw — it yields undefined, exactly as if nothing had been passed. So tryDecode(bytes, "gbk") used to silently drop the encoding the caller named and fall back to guessing, producing a decoding that could differ from the one requested with nothing to signal it. null is refused for the same reason: it is what a missing value looks like in JSON and database rows, so accepting it would hide the caller's bug.
Development
npm install
npm run typecheck # tsc --noEmit for src and test
npm test # run the test suite
npm run build # src/ → lib/
The suite covers the encoding algorithms themselves and drives the real filesystem and sandbox components for integration cases, including byte-exact round-trips for every supported encoding, BOM / CRLF fidelity, unmappable-character refusal, stale-write rejection and the sandbox fence.
To customize the plugin, use DSH's Creator mode for quick development.
Acknowledgements
The encoding layer began as an extraction from dsh-better-edit — BOM fidelity and the byte-exact round-trip of multi-byte encodings come from there. As a standalone plugin it was then rewritten and extended around one goal, "the bytes read back are the bytes written": the full range of Windows ANSI pages, a configurable guess list, evidence-based candidate ranking, and the ctx.fsEncoding service. Thanks to that project for the foundation.
Both plugins register read / write / edit on the scope layer, so they cannot be enabled at the same time (see Compatibility & conflicts above).
License
Comments
Comments live in GitHub Discussions. Sign in with GitHub to post or react.