· 5 min read
Here is a helper that looks completely fine:
function safeReadExcerpt(p, max = 800) {
try {
const content = fs.readFileSync(p, 'utf8');
return content.length > max ? content.slice(0, max) + '\n…(truncated)' : content;
} catch { return null; }
}Read the file, take the first 800 characters. It handles errors, it truncates, it has a sensible default. It shipped and nobody looked at it again.
The cost of a fixed-size excerpt scales with the size of the file.
The numbers
| file | time | RSS growth |
|---|---|---|
| 1 MB | 1 ms | +3 MB |
| 50 MB | 20 ms | +130 MB |
| 300 MB | 196 ms | +628 MB |
Roughly twice the file size, because the bytes land in a Buffer and then again as a UTF-16 string. To produce 800 characters.
Why it mattered here
Three multipliers turned a sloppy helper into a real problem.
The dashboard is single-threaded Node, so a 196ms read blocks every other request for its duration.
It runs three times per worker — task.md, handoff.md, status.md — across
every session in the coordination directory.
And those files are written by agents, not by the dashboard. Their size is not something the reader controls. A verbose agent writing a long handoff is normal behaviour, not an attack.
A safeStat helper already existed four lines above. Nobody had reached for it.
The fix, and the detail that makes it non-obvious
function safeReadExcerpt(p, max = 800) {
let fd;
try {
const st = fs.statSync(p);
if (!st.isFile()) return null;
const window = Math.min(st.size, max * 4 + 4);
const buf = Buffer.alloc(window);
fd = fs.openSync(p, 'r');
const bytesRead = fs.readSync(fd, buf, 0, window, 0);
const content = buf.subarray(0, bytesRead).toString('utf8');
return content.length > max ? content.slice(0, max) + '\n…(truncated)' : content;
} catch {
return null;
} finally {
if (fd !== undefined) { try { fs.closeSync(fd); } catch {} }
}
}Why max * 4? This is the bit worth slowing down for.
UTF-8 encodes a character in one to four bytes. The truncation decision counts
characters — content.length > max — as it always did. If you read exactly
max bytes, a file of Japanese text gives you ~267 characters, and you would
truncate early and silently. Reading max * 4 guarantees the window always holds
at least max characters when the file has them.
Slicing a byte window directly would also cut a multi-byte character in half and produce a replacement character at the boundary. Read a generous byte window, then slice by characters.
Proving the output didn't change
A faster function that returns something different is not a fix. So we kept the old implementation and diffed the two across cases chosen to break the new one:
- empty, short ASCII, exactly 800, 801, 100k ASCII
- multi-byte short, 801 multi-byte characters
- emoji on the boundary (500 × 🎉, surrogate pairs)
- mixed scripts, a missing path, a directory
Twelve cases, all identical.
Write the test as a memory bound
The obvious regression test is a timing assertion: "must finish in under 50ms." That test is flaky on a busy CI box and tells you nothing about why.
The property we actually care about is that the read does not scale with the file. So assert memory:
const big = 'a'.repeat(64 * 1024 * 1024);
withFile(big, (file) => {
const before = process.memoryUsage().rss;
const out = safeReadExcerpt(file);
const grewMb = (process.memoryUsage().rss - before) / 1048576;
assert.equal(out.length, 800 + '\n…(truncated)'.length);
assert.ok(grewMb < 8, `read grew RSS by ${grewMb.toFixed(0)}MB for an 800-char excerpt`);
});Memory growth is a property of the algorithm, not of how loaded the machine is.
And we checked the test actually catches the bug — a test that passes against the broken version is decoration. Against the old implementation, 64MB grows RSS by 97MB and the assertion fails. Against the new one, 0MB.
The general version
When a function returns a bounded amount of data, its cost should be bounded too. If it isn't, you have an input-size dependency hiding in a constant-size API.
The shape to look for is any readFileSync followed by a slice, split('\n')[0],
a header parse, or a JSON.parse you only need one field from. Each is a
constant-size answer paid for at input-size cost.
Fixed in ECC 2.24.8. The dashboard runs on localhost, reads only local files, and sends nothing anywhere.