Skip to main content
Kodelyth ECC
architecture

The Memory That Could Never Be Recalled

BM25 scores are not normalised, so a fixed minimum score is a bet on corpus size. One memory in our store was mathematically unreachable at any size, and the search returned nothing while working perfectly.

· 5 min read

BM25 IDF falls toward zero as a term appears in more documents, while a fixed 0.5 threshold stays flat — so common terms are discarded at every corpus size.

ECC keeps a local memory store: solutions and gotchas captured from past sessions, retrieved with BM25. Search had a sensible-looking default:

function search(query, { limit = 5, minScore = 0.5 } = {}) { /* … */ }

Discard anything scoring below 0.5. Reasonable, if scores run 0 to 1.

They don't.

BM25 does not produce a probability

The IDF term is:

log(1 + (N - df + 0.5) / (df + 0.5))

N is the number of documents, df is how many contain the term. The output is unbounded above and has no fixed scale. A score of 0.5 means nothing on its own — it depends on corpus size, document length, term rarity, and query length all at once.

Now follow df upward. As a term appears in more documents, (N - df) shrinks toward zero, and the IDF shrinks with it. When a term appears in every document, IDF collapses toward zero and so does the score it contributes.

That is correct behaviour. A word in every document tells you nothing about which document you want.

But against a fixed floor, correct behaviour becomes a bug: those documents score above zero, below 0.5, and get thrown away. At every corpus size. Forever.

One memory in the store was unreachable by any query. Not ranked low — returned never.

Why nobody noticed

This is the part worth dwelling on, because the failure is silent by design.

Search returned results. Good results, usually. Nothing errored, nothing logged, no test failed. The system behaved exactly like a search engine that had simply found nothing relevant — which is indistinguishable, from the outside, from a search engine that found something and discarded it.

The absence of a result is not an event. It leaves no trace to grep for.

The fix is one line and a different idea

const RELATIVE_SCORE_FLOOR = 0.25;
 
const topScore = sorted.length ? sorted[0][1] : 0;
const floor = minScore ?? topScore * RELATIVE_SCORE_FLOOR;

Keep anything within 25% of the best match. The floor now scales with the corpus instead of fighting it. A query where everything scores 0.3 keeps its best hits; a query where the top hit scores 40 discards the noise at 2.

Note minScore ?? … rather than ||. An explicitly passed floor is still honoured exactly — our doctor command passes 0.1 and must keep getting that — and ?? preserves a deliberate 0, which || would silently replace.

Guarding against the obvious overcorrection

A relative floor has a failure mode in the other direction: always return something. If the top hit is garbage, 25% of garbage is still garbage.

So the test suite pins both ends:

test('a term in every memory is still recallable', () => {
  // df === N drives IDF toward zero. Correct BM25, fatal under an absolute floor.
  assert.ok(res.hits >= 1);
});
 
test('the relative floor does not become "always return something"', () => {
  assert.equal(store.recall('quantum mechanics unrelated').length, 0);
  assert.equal(store.recall('zzzzzqqqq').length, 0);
});

Both directions, or the fix is just a different bug.

One of those tests needed a child process with its own store directory, because corpus size is the variable under test and the shared suite store has whatever earlier tests left in it. If your relevance test shares a corpus with other tests, it is testing something other than what you think.

The general lesson

A threshold is only meaningful against a bounded scale. BM25, TF-IDF, cosine on unnormalised vectors, and most raw ranking scores have no such bound.

If you are about to write if (score > 0.5), first answer: what is the maximum? If the honest answer is "depends", your threshold is a bet on data you have not seen yet — and the day it loses, nothing will tell you.

Rank, then cut relative to the top. Or normalise first, and know what you normalised against.


Fixed in ECC 2.24.2. The store is ~400 lines of dependency-free Node in scripts/memory, and everything stays on your disk.