Recall Benchmark — quality, measured
cachly's claim is structural: a memory that ranks by lesson quality (outcome, confidence, proven-ness, human review) surfaces the lesson that actually helps before a text-similar failed attempt. A flat-file memory — an LLM reading its own /memories directory — has none of those signals.
These numbers are reproducible from the open-source MCP package (npm run bench:gate) and a CI gate re-runs them on every commit.
What that gate cannot do: both corpora here are small — 17 lessons at home, a sample corpus externally. The two worst retrieval bugs we have shipped only appeared past a few hundred records, and this gate was blind to both. Treat it as a regression watchdog for known numbers, not as proof of quality. The honest measurement is against your own real corpus; the script for that ships in the same package (src/bench/korpus-aus-brain.ts).
Methodology
Three rankers are compared head-to-head over the same candidate set, so the comparison isolates ranking quality, not retrieval. Gold answers are the lessons that actually solved a problem; text-similar failed attempts act as distractors. Metrics are standard IR: Precision@1, Precision@3, Recall@3, MRR, nDCG@5.
| Ranker | What it models |
|---|---|
| flatfile | Naive term-overlap, no IDF, no quality signal — models an LLM reading its own /memories directory |
| baseline | Raw BM25+ keyword ranking — a solid classical retrieval baseline |
| cachly | BM25+ over a wider top-25 candidate pool, then quality-aware rerank (outcome, confidence, proven-ness, human review) |
Two corpora: a home fixture corpus (17 lessons · 13 queries) and a third-party-labeled external corpus — both fixed and versioned in the repo, runs are fully in-memory and deterministic.
Results
Measured 2026-06-07 · reproducible via npm run bench
Home fixture corpus (17 lessons · 13 queries)
| Metric | flatfile | baseline (BM25+) | cachly | vs flatfile |
|---|---|---|---|---|
| Precision@1 | 76.9% | 69.2% | 69.2% | +0.0% |
| Recall@3 | 100.0% | 96.2% | 100.0% | +4.0% |
| MRR | 87.2% | 84.6% | 83.3% | -1.5% |
| nDCG@5 | 89.9% | 88.8% | 88.1% | -0.8% |
vs. our own keyword search (BM25-style scoring — not a third-party BM25 library; the baseline is us, not the field) on this fixture set: +0.0% Precision@1 · −1.5% MRR · +4.0% Recall@3. On 17 lessons the two rankers tie on first place; the lift cachly claims is structural and only becomes visible on a corpus large enough to contain distractors.
External labeled corpus
| Metric | cachly | CI gate floor |
|---|---|---|
| Precision@1 | 78.6% | 71.0% |
| Recall@3 | 92.9% | 92.0% |
| MRR | 87.4% | 82.0% |
| nDCG@5 | 89.9% | 85.0% |
Where the lift comes from
Three mechanisms, each load-bearing — remove any one and the numbers drop:
- 1
Document-side cross-lingual expansion is disabled
Indexing a document with all ~28 multilingual synonyms of every common word (error → エラー, 错误, خطأ, …) made BM25 count them as exact matches and inflated any document containing a common word ~28×, burying topic-specific lessons. Queries still expand — a Japanese query still retrieves English lessons.
- 2
Wider candidate pool (top-25)
The reranker can only rescue a relevant lesson it can see; BM25 vocabulary mismatch sometimes ranks the right lesson 11–25. A 25-deep pool lets quality pull it back into the top 3–5.
- 3
Score compression: measured, then removed
Until 19 August 2026 we compressed BM25 scores with score^0.3 before applying the quality multiplier, and this section recommended it. On the 17-lesson fixture set it looked like the biggest single win. On 498 real lessons it halved the hit rate: 15% Precision@1 with ^0.3, 30% with ^1.0. The entire gain on the test bench was the damage on real data. We now use ^1.0. We leave this here rather than delete it, because a refuted mechanism with a number is worth more to someone building their own retrieval than three confirmed ones without.
Token cost — the other half
Recall quality is one half of the value; token cost is the other. On the same corpus we count the context input tokens a targeted recall sends per call versus the "paste everything" approach (a flat-file / CLAUDE.md dump re-sent into the prompt each call). Counted with gpt-tokenizer (cl100k_base).
| Approach | Context tokens / call |
|---|---|
| Paste everything (flat-file / dump) | 939 |
| cachly targeted recall (top-3) | ~57 |
| Context-token reduction | ~94% |
The reduction grows with the size of your knowledge base — re-sending everything scales linearly, targeted recall does not. This is retrieval efficiency only: an honest input-token number, not a blanket "X% lower bill". Output tokens, multi-turn dynamics, and semantic-cache hit rate are out of scope here.
Reproduce it yourself
The harness, both corpora, and the gate ship inside the open-source MCP package (@cachly-dev/mcp-server, src/bench/) — no network, no real Redis, deterministic:
npm run bench # head-to-head on the built-in fixture corpus npm run bench:external # third-party-labeled external corpus npm run bench:gate # CI gate: home + external, fails on regression npm run bench:cost # context-token cost vs paste-everything
CI regression gate: every commit runs both corpora and asserts cachly's metrics stay at or above committed floors. Floors move only deliberately, with a bench run in the PR — a ranking change can never silently degrade recall.
Questions about the methodology? Read how cachly's memory works or open an issue on the repo.