Reproducible retrieval benchmarks · including the results that went against us
We are not trying to be better. The claim is much cheaper and usable at context lengths where the alternative cannot run at all. Below is everything three people asked us to measure against that claim — including the results that went against us.
Every number is reproducible from github.com/staccDOTsol/supercontext. Model: gemini-2.5-flash via OpenRouter. Single runs at temperature 0.
13.7×
faster to first answer at 1M tokens
98×
fewer prompt tokens, same answers
24/24
answered at 1M–5M where the model refused outright
0–4
lost to bge-base on nDCG — the trade, priced in
k=8
the default we shipped dropped the answer under decoys — a config bug, since fixed
J=0
a question sharing zero vocabulary with its answer defeats the ranker — and more context does not help
The first three are the pitch. The fourth is the trade we made on purpose and would make again — quality parity was never claimed. The fifth is the one real defect these benchmarks found — a settings value, not an architecture, since raised in production. The sixth is the measured limit of the whole recall-recovery story: leCore's ranker is lexical, and at exactly zero question–answer overlap no top_k recovers it (§08).
01done·this is the pitch
A FAISS index does not exist until an embedding model has read every document. Latency-only comparisons skip that step. This measures model load + index build + one query — what a user actually waits for on a cold corpus.
Time to first answer, 1M-token corpus (11,940 docs) — lower is better
HNSW makes queries sub-linear, exactly as Shaw said — but it does not touch the dominant term. Graph construction is ~0.3 s on top of ~182 s of embedding.
| corpus | docs | leCore index | flat index | HNSW index | leCore query | HNSW query |
|---|---|---|---|---|---|---|
| 50k | 597 | 0.05 s | 10.60 s | 10.60 s | 2.39 ms | 31.23 ms |
| 200k | 2,388 | 0.57 s | 42.38 s | 42.41 s | 0.86 ms | 31.03 ms |
| 1M | 11,940 | 13.31 s | 182.03 s | 182.28 s | 4.86 ms | 35.83 ms |
02done·we lose
Shaw asked us to stop quoting our own test and run the standard one. We did, then ran three real embedding models through the same harness so it isn't ours-versus-literature. bge-base beats leCore on all four tasks. "BM25 is a floor" was correct.
nDCG@10 by task — higher is better
SciFact
ArguAna
NFCorpus
SCIDOCS
| system | SciFact | ArguAna | NFCorpus | SCIDOCS | FiQA |
|---|---|---|---|---|---|
| leCore BM25 + expand | 0.6705 | 0.4314 | 0.3167 | 0.1577 | 0.2378 |
| all-MiniLM-L6-v2 | 0.6451 | 0.5017 | 0.3159 | 0.2164 | — |
| e5-large-v2 | 0.7221 | 0.4642 | 0.3715 | 0.2050 | — |
| bge-base-en-v1.5 | 0.7404 | 0.6375 | 0.3735 | 0.2172 | — |
removed from ship
Our own dense VSA encoder scored 0.4160 on SciFact and made the hybrid worse at every fusion weight (0.05 → .6596, 0.15 → .6307, 0.30 → .6057, against .6705 for BM25 alone). We were shipping the worst of three configurations. It has been removed.
03done·this is the pitch
Scored three ways, never two: HIT, MISS, and UNRUNNABLE — the endpoint refused the request. Collapsing a refusal into a wrong answer would let a memory system claim it "beat" a model that was never allowed to compete.
Cells answered correctly, 900k–5M tokens, n=6 per tier. Beyond 1M the control returns "The input token count exceeds the maximum number of tokens allowed." Tokens the model reads stay flat: a 60k corpus → ~1,458 read; a 5M corpus → 3,401 read.
Totals across all 24 cells
Bars are prompt tokens spent, linear scale — 98× fewer for 24 hits against 5.
04done·negative·we overstated it — self-audited 2026-08-14
Two identical 303M language models — same seed, same schedule, same step count, RoPE in both arms; only the attention operator differs. We published this as 41.5% worse for HRR. Then we re-ran it against our own methodology, and 41.5% became 32.7%. About a fifth of our headline negative was our configuration, not the operator. The verdict is unchanged — the operator still loses — but the number we shipped was inflated, so here is the corrected one.
Fairness rerun · validation loss after 7,629 steps / 999,948,288 tokens, each arm at its own tuned learning rate — lower is better
bits/byte proxy: softmax 4.5436 · gated-HRR 6.0312. Both arms COMPILE=0, micro 16 / accum 8, same 7B-token bin, same seed, identical step count.
32.7%
worse for HRR, with every confound removed — down from the 41.5% we published. The prior measurement at 12M params was a ~14% gap; scale still widens it, just by less than we said.
3.0×
slower per token, not 11× as previously published — 151,269 vs 50,797 tok/s, after wiring the repo's own Triton scan.
10.2×
from the Triton scan alone, which also cut memory 74.9 → 41.8 GB.
Auditing our own headline negative — five confounds, all ours
correctionThe published pair had five asymmetries between the arms, and every one of them was introduced by us: torch.compile on for softmax and off for HRR; a learning rate tuned for neither and inherited from softmax; and a mismatched micro-batch / gradient-accumulation split that, because the loader strides by micro×accum, also desynchronised the data stream — the two arms were not reading the same tokens in the same order. The rerun sets COMPILE=0 on both, micro 16 / accum 8 on both, the same 7B-token bin, the same seed, and lands both arms on exactly 7,629 steps and 999,948,288 tokens.
41.5% → 32.7%
the HRR penalty as published, and as re-measured. The absolute gap goes 1.3236 → 1.0311 nats — 22% of it was artifact.
0.333
val loss HRR recovered from fair configs plus a tuned learning rate — 4.5138 → 4.1805. Softmax moved 0.041 (3.1902 → 3.1494).
still behind
HRR loses by a third under conditions we can no longer blame. The correction is to the magnitude, not the direction.
New finding: HRR's optimal learning rate is 3.3× softmax's
measuredSix learning rates, one per arm, four A100s, identical budget — 1,148 steps / ~150.5M tokens each, same schedule and warmup. HRR's best is 2e-3; softmax's is 6e-4. At the same budget, running HRR at softmax's 6e-4 costs it 1.14 nats of validation loss. That is what the original run did — so part of what we published as an architecture penalty was a hyperparameter one.
| HRR learning rate | val loss @ 150.5M tok | bpb proxy | note |
|---|---|---|---|
| 3e-4 | 7.4070 | 10.6861 | |
| 4.5e-4 | 7.5033 | 10.8250 | worst of six |
| 6e-4 | 7.4224 | 10.7083 | softmax's LR ← the published run |
| 9e-4 | 6.4272 | 9.2725 | non-monotone: 1.2e-3 is worse |
| 1.2e-3 | 6.8328 | 9.8577 | |
| 2e-3 | 6.2822 | 9.0633 | best — and the largest tested |
Read this honestly: the minimum sits at the edge of the swept range, so 2e-3 is a floor on HRR's optimum, not a located one — the true optimum may be higher and we have not measured it. The curve is also noisy at the small end (4.5e-4 scores worse than 3e-4, 1.2e-3 worse than 9e-4), which is one sweep per point at ~150M tokens, not a seed-averaged one. All six are also mid-decay: the schedule was written for 200M tokens and every point was cut by the same 1.5-hour wall clock at ~150.5M, so these are ranking numbers, not converged ones — which is why 6e-4 reads 7.4224 here and 4.1805 is nowhere near it (that arm ran 6.6× longer). The rerun used 2e-3 because it was the best rate measured, not because it was proven optimal.
caveats we own
The rerun closed the compile, learning-rate and batch-split asymmetries. Two caveats survive it and one is new. Both arms are still undertrained — ~1B tokens is roughly 1/6 Chinchilla-optimal for 303M, so this is the gap at 1B tokens, not the asymptotic gap. The rerun used a 7B-token bin where the original used 5B, so its validation tail is not the original's: compare the 3.1494 and 4.1805 to each other, and treat every cross-run absolute as indicative only. The internal gap is the experiment; the softmax arm moving 0.041 between runs is the size of that bin effect on this axis. A 6B-token Chinchilla-scale pair is running now — softmax past 2.59B of 6B tokens at the time of writing, HRR configured at lr 2e-3. It is in flight, not a result; nothing on this page depends on it and nothing should be read into it until both arms finish.
Follow-up: LM NIAH on this pair — a minor negative result
negativePlant "The secret code is XXXX" at start/middle/end of filler and score the logit margin of the true code token against a random wrong one — >0 means the needle won. 3 seeds per cell, both arms RoPE, training block 1024. Softmax retrieves inside its training length, decays at 2048, and is dead by 4096–8192. HRR never localizes at any length. Its lone positive cells (+0.16 at 8192) are position-insensitive — near-identical at start, middle and end (0.159 / 0.159 / 0.151) — a uniform bias toward the answer token, not retrieval. At 8192 the honest read is both arms fail.
| context | softmax margin · start / mid / end | HRR margin · start / mid / end |
|---|---|---|
| 512 | +4.84 +4.82 +5.79 | −0.05 −0.05 +0.06 |
| 1024 | +1.33 +4.85 +4.71 | −0.14 −0.14 −0.10 |
| 2048 | +0.13 +1.72 +0.47 | −0.28 −0.28 −0.28 |
| 4096 | +0.13 +0.07 +0.08 | −0.59 −0.59 −0.60 |
| 8192 | −0.31 −0.30 −0.37 | +0.16 +0.16 +0.15 ← bias, not retrieval |
Mandatory caveat: these cells were scored on the published pair — hrr val 4.5138 vs softmax 3.1902, 41.5% worse — so they confound the operator with plain LM quality: a worse language model fails NIAH for reasons that have nothing to do with attention. The fairness rerun above did not fix that. It was the matched-loss attempt, and it did not produce matched loss: 4.1805 vs 3.1494 is still 32.7% apart. So the confound stands, narrower. We have not re-scored this table on the rerun checkpoints — that is unmeasured, not unchanged — and the caveats box above applies in full.
Sources: the arms' own final JSONs — results/runpod_303m_rescue/final_softmax.json (3.149412655830383) and the HRR arm's out/final_hrr.json on its Vast box, mirrored to results/hrr_vs_softmax/303m_hrr_vast/ (4.180486848950386). Sweep logs: sweeps303/sweep_hrr_*.log, six files, one per rate. The superseded pair is 303m-results/final_softmax.json and is kept, not deleted. Checkpoints pushed to staccs/lecore-303m-rerun.
05done·found a real config bug
Cotten asked for this twice: plant decoys that carry the full query vocabulary but are not the answer. He was right — at the shipped top_k=8 it breaks, and that turned out to be a settings value rather than an architecture problem.
On the zero-decoy corpus a dumb exact-substring counter — no idf, no length normalisation, no saturation — scores precision@1 = 1.00, identical to leCore. A benchmark that a substring counter aces is not measuring a ranker. It is measuring the corpus's lexical confusability, and ours was ~zero.
Does the answer survive decoys? · 200 docs, 32 seeds/cell, k=8 — the default we shipped · scattered decoys, the hard case
0 decoys
2 decoys
5 decoys
10 decoys
25 decoys — recall breaks here
50 decoys
100 decoys
Precision@1 dies at two decoys — 1.00 → 0.19. The surviving envelope is recall, not precision: the answer still reaches the model's context up to ~10 competing passages (recall@8 = 0.91) and breaks between 10 and 25 (0.22). Past that, at this k, the answer is not in the context at all — so no amount of model quality recovers it, and only a bigger top_k does. Mean margin goes negative (−0.44 at 25 decoys): the best decoy outscores the needle.
We were never trying to out-rank bge-base. So the right answer to "the needle fell to rank 12" is not a better ranker — it is a bigger top_k. Recall is fully recoverable at every decoy level measured, and the k required is roughly the number of competing passages.
top_k needed to bring recall back to 1.0 — swept over 8, 12, 16, 24, 32, 48, 64
25 competing
k = 24
recall 0.22 at k=8 → 1.00
50 competing
k = 48
recall 0.12 at k=8 → 1.00
100 competing
k = 64
recall 0.25 at k=8 → 1.00
and we can afford it — prompt tokens on the 60k–500k sweep, linear scale
Raising k from 8 to 64 is 8× the retrieved tokens — 68,057 → ~544k on the 60k–500k sweep — which is still 12.3× fewer than the control's 6,676,718. Precision@1 was never the product. The product is that the answer is in the context at a price the alternative cannot match.
the honest limit
This is a scaling relationship, not a free lunch. Token cost grows with the number of confusable passages — a property of the query rather than of corpus size, but whether confusability grows with corpus size is not measured here, and it is the thing that would break this. That bench has since run on real corpora — §07: recall recovers within k ≤ 64 on all three BEIR sets, and at 2.68M docs the crossover survives at k = 128 — exactly the production cap, with zero headroom. And §08 now marks the lever's hard boundary: k-recovery requires at least one shared content word — at measured zero question–answer overlap, recall stays flat no matter the k.
what changed in prod
The shipped default was top_k=8 with a hard cap of 32 — both too low; the cap alone foreclosed the 100-decoy case. Fixed and deployed: default 16, cap 128, verified live.
leCore's BM25 is genuinely ranking rather than guessing. tied@top is ≈1.0 for leCore against 5–85 for the substring floor, so the floor's occasional wins are argsort index luck while leCore's are real separation. Clustering also helps, because a clustered decoy lands inside the answer's own chunk instead of competing with it as a separate document.
100 decoys, half-clustered — leCore vs the substring floor · 4.7× separation
mean rank of the answer — lower is better · scale 0–32
recall@20 — higher is better · scale 0–1.0
the honest reading
Our headline needle number was inflated by vocabulary isolation. A bag-of-words model has no representation for contains the answer versus is about the answer. So precision@1 is not something BM25 can be made to deliver — and it is not something we need. The adversarial bench did not kill the thesis; it killed the precision-retrieval thesis and revealed a viable cost/recall tradeoff in its place. The claim is recall@k, bought with a bigger top_k on the one axis we are cheap on — never precision@1.
Two arms do hold precision under exactly these decoys: the pure VSA path, profiled next — §06 is its first outright win, published beside its losses — and the quantized semantic encoder, measured under the same protocol in §10.
06done·its first outright win·wrapper was 4,200× the floor — fix deployed & re-measured
This bench exists because a critic asked. Cotten, publicly: "rewrite the benchmark so it never skips to BM25; always profile the VSA over 1024d and pay attention to _Index — btw the _Index wrapper isn't optimized." So the pure VSA path is now a first-class, never-skipped row, and every number ships regardless of direction. He was right about the wrapper. And the same profile handed the VSA — the arm §02 removed from the hybrid — its first outright win, on the exact adversarial bench that broke BM25 in §05.
Where the time goes — 100k docs × 1024d, cosine top-k, Apple M5 Pro CPU · log scale (1 ms → 30 s), lower is better
The math is milliseconds. The shipped wrapper rebuilt the HoloForest on every query — the one-time forest build is 19.6–20.4 s, essentially the whole 20.5 s, ~4,200× the floor. Cotten's 518 ms sat 40× below that deployed number: his complaint was charitable. The fix was a cache — build the forest once, on the first vec-path query, invalidated by any write — and it shipped and deployed the same day (commit ee813a7), then re-measured: local end-to-end vec recall 2.09 s → 2.4 ms at 5k items, 21.82 s → 5.9 ms at 20k; live, a 5,000-item forest context answers at 117 ms p50 (~100 ms of that is network) after a one-time 3.5 s first-query build — every query used to pay that 3.5 s. The BM25 path did not regress: 116–123 ms warm vs 102/97 ms the same morning, within jitter. Cached and uncached return identical ids and scores by construction, verified exact-equal at 5k and 20k. Encoder cost for scale: ~0.8 ms per 600-char chunk (0.72–0.87 across thread arms, ~1,150–1,380 chunks/s), query encode 0.08 ms.
Same protocol as §05 — 800 lines, 32 seeds per cell, k=8, scattered decoys the hard case. The idf-weighted dense superposition holds rank under exactly the term-collision that inverts BM25's scoring.
P@1 · 2 decoys, scattered
VSA margin +0.059 — the needle outscores the best decoy
R@8 · 25 decoys, scattered
mean rank 7.4 vs 12.0 — 3× more often in a k=8 context
P@1 · 25 decoys, clustered
clustered VSA P@1 = 1.00 at 2, 10 and 25 decoys alike
| scattered decoys | system | P@1 | R@8 | mean rank | mean margin |
|---|---|---|---|---|---|
| 0 | leCore VSA | 1.00 | 1.00 | 1.00 | n/a |
| leCore BM25 | 1.00 | 1.00 | 1.00 | n/a | |
| 2 | leCore VSA | 0.81 | 1.00 | 1.22 | +0.059 |
| leCore BM25 | 0.19 | 1.00 | 2.09 | −0.014 | |
| 10 | leCore VSA | 0.09 | 1.00 | 3.22 | −0.273 |
| leCore BM25 | 0.09 | 0.91 | 5.69 | −0.348 | |
| 25 | leCore VSA | 0.03 | 0.69 | 7.41 | −0.460 |
| leCore BM25 | 0.03 | 0.22 | 11.97 | −0.444 |
The honest read of the win: at 10+ scattered decoys the VSA's precision@1 collapses too — 0.09 → 0.03. What it keeps is recall: at 25 scattered decoys the answer is still inside a k=8 context 69% of the time against BM25's 22%, before any of §05's k-raising. Clustered decoys cost it nothing at all: P@1 = 1.00 at 2, 10 and 25, where BM25 falls 1.00 → 0.72 → 0.25.
the losses, republished unfiltered
This is not "the VSA is good now." Its ranking generality is unchanged: SciFact nDCG@10 0.4160 against leCore BM25's 0.6705 and bge-base's 0.7404 — the same score that got the arm removed from the hybrid in §02. And at zero lexical overlap it is still not retrieval: J=0 recall@64 = 0.0964, reproduced here to the fourth decimal against §08's recorded run, position skew and all.
BM25 — anchored recall
Any shared content word and it beats bge-base at deployed k (§08, J=.257: .958 at k=8) — on the cost axis nothing else touches (§01, §03).
VSA — decoy resistance
Holds the needle under the term-collision that inverts BM25 — P@1 .81 vs .19 at two scattered decoys, 1.00 clustered throughout (this section).
quantized bge — zero overlap
The only system not flat at J=0 (.615 at k=8 → .846 at k=64, §08). Since measured — §10: quantized with leCore's own machinery to within .002 nDCG of full bge, priced on CPU, shipped env-gated (default off). The VSA still beats it at 25 scattered decoys (R@8 .69 vs .59).
Method: local Apple M5 Pro (18 cores), Python 3.14.6 / numpy 2.5.0, $0 spent, wall clock 137 s. Load average was 5–6 throughout and is recorded per-section in results.json — a timing-contamination flag we publish rather than hide. Two thread arms, pinned single-thread and free, both reported; the ladder above uses the free arm (the single-thread floor was 3.7 ms, the single-thread wrapper 21.3 s — same shape). At 1k docs the auto index is exact and the wrapper is 2–4 ms; at 12k docs the forest kicks in and the wrapper is already 1.7–2.3 s. The rebuild dominates everywhere the forest exists.
07done·retires the §5 caveat·knife-edge at 2.68M docs
§5's decoys were templated strings we wrote ourselves, which left one open question: does recall recovers with k survive corpora where the confusable passages are real? So: three BEIR corpora with real judged queries, leCore BM25 against GPU-embedded baselines in the same harness, hit@k swept from 1 to 200. The question is the crossover: at what k does leCore's hit rate reach what bge-base delivers at k=10? Answer: k = 32–64 on all three, at 3.2–6.4× the retrieved tokens. And below, this section's own open question re-run at scale: the same crossover at 2.68M documents lands at k = 128 — exactly the production cap.
hit@k vs k — SciFact, the hardest crossover of the three · higher is better
k — passages retrieved per query · SciFact, 5,183 docs, 300 judged queries
leCore reaches .893 at k=64, clearing bge-base's .883 at k=10 — on the corpus where that took the most context. Real corpora land at gentler k than §5's worst-case templated decoys, where the required k was roughly the full count of competing passages.
| corpus | docs / queries | bge hit@10 | leCore hit@10 | crossover k | leCore hit@k | token cost |
|---|---|---|---|---|---|---|
| SciFact | 5,183 / 300 | 0.883 | 0.820 | k = 64 | 0.893 | 6.4× · ~17,857 tok/q |
| NFCorpus | 3,633 / 323 | 0.746 | 0.681 | k = 32 | 0.755 | 3.2× · ~9,725 tok/q |
| ArguAna | 8,674 / 1,401 | 0.884 | 0.685 | k = 48 | 0.894 | 4.8× · ~10,431 tok/q |
| NQ (at scale, below) | 2,681,468 / 3,452 | 0.781 | 0.475 | k = 128 — the cap | 0.784 | 12.8× · ~14,142 tok/q |
raising k helps them too
Do not read this as parity. Give bge-base the same k=64 and it scores .950 / .854 / .991 on the three corpora — more context helps everyone. What this bench establishes is that the recall-recovery lever survives real corpora, and what it costs: 3.2–6.4× tokens to match bge's k=10. The sellable number is cost per hit without a GPU or an embedding pass, not "we catch up."
inside the shipped cap
Production now defaults to top_k=16 with a cap of 128 — every crossover k measured here (32–64) sits inside it. The old cap of 32 would have foreclosed SciFact's k=64 outright. The 2.68M-doc run below lands exactly on the cap — inside it by zero.
Method and the honest fine print: clean Vast RTX 4090, load average 1.5, embedding on GPU. bge-HNSW matched exact flat search on every corpus — identical on SciFact and ArguAna, within 0.007 hit on NFCorpus — so there is no ANN recall loss to exploit at these sizes. And on this box the embedders' index build was as fast or faster than our BM25 build (bge 11.3–15.5 s GPU vs leCore 15.8–50.5 s CPU): the §01 cost-to-first-answer win assumes no GPU is present. One boundary on everything above: the ranker is lexical — §08 measures the case where a question shares zero vocabulary with its answer, and the k lever does not survive it.
The three corpora above top out at 8,674 documents — §12's stated "not built" was whether crossover k grows with corpus size. It does. On Natural Questions — 2,681,468 docs, 3,452 judged queries, same harness, same crossover definition — bge-base's hit@10 is .781, leCore's is .475, and leCore does not reach .781 until k = 128 — exactly the production cap (hit@128 = .784, at 12.8× the tokens, ~14,142 tok/query). The trend across everything measured: 3.6–8.7k docs → k = 32–64; 2.68M docs → k = 128. Zero headroom left.
hit@k vs k — Natural Questions, 2.68M docs · higher is better
k — passages retrieved per query · Natural Questions, 2,681,468 docs, 3,452 judged queries
leCore's curve meets bge-base's k=10 bar at the last point on the grid — .784 at k=128, against the bar of .781. One grid point later and there would have been no crossover to report inside the shipped cap. Past a few million documents the lexical lane cannot reach embedding-grade hit rates at sane k — that boundary now has a number, and it is where the semantic lane (§06's third lane, the quantized-bge stage — since measured in §10 and shipped env-gated, default off) takes over.
| system | hit@10 | hit@128 | one-time index | of which embedding | query |
|---|---|---|---|---|---|
| leCore BM25 (O(tokens) rebuild) | 0.475 | 0.784 | 309.8 s CPU | 0 s | 24.16 ms |
| bge-base · FAISS flat | 0.781 | 0.960 | 2,932.9 s | 2,927.6 s GPU | 6.37 ms |
| bge-base · HNSW | 0.774 | 0.950 | 3,668.7 s | shared with flat | 0.16 ms |
| MiniLM · FAISS flat | 0.683 | 0.931 | 976.8 s | 973.1 s GPU | 4.82 ms |
the index still costs nothing to embed
2.68M documents indexed in 309.8 s on CPU with zero embedding against bge-base's 2,932.9 s of GPU work — the §01 cold-start win holds at three orders of magnitude more docs. But the recurring side turned: matching bge's hit@10 now takes every one of the 128 passages the production cap allows.
the amortization math, from this run's own costs
leCore's premium at crossover is (128 − 10) × 110.5 ≈ 13,037 extra tokens per query, every query; bge pays its 2,927.6 s encode once. Break-even: Q* = C_embed ÷ (Δtok × p_tok) ≈ 250 queries at $0.40/hr GPU and $0.10/M input tokens, ~25 at $1.00/M. The prices are assumptions, the inputs are measured — and the honest reading is that at this scale the embedder amortizes within tens-to-hundreds of queries. The cost win at multi-million-doc scale is confined to cold-start, low-volume, or no-GPU workloads.
The fine print, all of it. The shipped BM25 build cannot run at this scale: its postings loop is O(vocab × N) — ~10¹² iterations at 2.68M docs — and does not terminate. The bench used an O(total tokens) rebuild whose postings are bit-identical (same idf, same per-(term,doc) weight expression, same ordering; scores() output identical). The 309.8 s is that rebuild, and the shipped build's non-termination is itself a reported finding — the fourth config-class engineering defect these benches have found, after top_k=8, the cap of 32, and §06's per-query forest rebuild. The box was not clean: RTX 4090, load average 6–24 across the arms, recorded per-arm in the JSON — timings carry that flag; hit rates are load-immune. e5-large was skipped at this scale (3× bge's encode cost for the same question), and only NQ has run — HotpotQA at 5.2M docs is the next tier and is not measured.
08done·negative at J=0 — the honest limit
The loss, plainly: leCore's ranker is lexical. A question sharing zero vocabulary with its answer defeats it, and buying more context does not help. 64 needle/question pairs per overlap level, with the zero enforced in code, not by eyeballing — stemmed content-token sets must not intersect, no two content tokens may share a 4-character prefix, no shared numbers. Needles were embedded in ~222k tokens of real fineweb prose (~1.9k chunks), 3 positions × 2 seeds, through the deployed 600-char chunker. At measured Jaccard 0.000, leCore BM25 goes R@8 = .005 → R@64 = .023 — flat. Unlike §05 and §06, raising k recovers nothing, because there is no lexical anchor to rank on.
recall@k at measured zero question–answer overlap · higher is better · each group is one system swept over k
leCore BM25 · J = 0.000 — flat, k recovers nothing
leCore VSA · J = 0.000 — argsort luck, position-skewed
bge-base · J = 0.000 — the reference struggles too
leCore BM25 · J = 0.257 (~one word in four) — ahead of bge at every k
The measured boundary: the cliff is at exactly zero. One shared content word in four (measured J = .257) restores full recall at the deployed k — .96 at k=8, 1.00 at k=64 — ahead of bge-base at every k. And even the semantic reference struggles at true zero: bge-base reaches only .85 at k=64, not 1.0.
| measured overlap | system | R@8 | R@16 | R@32 | R@64 |
|---|---|---|---|---|---|
| J = 0.000 | leCore BM25 | 0.005 | 0.005 | 0.008 | 0.023 ← flat |
| leCore VSA | 0.008 | 0.023 | 0.052 | 0.096 ← argsort luck | |
| bge-base | 0.615 | 0.706 | 0.784 | 0.846 ← the reference struggles too | |
| J = 0.257 (~25%) | leCore BM25 | 0.958 | 0.987 | 0.995 | 1.000 ← ahead of bge at every k |
| bge-base | 0.881 | 0.913 | 0.958 | 0.968 | |
| J = 0.476 (~50%) | leCore BM25 | 1.000 | 1.000 | 1.000 | 1.000 |
| bge-base | 0.942 | 0.958 | 0.966 | 0.976 |
the k lever has a boundary
This bounds §05 and §07's recall-recovers-with-k story: the lever works when any lexical anchor exists, and does not exist at zero. Also honest: leCore's VSA arm's .10 at k=64 is position-skewed — .16 / .10 / .03 at start / middle / end — which is argsort luck, not retrieval. The changelog named this exact case as the likeliest failure before the bench was run; it is published here either way.
one word in four is enough
The failure needs exactly zero. At measured J = .257 leCore BM25 is .958 at k=8 → 1.000 at k=64, ahead of bge-base at every k, and at J = .476 everything saturates near 1.00. Real questions that share literally no content word with their answer are a corner of the space — but it is a corner we cannot serve, and now it is measured.
Method and the caveat to carry: pairs generated by Qwen2.5-7B-Instruct, then gate-enforced with rejects counted and published — at the zero level 65 rejected for stem overlap, 15 for shared prefixes, 8 for shared digits; 0 needles destroyed by chunking; 384 trials at J=0 and 378 at each other level. The no-shared-digit rule sometimes forces zero-overlap questions to alter quantities (a "50 new sensors" needle asked about as "30 extra devices"). That is fine for retrieval scoring — only chunk identity is scored — but the J=0 pairs are not quotable as QA pairs.
This gradient is not a synthetic artifact. §09 reproduces it on real code: paraphrased questions about a 47k-token repo retrieve perfectly, and the same question type collapses to .30–.40 on a 4.3M-token repo — while identifier-anchored questions stay the lexical ranker's case throughout. And the one system not flat at J=0 is now leCore's own lane: §10 quantizes bge-base with our machinery — recall@64 at J=0 identical to full bge — and prices exactly where it can afford to run.
09done·mixed — §08's gradient, on real code
The product's actual emerging user is an agent throwing a codebase at zoo_ask — bind the repo over MCP, then ask questions about it. This bench measures that path, factually, for the first time. Three Python repos at pinned tags — requests v2.32.3 (~47k tokens, 18 files), fastapi 0.115.12 with its tests (~625k, 360 files), django 5.1.7 (~4.3M, 2,107 files) — and 82 questions generated by Qwen2.5-7B against sampled ground-truth functions, gated by an answerable-from-chunk judge, in three types: identifier-anchored (must name the function), behavioral paraphrase (programmatically gated to contain no identifier from the code — the semantic stress), and cross-file ("what calls X and what happens to its return value"). Gold answers are (path, line-range) regions, so the production 600-char chunker and function-level chunking score against the same truth.
The claim shape this backs: identifier and cross-file questions are the lexical ranker's case, at zero embedding cost — cross-file asks go .90 at k=8 → 1.00 by k=64, identifier .73 → .85, and leCore indexes the 4.3M-token django tree in ~3.0 s on CPU where bge-base needs ~40–42 s of GPU embedding per chunking. Paraphrase questions need a semantic stage — and bge-base is not dominated here, stated plainly below.
hit-any@k by question type · pooled over the three repos · 600-char chunking · higher is better
leCore · cross-file — the lexical ranker's best case
leCore · identifier-anchored
bge-base · paraphrase — bge wins this outright
leCore · paraphrase — the semantic stress, and our loss
Identifiers are lexical anchors and the ranker treats them as such: the two anchored question types climb into the .85–1.00 band inside the deployed k range, for zero embedding work. The paraphrase line is the honest one: pooled it stalls at .50 at k=64 while bge-base reaches .69 — and the pooled number hides a gradient, drawn next.
paraphrase questions, hit-any@64 vs repo size · 600-char chunking · §08's gradient on real code
requests · ~47k tokens
fastapi · ~625k tokens
django · ~4.3M tokens
Same shape as §08's overlap cliff, produced by real code instead of a synthetic gate: when a question shares no identifier with its answer, recall decays with corpus size, and at 4.3M tokens k=64 recovers only .30 (leCore, from .10 at k=8). bge-base decays too — .70, not 1.00 — but it is not dominated: it wins paraphrase outright, and under function-level chunking it also edges identifier asks (.88 vs .81 pooled at k=64). The sell here is cost-per-hit and zero-embed indexing, never better recall.
| system | chunking | question type | h@8 | h@16 | h@32 | h@64 |
|---|---|---|---|---|---|---|
| leCore BM25 | 600-char | identifier | 0.73 | 0.81 | 0.81 | 0.85 |
| 600-char | paraphrase | 0.31 | 0.35 | 0.42 | 0.50 ← the semantic stress | |
| 600-char | cross-file | 0.90 | 0.93 | 0.97 | 1.00 | |
| function | identifier | 0.69 | 0.73 | 0.81 | 0.81 | |
| function | paraphrase | 0.31 | 0.38 | 0.46 | 0.54 | |
| function | cross-file | 0.83 | 0.90 | 0.93 | 1.00 | |
| bge-base | 600-char | identifier | 0.62 | 0.69 | 0.69 | 0.81 |
| 600-char | paraphrase | 0.42 | 0.46 | 0.54 | 0.69 ← bge wins this outright | |
| 600-char | cross-file | 0.97 | 1.00 | 1.00 | 1.00 | |
| function | identifier | 0.81 | 0.85 | 0.85 | 0.88 ← edges leCore's 0.81 | |
| function | paraphrase | 0.54 | 0.58 | 0.58 | 0.58 | |
| function | cross-file | 1.00 | 1.00 | 1.00 | 1.00 |
Cross-file scored strictly — definition and caller both in the top-k — at k=64: leCore .77 (600-char) / .83 (function); bge .97 (600-char) but only .67 (function). Chunking is a real variable, not a footnote: function-level chunking helps bge on identifier asks (django h@8 .80 vs .40) and mildly helps leCore's paraphrase, while 600-char windows serve strict cross-file better under bge.
the anchored case costs nothing to index
leCore builds the django index — 4.3M tokens, ~30k chunks — in ~3.0 s on CPU; bge-base pays ~40–42 s of GPU embedding per chunking before the first query. On the questions agents actually ask by name — identifier and cross-file — that free index reaches .85–1.00 inside the deployed k range. This is the zoo_ask bind-your-repo path, measured — the MCP's real user, not a proxy corpus.
at scale, an identifier anchors every mention
The fastapi curiosity: identifier asks plateau for both systems — leCore holds at .70 from k=16 through 64 under both chunkings; bge sits at .60–.70 through k=32 and reaches only .80 at 64. Our hypothesis, not separately verified: fastapi ships its tests, and test files citing the name crowd the definition out of the top-k — the ranker finds every mention, not the one that matters.
Method and every caveat: one Vast RTX 4090 (~55 min, $0.53 total across three boxes, two of which never ran), destroyed after. n = 6–10 questions per cell — single cells carry ~±0.15 noise, so quote the gradients and pooled rows, never one cell. All three repos are Python. Questions are one model's style (Qwen2.5-7B; the answerable gate passed ~60–80% of candidates). bge truncates chunks at 512 tokens — its real behavior — and function chunks were capped at 6,000 chars. The BM25 postings build was re-ordered doc-major to O(total tokens) — the same rebuild class §07's 2.68M-doc run needed — with score bit-identity vs the shipped build asserted on 323 chunks × 10 probes. Source: results/code_bench/ — questions.json with gold regions, results.json, summary.json, corpus_stats.json.
10done·one negative inside — published
§08 ended at the cliff: at zero question–answer overlap only bge-base retrieves anything, and §06 named a quantized-bge semantic stage as the third lane — staged, not shipped, unmeasured. This is the measurement. We ran leCore's own requantize machinery (holographic_refactor.quantize_group, a budgeted per-tensor bit ladder with embedding cosine as the fidelity metric) over bge-base-en-v1.5, then made the result take every exam family it will face: MTEB before/after on four tasks — GPU and CPU — semantic NIAH before/after, FAISS flat + HNSW, §05's adversarial protocol with the quantized encoder as a fourth arm, and cost-to-first-answer against leCore BM25 on CPU only.
the artifact — budgeted requant, 1% encoder-fidelity budget · ladder 3/4/5/6/8 bits · group 64 · 126 s
7.62
mean bits over the quantized set
132 MB
packed estimate, vs 438 MB fp32
.9894
holdout embedding cosine (calib .9900)
11 / 73
tensors refused every rung — stayed fp32
The encoder resisted. Layers 0–1 took 3-bit wholesale (cosine ≥ .995); the budget then walked the ladder up through the middle of the stack, and 11 tensors — clustered in layers 7–9 — refused even 8-bit. The same machinery pointed at a 9B decoder (a separate, still-unfinished run) was landing whole tensors at ~3 bits/weight: encoder embedding geometry tolerates far less weight noise than decoder next-token loss. Uniform arms for calibration: 4-bit everywhere is 93.4 MB at SciFact .72842 — 98.4% of full, the sweet spot only if a packed runtime existed; 3-bit everywhere is 82.7 MB at .6922, visibly damaged. Shipped conclusion: the budgeted artifact. Home: huggingface.co/staccs/lecore-bge-assimilated.
nDCG@10, full vs quantized — the bars are the point: you cannot see the difference
SciFact Δ −.0016
ArguAna Δ +.0000
NFCorpus Δ −.0019
SCIDOCS Δ −.0020 — the worst delta of the four
Worst delta across the four tasks: .0020 (SCIDOCS); ArguAna is a rounding hair higher. The CPU arm re-ran SciFact and NFCorpus and scored bit-identical to the GPU rows. Semantic NIAH at J=0: recall@64 = .8464 for full and quantized — identical (recall@8 .615 vs .604). FAISS hit@10/@64 against §07's full-bge run: within .003 everywhere comparable, and HNSW diverged from exact flat search by at most .0093 (NFCorpus; 0 and .001 elsewhere). SciFact full-bge was re-measured in this run (.74039, matching §02's .7404); the other three baselines are §02's recorded same-harness numbers. The quantized encoder is the same retriever.
| decoys | arrangement | qbge P@1 | BM25 P@1 | qbge R@8 | BM25 R@8 |
|---|---|---|---|---|---|
| 2 | scattered | 0.81 | 0.19 | 1.00 | 1.00 |
| 10 | scattered | 0.19 | 0.09 | 1.00 | 0.91 |
| 25 | scattered | 0.00 | 0.03 | 0.59 | 0.25 |
| 100 | scattered | 0.16 | 0.06 | 0.28 | 0.25 |
| 25 | clustered | 0.66 | 0.25 | 1.00 | 1.00 |
| 100 | clustered | 0.72 | 0.03 | 0.81 | 0.25 |
Same protocol as §05, 32 seeds per cell, k=8. It holds precision exactly where BM25 dies — 0.81 vs 0.19 at two scattered decoys, 0.72 vs 0.03 at 100 clustered — then collapses at 25+ scattered like everything else (P@1 0.00, R@8 0.59). And §06's VSA keeps its crown: at 25 scattered decoys the VSA's R@8 = 0.69 against the quantized encoder's 0.59 — the three lanes each still hold a bench where they are best.
the negative, published plainly
A fourth arm unioned BM25's candidates with the quantized encoder's before ranking. It added nothing in any cell — precision and recall identical to the encoder alone across all nine measured cells, and at 100 scattered decoys its mean rank was marginally worse (26.6 vs 25.1). On this bench you can pick a lane; you cannot average them.
| corpus tokens | leCore TTFA | qbge-CPU TTFA | ratio | leCore query | qbge query |
|---|---|---|---|---|---|
| 50,000 | 0.06 s | 20.3 s | 343× | 0.30 ms | 43.9 ms |
| 200,000 | 0.52 s | 75.7 s | 145× | 0.90 ms | 139.9 ms |
| 1,000,000 | 15.6 s | 381.3 s | 24.4× | 4.96 ms | 147.2 ms |
Time-to-first-answer = index (and for the encoder, embed) + one query, 8 CPU threads, no GPU anywhere. The gap narrows with corpus size and is still 24.4× at 1M tokens. Query-side the encoder is fine — 27–31 ms batch-1 median at 4–8 threads (192-thread oversubscription regresses: 301 ms median, 1.6 s p90 — pin your threads). Real-corpus probes embedded 3.3–7.3 docs/s, putting full CPU embeds of the four BEIR corpora at an estimated 18–87 minutes each. And the caveat that kills a free lunch: requantize buys size, not CPU speed — the artifact stores dequantized fp32 snapped to the quantized grid, so it embeds at exactly normal bge-base CPU speed, and the 132 MB packed figure is an estimate, not a shipped file; no packed loader is built.
the integration verdict
Rerank mode is cheap — top-k × ~31 ms — but provably cannot fix J=0: §08 measured BM25's top-64 containing the needle 2.3% of the time at zero overlap, and reordering a candidate set that lacks the answer 97.7% of the time reorders failure. Candidates must be generated semantically — union at recall time, or embedding at bind time — which costs the corpus embed: 75 s at 200k tokens, 381 s at 1M, measured above. So the stage ships exactly as SEMANTIC_STAGE=off|rerank|union in the sidecar — committed, default off, union capped at 60k items, no score fusion anywhere (the recorded fusion experiment degraded monotonically at every weight, §02). A sensible opt-in for corpora around the low hundreds of thousands of tokens, where the one-time embed is about a minute; at the multi-million-token product ceiling the lexical and VSA lanes remain the only affordable first stage.
Method and fine print: two runs on a shared 192-core EPYC 9454 host, load average 10–18 throughout, recorded per-row and per-task in the JSONs — every timing carries that flag; every score is load-immune. Before/after arms swapped the quantized weights into the HF cache blob and restored the original from a backup after each arm (the swap is logged; no separate checksum chain was recorded). The Unicron MP filter itself did nothing, as predicted — 71 of 73 tensors heavy-tail passthrough, 2 guarded, 0 filtered — requantization is the lever. Sources: results/bge_assim/ and results/bge_matrix/.
11the audit
Three people told us we were wrong about something. Read each against the actual claim: much cheaper, and usable where the alternative cannot run at all. A confirmed counterclaim is only a defeat if it contradicts something we asserted — several of these are true and were priced in from the start.
| claim | who | verdict | does it hit the pitch? |
|---|---|---|---|
| BM25 is not SOTA — it's ~free and runs on a Raspberry Pi | Shaw | CONFIRMED — bge-base wins 4/4 | NO — quality parity was never claimed. This is the trade. |
| Compare against HNSW, not flat FAISS | Shaw | CONFIRMED — HNSW is sub-linear as stated | NO — HNSW fixes query time, not the O(corpus) embed we skip. |
| Embeddings = quality; ANN = performance | Shaw | CONFIRMED — exactly the split found | NO — we compete on the third axis, cost. |
| leCore recall trails FAISS | Cotten | CONFIRMED — our dense arm scored 0.416 | NO — and we deleted that arm. |
| Test adversarially, with near-matches | Cotten | CONFIRMED at k=8 — P@1 1.00 → 0.19 at two decoys | YES, briefly — fixed by raising k. R@24 = 1.00 at 25 decoys; holds on real BEIR corpora at crossover k ≤ 64, and at 2.68M docs at k = 128 — exactly the cap (§07). |
| You are ignoring leCore | Moose | CONFIRMED twice — chunker and encoder hand-rolled | Process, not product — both now use leCore. |
| 518 ms per query | Cotten | EXPLAINED — §06: prebuilt index 2–16 ms, deployed wrapper was 20.5 s. His number was charitable | Deployment, not algorithm — the wrapper rebuilt the forest per query. Fixed and re-measured (§06): forest cached, deployed 5k-item context 3.5 s per query → 117 ms p50; local 20k recall 21.82 s → 5.9 ms. |
| Never skip the VSA — profile it over 1024d; the _Index wrapper isn't optimized | Cotten | CONFIRMED on the wrapper — ~4,200× the math floor (§06) | NO — and the profile he asked for handed the VSA its first outright win: decoy resistance (§06). |
| HRR attention is 41.5% worse than softmax at 303M — our own published number | us | OVERSTATED BY US — 32.7% on the fairness rerun (§04) | NO — the operator still loses, so the verdict stands. But ~a fifth of the gap was our five confounds, and HRR's optimal LR is 3.3× softmax's. We found this by re-running our own negative. |
| Embeddings win everywhere | implied | REFUTED on cost — 13.32 s vs 182.90 s TTFA | This is the pitch. |
the two claims that survive intact are the only two we made
Much cheaper
13.7× faster to first answer at 1M tokens, 98× fewer prompt tokens, with enough headroom to raise top_k 8× and still be 12.3× cheaper than sending the corpus.
Much longer
24/24 answered at 1M–5M where the control returns UNRUNNABLE at every tier.
The adversarial bench is the one that genuinely landed, and only at the shipped top_k=8: a config bug, not an architecture defeat. The real-corpora shotgun (§07) then confirmed the recovery lever on BEIR data at k ≤ 64, and its 2.68M-doc rerun answered the corpus-size question: the crossover survives by a single grid point — k = 128, exactly the production cap — while the same run's cost fields put the embedder's amortization at tens-to-hundreds of queries. The semantic NIAH (§08) then drew the lever's hard boundary: it requires at least one shared content word, and at measured zero overlap raising k recovers nothing — while one word in four restores full recall at the deployed k, ahead of bge-base. That boundary was named in the changelog as the likeliest failure before it was run, and it is published either way. And the code bench (§09) took the whole picture to the product's real workload: on repositories, identifier and cross-file questions are the lexical lane's case at zero embedding cost, the paraphrase gradient reproduces §08's cliff on real code, and bge-base is not dominated — it wins paraphrase outright, which is exactly where the staged semantic lane points. §10 then built and measured that lane with our own machinery: within .002 nDCG of full bge on all four MTEB tasks, identical at the J=0 cliff, 343× → 24.4× behind leCore on CPU cost-to-first-answer from 50k to 1M tokens — the trade priced end to end, the union negative published, and the stage shipped env-gated, default off. The newest row is the one nobody asked for: we re-ran our own 303M negative with our five confounds removed and found we had overstated it — 41.5% is 32.7% (§04). The operator still loses, so the verdict is unchanged; the magnitude was partly our methodology, and correcting a number that was already against us is the same discipline as publishing it in the first place.
12open questions & closed records
Crossover k at 100k+ docs — done
Closed 2026-08-14 by the 2.68M-doc shotgun (§07). Crossover k does grow with corpus size — 3.6–8.7k docs → k = 32–64; 2.68M docs → k = 128, exactly the production cap, zero headroom. Past a few million docs the lexical lane cannot reach embedding-grade hit rates at sane k; the thing that could break the cost/recall tradeoff now has a measured boundary instead. The row stays here as the record that it was once open.
Does a bigger k eat the cost win? — answered at 2.68M docs
Closed 2026-08-14, from the same run's own cost fields (§07): leCore pays ~13,037 extra tokens on every query at the NQ crossover; bge pays its 2,927.6 s GPU encode once. Break-even Q* = C_embed ÷ (Δtok × p_tok) ≈ 25–250 queries depending on LLM input price — at multi-million-doc scale the embedder amortizes fast, and the cost win narrows to cold-start, low-volume, or no-GPU workloads. The prices are assumptions; the formula and inputs are measured. The row stays as the record it was once open.
Loss-matched LM NIAH rerun — ran 2026-08-14, did not match
Closed 2026-08-14, and it closed against the premise. The rerun landed (§04): compile symmetric, batch split symmetric, HRR's learning rate swept and tuned to 2e-3, both arms on the same bin, seed and 7,629 steps. It moved HRR 4.5138 → 4.1805 and cut the penalty 41.5% → 32.7% — but it did not produce a loss-matched pair, so the LM NIAH confound is narrowed, not removed. Loss-matching these two at 303M is not reachable by configuration; it would need a compute-matched-not-token-matched design, which is not built. The follow-up table in §04 has also not been re-scored on the new checkpoints — unmeasured, not unchanged. The row stays as the record it was once open.
Code bench follow-ups
§09's mention-crowding explanation for the fastapi identifier plateau is a hypothesis, not a measurement — verifying it means scoring definition-vs-mention ranks directly. All three repos are Python; non-Python corpora are unmeasured. The semantic stage itself is no longer the open item — §10 measured it component-by-component (2026-08-14) — but running it over §09's paraphrase set, the exact cells bge wins, is still the obvious next bench.
Forest cache in prod, then re-measure — done
Closed 2026-08-14, same day it opened. The cache shipped, deployed, and was re-measured live: a 5,000-item forest context went from 3.5 s per query to 117 ms p50; local vec recall at 20k items from 21.82 s to 5.9 ms. Full before/after now lives in §06. The row stays here as the record that it was once open.
Method notes
n=9 per cell at 60k–500k, n=6 past it, single runs at temperature 0. Cost benchmarks run on an unloaded box — an earlier set was discarded after load average 67 inflated FAISS index times by 3×. Only the harness is parallelised: BM25 fits single-threaded because that is how leCore ships it, and the embedder gets its own batching because that is how it ships. Model: gemini-2.5-flash via OpenRouter. elizaOS's own harness could not be run — bun install fails on three private @elizaos/* packages — so the numbers are ours, our runner, their methodology.