Reproducible retrieval benchmarks · including the results that went against us

What we measured, and what it refuted

We are not trying to be better. The claim is much cheaper and usable at context lengths where the alternative cannot run at all. Below is everything three people asked us to measure against that claim — including the results that went against us.

Every number is reproducible from github.com/staccDOTsol/supercontext. Model: gemini-2.5-flash via OpenRouter. Single runs at temperature 0.

pitch

13.7×

faster to first answer at 1M tokens

pitch

98×

fewer prompt tokens, same answers

pitch

24/24

answered at 1M–5M where the model refused outright

we lose

0–4

lost to bge-base on nDCG — the trade, priced in

bug found

k=8

the default we shipped dropped the answer under decoys — a config bug, since fixed

we lose

J=0

a question sharing zero vocabulary with its answer defeats the ranker — and more context does not help

The first three are the pitch. The fourth is the trade we made on purpose and would make again — quality parity was never claimed. The fifth is the one real defect these benchmarks found — a settings value, not an architecture, since raised in production. The sixth is the measured limit of the whole recall-recovery story: leCore's ranker is lexical, and at exactly zero question–answer overlap no top_k recovers it (§08).

01done·this is the pitch

Cost to first answer

A FAISS index does not exist until an embedding model has read every document. Latency-only comparisons skip that step. This measures model load + index build + one query — what a user actually waits for on a cold corpus.

Time to first answer, 1M-token corpus (11,940 docs) — lower is better

leCore BM25 13.32 s
FAISS flat 182.90 s
FAISS HNSW 183.19 s

HNSW makes queries sub-linear, exactly as Shaw said — but it does not touch the dominant term. Graph construction is ~0.3 s on top of ~182 s of embedding.

corpus docs leCore index flat index HNSW index leCore query HNSW query
50k 597 0.05 s 10.60 s 10.60 s 2.39 ms 31.23 ms
200k 2,388 0.57 s 42.38 s 42.41 s 0.86 ms 31.03 ms
1M 11,940 13.31 s 182.03 s 182.28 s 4.86 ms 35.83 ms

02done·we lose

Retrieval quality, MTEB

Shaw asked us to stop quoting our own test and run the standard one. We did, then ran three real embedding models through the same harness so it isn't ours-versus-literature. bge-base beats leCore on all four tasks. "BM25 is a floor" was correct.

nDCG@10 by task — higher is better

leCore BM25 bge-base

SciFact

leCore.6705
bge-base.7404

ArguAna

leCore.4314
bge-base.6375

NFCorpus

leCore.3167
bge-base.3735

SCIDOCS

leCore.1577
bge-base.2172
0nDCG@100.80
system SciFact ArguAna NFCorpus SCIDOCS FiQA
leCore BM25 + expand 0.6705 0.4314 0.3167 0.1577 0.2378
all-MiniLM-L6-v2 0.6451 0.5017 0.3159 0.2164 —
e5-large-v2 0.7221 0.4642 0.3715 0.2050 —
bge-base-en-v1.5 0.7404 0.6375 0.3735 0.2172 —

removed from ship

Our own dense VSA encoder scored 0.4160 on SciFact and made the hybrid worse at every fusion weight (0.05 → .6596, 0.15 → .6307, 0.30 → .6057, against .6705 for BM25 alone). We were shipping the worst of three configurations. It has been removed.

03done·this is the pitch

Past the model's window

Scored three ways, never two: HIT, MISS, and UNRUNNABLE — the endpoint refused the request. Collapsing a refusal into a wrong answer would let a memory system claim it "beat" a model that was never allowed to compete.

900K 6 / 6 leCore HIT 5 / 6 control HIT
1M 6 / 6 leCore HIT unrunnable control refused · 0 / 6
2M 6 / 6 leCore HIT unrunnable control refused · 0 / 6
5M 6 / 6 leCore HIT unrunnable control refused · 0 / 6

Cells answered correctly, 900k–5M tokens, n=6 per tier. Beyond 1M the control returns "The input token count exceeds the maximum number of tokens allowed." Tokens the model reads stay flat: a 60k corpus → ~1,458 read; a 5M corpus → 3,401 read.

Totals across all 24 cells

leCore HIT 24 · UNRUNNABLE 0 · 80,300 prompt tokens
control HIT 5 · UNRUNNABLE 19 · 4,836,442 prompt tokens

Bars are prompt tokens spent, linear scale — 98× fewer for 24 hits against 5.

04done·negative·we overstated it — self-audited 2026-08-14

303M matched pair — HRR vs softmax

Two identical 303M language models — same seed, same schedule, same step count, RoPE in both arms; only the attention operator differs. We published this as 41.5% worse for HRR. Then we re-ran it against our own methodology, and 41.5% became 32.7%. About a fifth of our headline negative was our configuration, not the operator. The verdict is unchanged — the operator still loses — but the number we shipped was inflated, so here is the corrected one.

Fairness rerun · validation loss after 7,629 steps / 999,948,288 tokens, each arm at its own tuned learning rate — lower is better

softmax · 6e-4 3.1494
gated-HRR · 2e-3 4.1805
0val loss (nats)5.0

bits/byte proxy: softmax 4.5436 · gated-HRR 6.0312. Both arms COMPILE=0, micro 16 / accum 8, same 7B-token bin, same seed, identical step count.

32.7%

worse for HRR, with every confound removed — down from the 41.5% we published. The prior measurement at 12M params was a ~14% gap; scale still widens it, just by less than we said.

3.0×

slower per token, not 11× as previously published — 151,269 vs 50,797 tok/s, after wiring the repo's own Triton scan.

10.2×

from the Triton scan alone, which also cut memory 74.9 → 41.8 GB.

Auditing our own headline negative — five confounds, all ours

correction

The published pair had five asymmetries between the arms, and every one of them was introduced by us: torch.compile on for softmax and off for HRR; a learning rate tuned for neither and inherited from softmax; and a mismatched micro-batch / gradient-accumulation split that, because the loader strides by micro×accum, also desynchronised the data stream — the two arms were not reading the same tokens in the same order. The rerun sets COMPILE=0 on both, micro 16 / accum 8 on both, the same 7B-token bin, the same seed, and lands both arms on exactly 7,629 steps and 999,948,288 tokens.

41.5% → 32.7%

the HRR penalty as published, and as re-measured. The absolute gap goes 1.3236 → 1.0311 nats — 22% of it was artifact.

0.333

val loss HRR recovered from fair configs plus a tuned learning rate — 4.5138 → 4.1805. Softmax moved 0.041 (3.1902 → 3.1494).

still behind

HRR loses by a third under conditions we can no longer blame. The correction is to the magnitude, not the direction.

New finding: HRR's optimal learning rate is 3.3× softmax's

measured

Six learning rates, one per arm, four A100s, identical budget — 1,148 steps / ~150.5M tokens each, same schedule and warmup. HRR's best is 2e-3; softmax's is 6e-4. At the same budget, running HRR at softmax's 6e-4 costs it 1.14 nats of validation loss. That is what the original run did — so part of what we published as an architecture penalty was a hyperparameter one.

HRR learning rate val loss @ 150.5M tok bpb proxy note
3e-4 7.4070 10.6861
4.5e-4 7.5033 10.8250 worst of six
6e-4 7.4224 10.7083 softmax's LR ← the published run
9e-4 6.4272 9.2725 non-monotone: 1.2e-3 is worse
1.2e-3 6.8328 9.8577
2e-3 6.2822 9.0633 best — and the largest tested

Read this honestly: the minimum sits at the edge of the swept range, so 2e-3 is a floor on HRR's optimum, not a located one — the true optimum may be higher and we have not measured it. The curve is also noisy at the small end (4.5e-4 scores worse than 3e-4, 1.2e-3 worse than 9e-4), which is one sweep per point at ~150M tokens, not a seed-averaged one. All six are also mid-decay: the schedule was written for 200M tokens and every point was cut by the same 1.5-hour wall clock at ~150.5M, so these are ranking numbers, not converged ones — which is why 6e-4 reads 7.4224 here and 4.1805 is nowhere near it (that arm ran 6.6× longer). The rerun used 2e-3 because it was the best rate measured, not because it was proven optimal.

caveats we own

The rerun closed the compile, learning-rate and batch-split asymmetries. Two caveats survive it and one is new. Both arms are still undertrained — ~1B tokens is roughly 1/6 Chinchilla-optimal for 303M, so this is the gap at 1B tokens, not the asymptotic gap. The rerun used a 7B-token bin where the original used 5B, so its validation tail is not the original's: compare the 3.1494 and 4.1805 to each other, and treat every cross-run absolute as indicative only. The internal gap is the experiment; the softmax arm moving 0.041 between runs is the size of that bin effect on this axis. A 6B-token Chinchilla-scale pair is running now — softmax past 2.59B of 6B tokens at the time of writing, HRR configured at lr 2e-3. It is in flight, not a result; nothing on this page depends on it and nothing should be read into it until both arms finish.

Follow-up: LM NIAH on this pair — a minor negative result

negative

Plant "The secret code is XXXX" at start/middle/end of filler and score the logit margin of the true code token against a random wrong one — >0 means the needle won. 3 seeds per cell, both arms RoPE, training block 1024. Softmax retrieves inside its training length, decays at 2048, and is dead by 4096–8192. HRR never localizes at any length. Its lone positive cells (+0.16 at 8192) are position-insensitive — near-identical at start, middle and end (0.159 / 0.159 / 0.151) — a uniform bias toward the answer token, not retrieval. At 8192 the honest read is both arms fail.

context softmax margin · start / mid / end HRR margin · start / mid / end
512 +4.84  +4.82  +5.79 −0.05  −0.05  +0.06
1024 +1.33  +4.85  +4.71 −0.14  −0.14  −0.10
2048 +0.13  +1.72  +0.47 −0.28  −0.28  −0.28
4096 +0.13  +0.07  +0.08 −0.59  −0.59  −0.60
8192 −0.31  −0.30  −0.37 +0.16  +0.16  +0.15 ← bias, not retrieval

Mandatory caveat: these cells were scored on the published pair — hrr val 4.5138 vs softmax 3.1902, 41.5% worse — so they confound the operator with plain LM quality: a worse language model fails NIAH for reasons that have nothing to do with attention. The fairness rerun above did not fix that. It was the matched-loss attempt, and it did not produce matched loss: 4.1805 vs 3.1494 is still 32.7% apart. So the confound stands, narrower. We have not re-scored this table on the rerun checkpoints — that is unmeasured, not unchanged — and the caveats box above applies in full.

Sources: the arms' own final JSONs — results/runpod_303m_rescue/final_softmax.json (3.149412655830383) and the HRR arm's out/final_hrr.json on its Vast box, mirrored to results/hrr_vs_softmax/303m_hrr_vast/ (4.180486848950386). Sweep logs: sweeps303/sweep_hrr_*.log, six files, one per rate. The superseded pair is 303m-results/final_softmax.json and is kept, not deleted. Checkpoints pushed to staccs/lecore-303m-rerun.

05done·found a real config bug

Adversarial near-match

Cotten asked for this twice: plant decoys that carry the full query vocabulary but are not the answer. He was right — at the shipped top_k=8 it breaks, and that turned out to be a settings value rather than an architecture problem.

The control row is the whole finding

On the zero-decoy corpus a dumb exact-substring counter — no idf, no length normalisation, no saturation — scores precision@1 = 1.00, identical to leCore. A benchmark that a substring counter aces is not measuring a ranker. It is measuring the corpus's lexical confusability, and ours was ~zero.

Does the answer survive decoys? · 200 docs, 32 seeds/cell, k=8 — the default we shipped · scattered decoys, the hard case

recall@8 — answer reached the model precision@1 — answer ranked first

0 decoys

recall@81.00
P@11.00

2 decoys

recall@81.00
P@1.19

5 decoys

recall@81.00
P@1.25

10 decoys

recall@8.91
P@1.09

25 decoys — recall breaks here

recall@8.22
P@1.03

50 decoys

recall@8.12
P@1.03

100 decoys

recall@8.25
P@1.06
0rate1.0

Precision@1 dies at two decoys — 1.00 → 0.19. The surviving envelope is recall, not precision: the answer still reaches the model's context up to ~10 competing passages (recall@8 = 0.91) and breaks between 10 and 25 (0.22). Past that, at this k, the answer is not in the context at all — so no amount of model quality recovers it, and only a bigger top_k does. Mean margin goes negative (−0.44 at 25 decoys): the best decoy outscores the needle.

…and the fix is to buy more context, which is the entire pitch

We were never trying to out-rank bge-base. So the right answer to "the needle fell to rank 12" is not a better ranker — it is a bigger top_k. Recall is fully recoverable at every decoy level measured, and the k required is roughly the number of competing passages.

top_k needed to bring recall back to 1.0 — swept over 8, 12, 16, 24, 32, 48, 64

25 competing

k = 24

recall 0.22 at k=8 → 1.00

50 competing

k = 48

recall 0.12 at k=8 → 1.00

100 competing

k = 64

recall 0.25 at k=8 → 1.00

and we can afford it — prompt tokens on the 60k–500k sweep, linear scale

leCore k=8 68,057
leCore k=64 ~544,000
control — send the corpus 6,676,718

Raising k from 8 to 64 is 8× the retrieved tokens — 68,057 → ~544k on the 60k–500k sweep — which is still 12.3× fewer than the control's 6,676,718. Precision@1 was never the product. The product is that the answer is in the context at a price the alternative cannot match.

the honest limit

This is a scaling relationship, not a free lunch. Token cost grows with the number of confusable passages — a property of the query rather than of corpus size, but whether confusability grows with corpus size is not measured here, and it is the thing that would break this. That bench has since run on real corpora — §07: recall recovers within k ≤ 64 on all three BEIR sets, and at 2.68M docs the crossover survives at k = 128 — exactly the production cap, with zero headroom. And §08 now marks the lever's hard boundary: k-recovery requires at least one shared content word — at measured zero question–answer overlap, recall stays flat no matter the k.

what changed in prod

The shipped default was top_k=8 with a hard cap of 32 — both too low; the cap alone foreclosed the 100-decoy case. Fixed and deployed: default 16, cap 128, verified live.

What does survive

leCore's BM25 is genuinely ranking rather than guessing. tied@top is ≈1.0 for leCore against 5–85 for the substring floor, so the floor's occasional wins are argsort index luck while leCore's are real separation. Clustering also helps, because a clustered decoy lands inside the answer's own chunk instead of competing with it as a separate document.

100 decoys, half-clustered — leCore vs the substring floor · 4.7× separation

mean rank of the answer — lower is better · scale 0–32

leCore6.66
floor31.22

recall@20 — higher is better · scale 0–1.0

leCore1.00
floor0.28

the honest reading

Our headline needle number was inflated by vocabulary isolation. A bag-of-words model has no representation for contains the answer versus is about the answer. So precision@1 is not something BM25 can be made to deliver — and it is not something we need. The adversarial bench did not kill the thesis; it killed the precision-retrieval thesis and revealed a viable cost/recall tradeoff in its place. The claim is recall@k, bought with a bigger top_k on the one axis we are cheap on — never precision@1.

Two arms do hold precision under exactly these decoys: the pure VSA path, profiled next — §06 is its first outright win, published beside its losses — and the quantized semantic encoder, measured under the same protocol in §10.

06done·its first outright win·wrapper was 4,200× the floor — fix deployed & re-measured

The VSA, profiled

This bench exists because a critic asked. Cotten, publicly: "rewrite the benchmark so it never skips to BM25; always profile the VSA over 1024d and pay attention to _Index — btw the _Index wrapper isn't optimized." So the pure VSA path is now a first-class, never-skipped row, and every number ships regardless of direction. He was right about the wrapper. And the same profile handed the VSA — the arm §02 removed from the hybrid — its first outright win, on the exact adversarial bench that broke BM25 in §05.

Where the time goes — 100k docs × 1024d, cosine top-k, Apple M5 Pro CPU · log scale (1 ms → 30 s), lower is better

Index.nearest · prebuilt, warm, k=1 2.3 ms
raw f32 matmul + argmax — the math floor 4.8 ms
Index.nearest · prebuilt, warm, k=64 15.7 ms
Cotten's independent measurement 518 ms
deployed sidecar wrapper · top-8 · before fix 20,536 ms was — now 117 ms p50 deployed (cached), measured on a 5,000-item forest context (~100 ms of it network; one-time 3.5 s first-query build). Local vec recall: 5k 2.09 s → 2.4 ms · 20k 21.82 s → 5.9 ms. The bar stands as the historical record.

The math is milliseconds. The shipped wrapper rebuilt the HoloForest on every query — the one-time forest build is 19.6–20.4 s, essentially the whole 20.5 s, ~4,200× the floor. Cotten's 518 ms sat 40× below that deployed number: his complaint was charitable. The fix was a cache — build the forest once, on the first vec-path query, invalidated by any write — and it shipped and deployed the same day (commit ee813a7), then re-measured: local end-to-end vec recall 2.09 s → 2.4 ms at 5k items, 21.82 s → 5.9 ms at 20k; live, a 5,000-item forest context answers at 117 ms p50 (~100 ms of that is network) after a one-time 3.5 s first-query build — every query used to pay that 3.5 s. The BM25 path did not regress: 116–123 ms warm vs 102/97 ms the same morning, within jitter. Cached and uncached return identical ids and scores by construction, verified exact-equal at 5k and 20k. Encoder cost for scale: ~0.8 ms per 600-char chunk (0.72–0.87 across thread arms, ~1,150–1,380 chunks/s), query encode 0.08 ms.

The win: decoy resistance

Same protocol as §05 — 800 lines, 32 seeds per cell, k=8, scattered decoys the hard case. The idf-weighted dense superposition holds rank under exactly the term-collision that inverts BM25's scoring.

P@1 · 2 decoys, scattered

VSA.81
BM25.19

VSA margin +0.059 — the needle outscores the best decoy

R@8 · 25 decoys, scattered

VSA.69
BM25.22

mean rank 7.4 vs 12.0 — 3× more often in a k=8 context

P@1 · 25 decoys, clustered

VSA1.00
BM25.25

clustered VSA P@1 = 1.00 at 2, 10 and 25 decoys alike

scattered decoys system P@1 R@8 mean rank mean margin
0leCore VSA1.001.001.00n/a
leCore BM251.001.001.00n/a
2leCore VSA0.811.001.22+0.059
leCore BM250.191.002.09−0.014
10leCore VSA0.091.003.22−0.273
leCore BM250.090.915.69−0.348
25leCore VSA0.030.697.41−0.460
leCore BM250.030.2211.97−0.444

The honest read of the win: at 10+ scattered decoys the VSA's precision@1 collapses too — 0.09 → 0.03. What it keeps is recall: at 25 scattered decoys the answer is still inside a k=8 context 69% of the time against BM25's 22%, before any of §05's k-raising. Clustered decoys cost it nothing at all: P@1 = 1.00 at 2, 10 and 25, where BM25 falls 1.00 → 0.72 → 0.25.

the losses, republished unfiltered

This is not "the VSA is good now." Its ranking generality is unchanged: SciFact nDCG@10 0.4160 against leCore BM25's 0.6705 and bge-base's 0.7404 — the same score that got the arm removed from the hybrid in §02. And at zero lexical overlap it is still not retrieval: J=0 recall@64 = 0.0964, reproduced here to the fourth decimal against §08's recorded run, position skew and all.

What this leaves: three lanes, each measured where it wins

BM25 — anchored recall

Any shared content word and it beats bge-base at deployed k (§08, J=.257: .958 at k=8) — on the cost axis nothing else touches (§01, §03).

VSA — decoy resistance

Holds the needle under the term-collision that inverts BM25 — P@1 .81 vs .19 at two scattered decoys, 1.00 clustered throughout (this section).

quantized bge — zero overlap

The only system not flat at J=0 (.615 at k=8 → .846 at k=64, §08). Since measured — §10: quantized with leCore's own machinery to within .002 nDCG of full bge, priced on CPU, shipped env-gated (default off). The VSA still beats it at 25 scattered decoys (R@8 .69 vs .59).

Method: local Apple M5 Pro (18 cores), Python 3.14.6 / numpy 2.5.0, $0 spent, wall clock 137 s. Load average was 5–6 throughout and is recorded per-section in results.json — a timing-contamination flag we publish rather than hide. Two thread arms, pinned single-thread and free, both reported; the ladder above uses the free arm (the single-thread floor was 3.7 ms, the single-thread wrapper 21.3 s — same shape). At 1k docs the auto index is exact and the wrapper is 2–4 ms; at 12k docs the forest kicks in and the wrapper is already 1.7–2.3 s. The rebuild dominates everywhere the forest exists.

07done·retires the §5 caveat·knife-edge at 2.68M docs

Shotgun on real corpora

§5's decoys were templated strings we wrote ourselves, which left one open question: does recall recovers with k survive corpora where the confusable passages are real? So: three BEIR corpora with real judged queries, leCore BM25 against GPU-embedded baselines in the same harness, hit@k swept from 1 to 200. The question is the crossover: at what k does leCore's hit rate reach what bge-base delivers at k=10? Answer: k = 32–64 on all three, at 3.2–6.4× the retrieved tokens. And below, this section's own open question re-run at scale: the same crossover at 2.68M documents lands at k = 128 — exactly the production cap.

hit@k vs k — SciFact, the hardest crossover of the three · higher is better

leCore BM25 bge-base
1.0 .75 .5
bge @10 = .883 — the bar crossover → .893 at k=64
1258101624324864100200

k — passages retrieved per query · SciFact, 5,183 docs, 300 judged queries

leCore @10 .820 bge-base @10 .883 leCore @64 .893 ← crossover bge-base @64 .950 leCore @200 .923 bge-base @200 .980

leCore reaches .893 at k=64, clearing bge-base's .883 at k=10 — on the corpus where that took the most context. Real corpora land at gentler k than §5's worst-case templated decoys, where the required k was roughly the full count of competing passages.

corpus docs / queries bge hit@10 leCore hit@10 crossover k leCore hit@k token cost
SciFact 5,183 / 300 0.883 0.820 k = 64 0.893 6.4× · ~17,857 tok/q
NFCorpus 3,633 / 323 0.746 0.681 k = 32 0.755 3.2× · ~9,725 tok/q
ArguAna 8,674 / 1,401 0.884 0.685 k = 48 0.894 4.8× · ~10,431 tok/q
NQ (at scale, below) 2,681,468 / 3,452 0.781 0.475 k = 128 — the cap 0.784 12.8× · ~14,142 tok/q

raising k helps them too

Do not read this as parity. Give bge-base the same k=64 and it scores .950 / .854 / .991 on the three corpora — more context helps everyone. What this bench establishes is that the recall-recovery lever survives real corpora, and what it costs: 3.2–6.4× tokens to match bge's k=10. The sellable number is cost per hit without a GPU or an embedding pass, not "we catch up."

inside the shipped cap

Production now defaults to top_k=16 with a cap of 128 — every crossover k measured here (32–64) sits inside it. The old cap of 32 would have foreclosed SciFact's k=64 outright. The 2.68M-doc run below lands exactly on the cap — inside it by zero.

Method and the honest fine print: clean Vast RTX 4090, load average 1.5, embedding on GPU. bge-HNSW matched exact flat search on every corpus — identical on SciFact and ArguAna, within 0.007 hit on NFCorpus — so there is no ANN recall loss to exploit at these sizes. And on this box the embedders' index build was as fast or faster than our BM25 build (bge 11.3–15.5 s GPU vs leCore 15.8–50.5 s CPU): the §01 cost-to-first-answer win assumes no GPU is present. One boundary on everything above: the ranker is lexical — §08 measures the case where a question shares zero vocabulary with its answer, and the k lever does not survive it.

At 2.68M documents, the crossover survives by a single k: it lands exactly on the cap

The three corpora above top out at 8,674 documents — §12's stated "not built" was whether crossover k grows with corpus size. It does. On Natural Questions — 2,681,468 docs, 3,452 judged queries, same harness, same crossover definition — bge-base's hit@10 is .781, leCore's is .475, and leCore does not reach .781 until k = 128 — exactly the production cap (hit@128 = .784, at 12.8× the tokens, ~14,142 tok/query). The trend across everything measured: 3.6–8.7k docs → k = 32–64; 2.68M docs → k = 128. Zero headroom left.

hit@k vs k — Natural Questions, 2.68M docs · higher is better

leCore BM25 bge-base
1.0 .5 0
bge @10 = .781 — the bar crossover → k=128 = the cap
1258101624324864100128

k — passages retrieved per query · Natural Questions, 2,681,468 docs, 3,452 judged queries

leCore @10 .475 bge-base @10 .781 leCore @128 .784 ← crossover bge-base @128 .960

leCore's curve meets bge-base's k=10 bar at the last point on the grid — .784 at k=128, against the bar of .781. One grid point later and there would have been no crossover to report inside the shipped cap. Past a few million documents the lexical lane cannot reach embedding-grade hit rates at sane k — that boundary now has a number, and it is where the semantic lane (§06's third lane, the quantized-bge stage — since measured in §10 and shipped env-gated, default off) takes over.

system hit@10 hit@128 one-time index of which embedding query
leCore BM25 (O(tokens) rebuild) 0.475 0.784 309.8 s CPU 0 s 24.16 ms
bge-base · FAISS flat 0.781 0.960 2,932.9 s 2,927.6 s GPU 6.37 ms
bge-base · HNSW 0.774 0.950 3,668.7 s shared with flat 0.16 ms
MiniLM · FAISS flat 0.683 0.931 976.8 s 973.1 s GPU 4.82 ms

the index still costs nothing to embed

2.68M documents indexed in 309.8 s on CPU with zero embedding against bge-base's 2,932.9 s of GPU work — the §01 cold-start win holds at three orders of magnitude more docs. But the recurring side turned: matching bge's hit@10 now takes every one of the 128 passages the production cap allows.

the amortization math, from this run's own costs

leCore's premium at crossover is (128 − 10) × 110.5 ≈ 13,037 extra tokens per query, every query; bge pays its 2,927.6 s encode once. Break-even: Q* = C_embed ÷ (Δtok × p_tok) ≈ 250 queries at $0.40/hr GPU and $0.10/M input tokens, ~25 at $1.00/M. The prices are assumptions, the inputs are measured — and the honest reading is that at this scale the embedder amortizes within tens-to-hundreds of queries. The cost win at multi-million-doc scale is confined to cold-start, low-volume, or no-GPU workloads.

The fine print, all of it. The shipped BM25 build cannot run at this scale: its postings loop is O(vocab × N) — ~10¹² iterations at 2.68M docs — and does not terminate. The bench used an O(total tokens) rebuild whose postings are bit-identical (same idf, same per-(term,doc) weight expression, same ordering; scores() output identical). The 309.8 s is that rebuild, and the shipped build's non-termination is itself a reported finding — the fourth config-class engineering defect these benches have found, after top_k=8, the cap of 32, and §06's per-query forest rebuild. The box was not clean: RTX 4090, load average 6–24 across the arms, recorded per-arm in the JSON — timings carry that flag; hit rates are load-immune. e5-large was skipped at this scale (3× bge's encode cost for the same question), and only NQ has run — HotpotQA at 5.2M docs is the next tier and is not measured.

08done·negative at J=0 — the honest limit

Semantic NIAH — zero overlap

The loss, plainly: leCore's ranker is lexical. A question sharing zero vocabulary with its answer defeats it, and buying more context does not help. 64 needle/question pairs per overlap level, with the zero enforced in code, not by eyeballing — stemmed content-token sets must not intersect, no two content tokens may share a 4-character prefix, no shared numbers. Needles were embedded in ~222k tokens of real fineweb prose (~1.9k chunks), 3 positions × 2 seeds, through the deployed 600-char chunker. At measured Jaccard 0.000, leCore BM25 goes R@8 = .005 → R@64 = .023 — flat. Unlike §05 and §06, raising k recovers nothing, because there is no lexical anchor to rank on.

recall@k at measured zero question–answer overlap · higher is better · each group is one system swept over k

leCore BM25 · J = 0.000 — flat, k recovers nothing

k=8.005
k=16.005
k=32.008
k=64.023

leCore VSA · J = 0.000 — argsort luck, position-skewed

k=8.008
k=16.023
k=32.052
k=64.096

bge-base · J = 0.000 — the reference struggles too

k=8.615
k=16.706
k=32.784
k=64.846

leCore BM25 · J = 0.257 (~one word in four) — ahead of bge at every k

k=8.958
k=16.987
k=32.995
k=641.000
0recall@k · 384 trials at J=0 · fineweb ~222k tokens per seed1.0

The measured boundary: the cliff is at exactly zero. One shared content word in four (measured J = .257) restores full recall at the deployed k — .96 at k=8, 1.00 at k=64 — ahead of bge-base at every k. And even the semantic reference struggles at true zero: bge-base reaches only .85 at k=64, not 1.0.

measured overlap system R@8 R@16 R@32 R@64
J = 0.000leCore BM250.0050.0050.0080.023 ← flat
leCore VSA0.0080.0230.0520.096 ← argsort luck
bge-base0.6150.7060.7840.846 ← the reference struggles too
J = 0.257 (~25%)leCore BM250.9580.9870.9951.000 ← ahead of bge at every k
bge-base0.8810.9130.9580.968
J = 0.476 (~50%)leCore BM251.0001.0001.0001.000
bge-base0.9420.9580.9660.976

the k lever has a boundary

This bounds §05 and §07's recall-recovers-with-k story: the lever works when any lexical anchor exists, and does not exist at zero. Also honest: leCore's VSA arm's .10 at k=64 is position-skewed — .16 / .10 / .03 at start / middle / end — which is argsort luck, not retrieval. The changelog named this exact case as the likeliest failure before the bench was run; it is published here either way.

one word in four is enough

The failure needs exactly zero. At measured J = .257 leCore BM25 is .958 at k=8 → 1.000 at k=64, ahead of bge-base at every k, and at J = .476 everything saturates near 1.00. Real questions that share literally no content word with their answer are a corner of the space — but it is a corner we cannot serve, and now it is measured.

Method and the caveat to carry: pairs generated by Qwen2.5-7B-Instruct, then gate-enforced with rejects counted and published — at the zero level 65 rejected for stem overlap, 15 for shared prefixes, 8 for shared digits; 0 needles destroyed by chunking; 384 trials at J=0 and 378 at each other level. The no-shared-digit rule sometimes forces zero-overlap questions to alter quantities (a "50 new sensors" needle asked about as "30 extra devices"). That is fine for retrieval scoring — only chunk identity is scored — but the J=0 pairs are not quotable as QA pairs.

This gradient is not a synthetic artifact. §09 reproduces it on real code: paraphrased questions about a 47k-token repo retrieve perfectly, and the same question type collapses to .30–.40 on a 4.3M-token repo — while identifier-anchored questions stay the lexical ranker's case throughout. And the one system not flat at J=0 is now leCore's own lane: §10 quantizes bge-base with our machinery — recall@64 at J=0 identical to full bge — and prices exactly where it can afford to run.

09done·mixed — §08's gradient, on real code

Bind your repo — code corpora

The product's actual emerging user is an agent throwing a codebase at zoo_ask — bind the repo over MCP, then ask questions about it. This bench measures that path, factually, for the first time. Three Python repos at pinned tags — requests v2.32.3 (~47k tokens, 18 files), fastapi 0.115.12 with its tests (~625k, 360 files), django 5.1.7 (~4.3M, 2,107 files) — and 82 questions generated by Qwen2.5-7B against sampled ground-truth functions, gated by an answerable-from-chunk judge, in three types: identifier-anchored (must name the function), behavioral paraphrase (programmatically gated to contain no identifier from the code — the semantic stress), and cross-file ("what calls X and what happens to its return value"). Gold answers are (path, line-range) regions, so the production 600-char chunker and function-level chunking score against the same truth.

The claim shape this backs: identifier and cross-file questions are the lexical ranker's case, at zero embedding cost — cross-file asks go .90 at k=8 → 1.00 by k=64, identifier .73 → .85, and leCore indexes the 4.3M-token django tree in ~3.0 s on CPU where bge-base needs ~40–42 s of GPU embedding per chunking. Paraphrase questions need a semantic stage — and bge-base is not dominated here, stated plainly below.

hit-any@k by question type · pooled over the three repos · 600-char chunking · higher is better

leCore · cross-file — the lexical ranker's best case

k=8.90
k=16.93
k=32.97
k=641.00

leCore · identifier-anchored

k=8.73
k=16.81
k=32.81
k=64.85

bge-base · paraphrase — bge wins this outright

k=8.42
k=16.46
k=32.54
k=64.69

leCore · paraphrase — the semantic stress, and our loss

k=8.31
k=16.35
k=32.42
k=64.50
0hit-any@k · 82 questions, pooled · tokens@k (mean, 600-char): ~1.2k at k=8 → ~9.5k at k=641.0

Identifiers are lexical anchors and the ranker treats them as such: the two anchored question types climb into the .85–1.00 band inside the deployed k range, for zero embedding work. The paraphrase line is the honest one: pooled it stalls at .50 at k=64 while bge-base reaches .69 — and the pooled number hides a gradient, drawn next.

paraphrase questions, hit-any@64 vs repo size · 600-char chunking · §08's gradient on real code

leCore BM25 bge-base

requests · ~47k tokens

leCore1.00
bge-base1.00

fastapi · ~625k tokens

leCore.40
bge-base.50

django · ~4.3M tokens

leCore.30
bge-base.70
0hit-any@64 · n=6–10 questions per point (±0.15) — trust the shape, not any single cell1.0

Same shape as §08's overlap cliff, produced by real code instead of a synthetic gate: when a question shares no identifier with its answer, recall decays with corpus size, and at 4.3M tokens k=64 recovers only .30 (leCore, from .10 at k=8). bge-base decays too — .70, not 1.00 — but it is not dominated: it wins paraphrase outright, and under function-level chunking it also edges identifier asks (.88 vs .81 pooled at k=64). The sell here is cost-per-hit and zero-embed indexing, never better recall.

system chunking question type h@8 h@16 h@32 h@64
leCore BM25600-charidentifier0.730.810.810.85
600-charparaphrase0.310.350.420.50 ← the semantic stress
600-charcross-file0.900.930.971.00
functionidentifier0.690.730.810.81
functionparaphrase0.310.380.460.54
functioncross-file0.830.900.931.00
bge-base600-charidentifier0.620.690.690.81
600-charparaphrase0.420.460.540.69 ← bge wins this outright
600-charcross-file0.971.001.001.00
functionidentifier0.810.850.850.88 ← edges leCore's 0.81
functionparaphrase0.540.580.580.58
functioncross-file1.001.001.001.00

Cross-file scored strictly — definition and caller both in the top-k — at k=64: leCore .77 (600-char) / .83 (function); bge .97 (600-char) but only .67 (function). Chunking is a real variable, not a footnote: function-level chunking helps bge on identifier asks (django h@8 .80 vs .40) and mildly helps leCore's paraphrase, while 600-char windows serve strict cross-file better under bge.

the anchored case costs nothing to index

leCore builds the django index — 4.3M tokens, ~30k chunks — in ~3.0 s on CPU; bge-base pays ~40–42 s of GPU embedding per chunking before the first query. On the questions agents actually ask by name — identifier and cross-file — that free index reaches .85–1.00 inside the deployed k range. This is the zoo_ask bind-your-repo path, measured — the MCP's real user, not a proxy corpus.

at scale, an identifier anchors every mention

The fastapi curiosity: identifier asks plateau for both systems — leCore holds at .70 from k=16 through 64 under both chunkings; bge sits at .60–.70 through k=32 and reaches only .80 at 64. Our hypothesis, not separately verified: fastapi ships its tests, and test files citing the name crowd the definition out of the top-k — the ranker finds every mention, not the one that matters.

Method and every caveat: one Vast RTX 4090 (~55 min, $0.53 total across three boxes, two of which never ran), destroyed after. n = 6–10 questions per cell — single cells carry ~±0.15 noise, so quote the gradients and pooled rows, never one cell. All three repos are Python. Questions are one model's style (Qwen2.5-7B; the answerable gate passed ~60–80% of candidates). bge truncates chunks at 512 tokens — its real behavior — and function chunks were capped at 6,000 chars. The BM25 postings build was re-ordered doc-major to O(total tokens) — the same rebuild class §07's 2.68M-doc run needed — with score bit-identity vs the shipped build asserted on 323 chunks × 10 probes. Source: results/code_bench/ — questions.json with gold regions, results.json, summary.json, corpus_stats.json.

10done·one negative inside — published

The semantic lane, measured

§08 ended at the cliff: at zero question–answer overlap only bge-base retrieves anything, and §06 named a quantized-bge semantic stage as the third lane — staged, not shipped, unmeasured. This is the measurement. We ran leCore's own requantize machinery (holographic_refactor.quantize_group, a budgeted per-tensor bit ladder with embedding cosine as the fidelity metric) over bge-base-en-v1.5, then made the result take every exam family it will face: MTEB before/after on four tasks — GPU and CPU — semantic NIAH before/after, FAISS flat + HNSW, §05's adversarial protocol with the quantized encoder as a fourth arm, and cost-to-first-answer against leCore BM25 on CPU only.

the artifact — budgeted requant, 1% encoder-fidelity budget · ladder 3/4/5/6/8 bits · group 64 · 126 s

7.62

mean bits over the quantized set

132 MB

packed estimate, vs 438 MB fp32

.9894

holdout embedding cosine (calib .9900)

11 / 73

tensors refused every rung — stayed fp32

The encoder resisted. Layers 0–1 took 3-bit wholesale (cosine ≥ .995); the budget then walked the ladder up through the middle of the stack, and 11 tensors — clustered in layers 7–9 — refused even 8-bit. The same machinery pointed at a 9B decoder (a separate, still-unfinished run) was landing whole tensors at ~3 bits/weight: encoder embedding geometry tolerates far less weight noise than decoder next-token loss. Uniform arms for calibration: 4-bit everywhere is 93.4 MB at SciFact .72842 — 98.4% of full, the sweet spot only if a packed runtime existed; 3-bit everywhere is 82.7 MB at .6922, visibly damaged. Shipped conclusion: the budgeted artifact. Home: huggingface.co/staccs/lecore-bge-assimilated.

nDCG@10, full vs quantized — the bars are the point: you cannot see the difference

full bge-base quantized (7.6-bit budget)

SciFact Δ −.0016

full.7404
quantized.7388

ArguAna Δ +.0000

full.6375
quantized.6375

NFCorpus Δ −.0019

full.3735
quantized.3716

SCIDOCS Δ −.0020 — the worst delta of the four

full.2172
quantized.2152
0nDCG@100.80

Worst delta across the four tasks: .0020 (SCIDOCS); ArguAna is a rounding hair higher. The CPU arm re-ran SciFact and NFCorpus and scored bit-identical to the GPU rows. Semantic NIAH at J=0: recall@64 = .8464 for full and quantized — identical (recall@8 .615 vs .604). FAISS hit@10/@64 against §07's full-bge run: within .003 everywhere comparable, and HNSW diverged from exact flat search by at most .0093 (NFCorpus; 0 and .001 elsewhere). SciFact full-bge was re-measured in this run (.74039, matching §02's .7404); the other three baselines are §02's recorded same-harness numbers. The quantized encoder is the same retriever.

Where it slots into the decoy bench

decoys arrangement qbge P@1 BM25 P@1 qbge R@8 BM25 R@8
2scattered0.810.191.001.00
10scattered0.190.091.000.91
25scattered0.000.030.590.25
100scattered0.160.060.280.25
25clustered0.660.251.001.00
100clustered0.720.030.810.25

Same protocol as §05, 32 seeds per cell, k=8. It holds precision exactly where BM25 dies — 0.81 vs 0.19 at two scattered decoys, 0.72 vs 0.03 at 100 clustered — then collapses at 25+ scattered like everything else (P@1 0.00, R@8 0.59). And §06's VSA keeps its crown: at 25 scattered decoys the VSA's R@8 = 0.69 against the quantized encoder's 0.59 — the three lanes each still hold a bench where they are best.

the negative, published plainly

A fourth arm unioned BM25's candidates with the quantized encoder's before ranking. It added nothing in any cell — precision and recall identical to the encoder alone across all nine measured cells, and at 100 scattered decoys its mean rank was marginally worse (26.6 vs 25.1). On this bench you can pick a lane; you cannot average them.

What the lane costs, CPU only

corpus tokens leCore TTFA qbge-CPU TTFA ratio leCore query qbge query
50,0000.06 s20.3 s343×0.30 ms43.9 ms
200,0000.52 s75.7 s145×0.90 ms139.9 ms
1,000,00015.6 s381.3 s24.4×4.96 ms147.2 ms

Time-to-first-answer = index (and for the encoder, embed) + one query, 8 CPU threads, no GPU anywhere. The gap narrows with corpus size and is still 24.4× at 1M tokens. Query-side the encoder is fine — 27–31 ms batch-1 median at 4–8 threads (192-thread oversubscription regresses: 301 ms median, 1.6 s p90 — pin your threads). Real-corpus probes embedded 3.3–7.3 docs/s, putting full CPU embeds of the four BEIR corpora at an estimated 18–87 minutes each. And the caveat that kills a free lunch: requantize buys size, not CPU speed — the artifact stores dequantized fp32 snapped to the quantized grid, so it embeds at exactly normal bge-base CPU speed, and the 132 MB packed figure is an estimate, not a shipped file; no packed loader is built.

the integration verdict

Rerank mode is cheap — top-k × ~31 ms — but provably cannot fix J=0: §08 measured BM25's top-64 containing the needle 2.3% of the time at zero overlap, and reordering a candidate set that lacks the answer 97.7% of the time reorders failure. Candidates must be generated semantically — union at recall time, or embedding at bind time — which costs the corpus embed: 75 s at 200k tokens, 381 s at 1M, measured above. So the stage ships exactly as SEMANTIC_STAGE=off|rerank|union in the sidecar — committed, default off, union capped at 60k items, no score fusion anywhere (the recorded fusion experiment degraded monotonically at every weight, §02). A sensible opt-in for corpora around the low hundreds of thousands of tokens, where the one-time embed is about a minute; at the multi-million-token product ceiling the lexical and VSA lanes remain the only affordable first stage.

Method and fine print: two runs on a shared 192-core EPYC 9454 host, load average 10–18 throughout, recorded per-row and per-task in the JSONs — every timing carries that flag; every score is load-immune. Before/after arms swapped the quantized weights into the HF cache blob and restored the original from a backup after each arm (the swap is logged; no separate checksum chain was recorded). The Unicron MP filter itself did nothing, as predicted — 71 of 73 tensors heavy-tail passthrough, 2 guarded, 0 filtered — requantization is the lever. Sources: results/bge_assim/ and results/bge_matrix/.

11the audit

Scorecard on the counterclaims

Three people told us we were wrong about something. Read each against the actual claim: much cheaper, and usable where the alternative cannot run at all. A confirmed counterclaim is only a defeat if it contradicts something we asserted — several of these are true and were priced in from the start.

claim who verdict does it hit the pitch?
BM25 is not SOTA — it's ~free and runs on a Raspberry Pi Shaw CONFIRMED — bge-base wins 4/4 NO — quality parity was never claimed. This is the trade.
Compare against HNSW, not flat FAISS Shaw CONFIRMED — HNSW is sub-linear as stated NO — HNSW fixes query time, not the O(corpus) embed we skip.
Embeddings = quality; ANN = performance Shaw CONFIRMED — exactly the split found NO — we compete on the third axis, cost.
leCore recall trails FAISS Cotten CONFIRMED — our dense arm scored 0.416 NO — and we deleted that arm.
Test adversarially, with near-matches Cotten CONFIRMED at k=8 — P@1 1.00 → 0.19 at two decoys YES, briefly — fixed by raising k. R@24 = 1.00 at 25 decoys; holds on real BEIR corpora at crossover k ≤ 64, and at 2.68M docs at k = 128 — exactly the cap (§07).
You are ignoring leCore Moose CONFIRMED twice — chunker and encoder hand-rolled Process, not product — both now use leCore.
518 ms per query Cotten EXPLAINED — §06: prebuilt index 2–16 ms, deployed wrapper was 20.5 s. His number was charitable Deployment, not algorithm — the wrapper rebuilt the forest per query. Fixed and re-measured (§06): forest cached, deployed 5k-item context 3.5 s per query → 117 ms p50; local 20k recall 21.82 s → 5.9 ms.
Never skip the VSA — profile it over 1024d; the _Index wrapper isn't optimized Cotten CONFIRMED on the wrapper — ~4,200× the math floor (§06) NO — and the profile he asked for handed the VSA its first outright win: decoy resistance (§06).
HRR attention is 41.5% worse than softmax at 303M — our own published number us OVERSTATED BY US — 32.7% on the fairness rerun (§04) NO — the operator still loses, so the verdict stands. But ~a fifth of the gap was our five confounds, and HRR's optimal LR is 3.3× softmax's. We found this by re-running our own negative.
Embeddings win everywhere implied REFUTED on cost — 13.32 s vs 182.90 s TTFA This is the pitch.

the two claims that survive intact are the only two we made

Much cheaper

13.7× faster to first answer at 1M tokens, 98× fewer prompt tokens, with enough headroom to raise top_k 8× and still be 12.3× cheaper than sending the corpus.

Much longer

24/24 answered at 1M–5M where the control returns UNRUNNABLE at every tier.

The adversarial bench is the one that genuinely landed, and only at the shipped top_k=8: a config bug, not an architecture defeat. The real-corpora shotgun (§07) then confirmed the recovery lever on BEIR data at k ≤ 64, and its 2.68M-doc rerun answered the corpus-size question: the crossover survives by a single grid point — k = 128, exactly the production cap — while the same run's cost fields put the embedder's amortization at tens-to-hundreds of queries. The semantic NIAH (§08) then drew the lever's hard boundary: it requires at least one shared content word, and at measured zero overlap raising k recovers nothing — while one word in four restores full recall at the deployed k, ahead of bge-base. That boundary was named in the changelog as the likeliest failure before it was run, and it is published either way. And the code bench (§09) took the whole picture to the product's real workload: on repositories, identifier and cross-file questions are the lexical lane's case at zero embedding cost, the paraphrase gradient reproduces §08's cliff on real code, and bge-base is not dominated — it wins paraphrase outright, which is exactly where the staged semantic lane points. §10 then built and measured that lane with our own machinery: within .002 nDCG of full bge on all four MTEB tasks, identical at the J=0 cliff, 343× → 24.4× behind leCore on CPU cost-to-first-answer from 50k to 1M tokens — the trade priced end to end, the union negative published, and the stage shipped env-gated, default off. The newest row is the one nobody asked for: we re-ran our own 303M negative with our five confounds removed and found we had overstated it — 41.5% is 32.7% (§04). The operator still loses, so the verdict is unchanged; the magnitude was partly our methodology, and correcting a number that was already against us is the same discipline as publishing it in the first place.

12open questions & closed records

Not built yet

Crossover k at 100k+ docs — done

Closed 2026-08-14 by the 2.68M-doc shotgun (§07). Crossover k does grow with corpus size — 3.6–8.7k docs → k = 32–64; 2.68M docs → k = 128, exactly the production cap, zero headroom. Past a few million docs the lexical lane cannot reach embedding-grade hit rates at sane k; the thing that could break the cost/recall tradeoff now has a measured boundary instead. The row stays here as the record that it was once open.

Does a bigger k eat the cost win? — answered at 2.68M docs

Closed 2026-08-14, from the same run's own cost fields (§07): leCore pays ~13,037 extra tokens on every query at the NQ crossover; bge pays its 2,927.6 s GPU encode once. Break-even Q* = C_embed ÷ (Δtok × p_tok) ≈ 25–250 queries depending on LLM input price — at multi-million-doc scale the embedder amortizes fast, and the cost win narrows to cold-start, low-volume, or no-GPU workloads. The prices are assumptions; the formula and inputs are measured. The row stays as the record it was once open.

Loss-matched LM NIAH rerun — ran 2026-08-14, did not match

Closed 2026-08-14, and it closed against the premise. The rerun landed (§04): compile symmetric, batch split symmetric, HRR's learning rate swept and tuned to 2e-3, both arms on the same bin, seed and 7,629 steps. It moved HRR 4.5138 → 4.1805 and cut the penalty 41.5% → 32.7% — but it did not produce a loss-matched pair, so the LM NIAH confound is narrowed, not removed. Loss-matching these two at 303M is not reachable by configuration; it would need a compute-matched-not-token-matched design, which is not built. The follow-up table in §04 has also not been re-scored on the new checkpoints — unmeasured, not unchanged. The row stays as the record it was once open.

Code bench follow-ups

§09's mention-crowding explanation for the fastapi identifier plateau is a hypothesis, not a measurement — verifying it means scoring definition-vs-mention ranks directly. All three repos are Python; non-Python corpora are unmeasured. The semantic stage itself is no longer the open item — §10 measured it component-by-component (2026-08-14) — but running it over §09's paraphrase set, the exact cells bge wins, is still the obvious next bench.

Forest cache in prod, then re-measure — done

Closed 2026-08-14, same day it opened. The cache shipped, deployed, and was re-measured live: a 5,000-item forest context went from 3.5 s per query to 117 ms p50; local vec recall at 20k items from 21.82 s to 5.9 ms. Full before/after now lives in §06. The row stays here as the record that it was once open.

Method notes

n=9 per cell at 60k–500k, n=6 past it, single runs at temperature 0. Cost benchmarks run on an unloaded box — an earlier set was discarded after load average 67 inflated FAISS index times by 3×. Only the harness is parallelised: BM25 fits single-threaded because that is how leCore ships it, and the embedder gets its own batching because that is how it ships. Model: gemini-2.5-flash via OpenRouter. elizaOS's own harness could not be run — bun install fails on three private @elizaos/* packages — so the numbers are ours, our runner, their methodology.