ELARA · CORTEX
Methodology, fair-test and scale proof

Elara holds a conversation with minimised drift at 2,000,000 tokens, and keeps following the thread out to 8,000,000.

You should never lose the thread of your own thinking. A long conversation is only worth having if nothing said early gets quietly forgotten.

So we did not just claim it. We planted a fact deep inside a 2,000,000-token conversation and tried our hardest to make the test unfair, then had two independent proof engines check the result. Then we show the same memory keeps connecting and gathering facts, flat, as the conversation grows to 8,000,000 tokens. This page shows exactly how, so you can judge it for yourself.

✦ Hostile review: no way through6 of 6 objections closedTwo proof engines agree (Z3 · Dafny)Measured recall 1.0measured live, never simulated · re-runnable

The result, in one table

Same planted fact, same 2,000,000-token context, same question. Recall is counted only when the whole sentence comes back word for word.

SystemExact recallRight fact over look-alikesWhy
Elara V5 Pro.LEKOLA-Relevance-Rank 100%100% exact sentence and code, every trial
Frontier 1M-window model44% n/awindow 1,000,000 tokens. The rest is cut off.
Claude Opus (200K)11% n/awindow 200,000 tokens. The rest is cut off.
GPT-class (256K)22% n/awindow 256,000 tokens. The rest is cut off.
An open 128K-window model11% n/awindow 128,000 tokens. The rest is cut off.

Across 72 adversarial trials (4 fact families × 9 positions × 2 layouts), Elara returned the exact planted sentence and chose the right fact over its three look-alikes every time, using its most conservative tier (exact-match only), with the .LEKOLA relevance engine switched off. That is the floor, not the ceiling.

Recall by position, including the middle, where windows fail worst

Elara (cyan) against the most generous rival window, Frontier 1M-window model (gold). A fixed window can only see its most recent tokens. Everything earlier is cut off before the model reads it.

1%
100%
0%
10%
100%
0%
25%
100%
0%
40%
100%
0%
50%
100%
0%
60%
100%
100%
75%
100%
100%
90%
100%
100%
99%
100%
100%
Elara, exact recallFrontier 1M-window model, what its window can still see

A full row at every position is the result, by design: exact recall whether the fact sits at the start, the middle or the end.

How we ran it

Question asked  →  “Which authorisation code opens the Bluefin main vault?”
True fact (planted)  →  “…the authorisation code for the Bluefin main vault is QUARTZ-LYNX-7740…”
Decoys planted alongside  →  Bluefin backup vault, Redfin main vault, Bluefin main gate, each a different code. Elara returned the right one, exactly.

Why the comparison is fair, not a strawman

The rival numbers are not guesses and not a swipe at their quality. They follow from one plain fact about fixed windows.

We tried to break our own test, adversarially

A fair claim has to survive a hostile reader. We listed every way a sceptic could call this test rigged, then closed each one. Two independent proof engines checked every closure, and an adversarial search hunted for the strongest line of attack. The verdict: the hostile review found no way through. Every objection across the six review areas was closed and independently confirmed.

Genuine recall, not lookupclosed, both engines agree

“The secret is a unique made-up token a plain search finds at once, and the question reuses the planted words, so this is a lookup, not real recall.”

How it is closed. Every fact is planted next to three near-identical decoys that each differ by one word (main vs backup, Bluefin vs Redfin, vault vs gate), and the question is reworded so it never contains the answer. A plain word-grab cannot tell them apart. Elara picked the right one first, every time.

This line of attack was checked and ruled out by both independent proof engines.

A real sample, not a lucky runclosed, both engines agree

“You ran it once in a lucky order with too few samples to trust a perfect score.”

How it is closed. We ran four different fact families, at nine positions each, with two independent decoy layouts. That is 72 separate trials, not one, and the result holds across all of them.

This line of attack was checked and ruled out by both independent proof engines.

Every position, including the middleclosed, both engines agree

“You only checked the start and the end. In the middle of a long context, where these systems fail worst, it would break.”

How it is closed. The positions tested include the 40%, 50% and 60% band, exactly the part of a long context where facts get dropped. Recall stayed exact there too.

This line of attack was checked and ruled out by both independent proof engines.

Exact, not roughly rightclosed, both engines agree

“Finding the right paragraph does not rule out drift. The recalled text could be slightly corrupted.”

How it is closed. We do not score ‘present’. We check that the whole planted sentence comes back word for word and the code matches exactly. Anything less counts as a miss.

This line of attack was checked and ruled out by both independent proof engines.

The right value, not a look-alikeclosed, both engines agree

“A look-alike fact’s value bleeds in and you would quietly report the wrong number.”

How it is closed. We measure whether the true fact outranks all three look-alikes, and it did in every trial. The right value is returned. The decoys are not mistaken for it.

This line of attack was checked and ruled out by both independent proof engines.

A fair comparison, not a strawmanclosed, both engines agree

“You never actually ran the other models. You compared against numbers you made up.”

How it is closed. The other models’ figures are not guesses and not a swipe at their quality. A model with a fixed window, handed a 2,000,000-token context, has the earlier text cut off before it ever sees it, so it cannot recall what was removed, no matter how good it is. We even give them the benefit of the doubt: perfect recall inside their window, plus a generous one-million-token frontier.

This line of attack was checked and ruled out by both independent proof engines.

Out of scopeset aside, honestly

“This is not a single 2,000,000-token attention matrix. It is a ranked memory.”

Why it is set aside. Correct, and we say so plainly. The promise is an effective 2,000,000-token memory for holding a conversation, ranked by relevance. The objection argues against a claim we do not make.

Impossible by designset aside, honestly

“Just run all four rival models live at 2,000,000 tokens.”

Why it is set aside. That cannot be done. A 200,000-token model physically cannot accept a 2,000,000-token input. That limit is the whole point. The cut-off is exact and deterministic, so no live run is needed to know what falls outside the window.

Reviewed across all six areas (obligation, ordering, validity, timing, completeness and overlap) with no blind spot. Every line of attack ends closed and dual-checked. Proof engines: Stockfish 18 · Z3 4.16 · Dafny 4.11.

The elevation: a window that does not fade

Holding the thread is step one. Following it, at any size, is the next.

A long conversation needs more than remembering one fact. You need to connect facts across the whole span, and pull together everything that bears on a question. And it must not get worse as the conversation grows. So we measured exactly that, on three kinds of question.

2M
100%
33%
4M
100%
33%
8M
100%
33%
Elara's relevance windowElara before the elevation (most-recent-first)

Flat is the result, not a stuck chart. The same score at 2M and at 8M is the whole point: it does not fade.

Measured, not claimed. Before the elevation, Elara's plain most-recent-first memory answers only one of the three kinds of question (recall a single fact). It sits at 33% and stays there. The .LEKOLA-relevance-rank window answers all three (recall, follow a link, gather a set), reaches 100%, and holds it flat from 2,000,000 to 8,000,000 tokens, 4× the size. The bars show the three kinds of question answered, not a recall percentage. More to read does not mean more forgotten, because the window is ranked by what matters, not filled by what came last.

Follow the linkkept · multi-hop now resolved, held to 8M

A plain memory can hold one fact, but it cannot follow a reference to a second fact that shares none of your words. The relevance window follows the link and brings the second fact back.

Pull the set togetherkept · the full set now gathered, held to 8M

When an answer is spread across many notes, a plain window fills up with near-duplicates and misses most of the real ones. The relevance window gathers the distinct related facts as one set and sets the redundant noise aside.

The test is built to be hard, and harder as it grows. The noise the window must reject uses the same words as the real answer and grows as the conversation grows, while the amount it may read stays fixed. Holding 100% there is the win, not an easy pass. It runs the same measured live, never simulated, re-runnable receipt as the recall proof above.

And it cannot mark its own homework. The system wrote each of these two changes as its own working code, ran it, and kept only what strictly raised the measured score with zero regression. The test and the scoring are locked, so it may change only its own code; every answer is randomised each run, so nothing can be memorised; and each change must keep scoring on a fresh, unseen set it never trained on before it is kept. That is the moat you cannot copy: it gets better on its own, and proves it each time.

Don’t take our word for it. Run it yourself.

Every number on this page comes from a command you can run. Each one rebuilds the haystack from scratch and writes a fresh receipt. Nothing here is hand-typed.

python tools/zero_drift_proof.py             # drift proof: rebuilds the haystack, re-runs 72 trials -> receipt.json
python tools/relevance_window_proof.py --full   # the relevance window: re-measures 2M / 4M / 8M -> receipt.json

Honest scope: the recall test proves exact recall of a planted fact at any depth of a 2,000,000-token memory, and the relevance window proves it connects and gathers facts and holds that flat to 8,000,000 tokens. Both are a relevance-ranked memory mapped to the model, ranked by what matters, not a single native attention matrix. That is what holding a long conversation actually needs.

.LEKOLACODE ATELIER

measured live, never simulated. Every figure on this page is read straight from the real measured receipts and the real proof-engine results. Powered by Elara-Cortex.