We don't claim, we measure. The mathematics is solved, not guessed. And we lead with where we win.
The biggest working memory of any model on the board, a reasoning layer that checks its own work, and frontier-class scores on the hardest public tests, every number from a live run you can check.
Where we're unbeaten
The biggest working memory on the board
Give the AI your whole repository and it stops guessing at the parts it can't see. Elara holds 2,000,000 relevance-ranked tokens at once, double every frontier leader, and eight to twelve times the rest. This is published spec, not a benchmark.
Unbeaten: 2,000,000 relevance-ranked tokens with minimised drift. Double the 1,000,000 of GPT-5.4, Gemini 3.1 Pro and Claude Opus 4.8, and 8 to 12 times the rest. You pay no more than Claude Opus 4.8, which gives you half of it.
Context windows are the providers' published figures (Anthropic, OpenAI, Google, xAI, DeepSeek; cross-checked on Artificial Analysis), as of June 2026. Elara V5 Pro is 1,000,000 tokens native, extended to about 2,000,000 effective. Prices are per million tokens; Elara V5 Pro's maximum-context tier is $5 in / $25 out, the same as Claude Opus 4.8.
The proof · hardest public science
GPQA Diamond
When your problem is genuinely hard, you need an answer that holds up, solved not guessed. This is the toughest public science test there is: 198 graduate-level questions across physics, chemistry and biology, where PhD experts in the field still only reach about 65%.
Where we rank: PhD-level science: 87.9, frontier-class. Clear of DeepSeek R1 (71.0) and OpenAI o3 (83.3), and a whisker off the top tier.
How we ran it. Every question went live to the production API and was graded against the gold key, no cherry-picking. Elara V5 Pro reached 87.9% on the clean answered set (n=141 of the 198-question Diamond split); 57 items hit an upstream provider error and are held out pending a full re-run. Comparator figures are the labs' published GPQA Diamond scores. Frontier-class alongside the top labs, honest that it isn't the single best. Measured, not asserted.
The proof · code that runs
HumanEval · code generation
Code that looks right but doesn't run wastes your day, so we score by running it: 164 Python problems, each answer executed against the hidden tests. It either runs or it does not.
Where we rank: real code, above Grok 4 (90.2), DeepSeek R1 (90.0) and Gemini 3.1 Pro (89.6). Only the two Claudes and GPT-5.4 edge ahead, by a fraction.
Measured: Elara V5 Pro passed 90.9% on the first try (149/164), every answer run against the canonical hidden tests, not eyeballed. This benchmark is near-saturated at the frontier; we publish it anyway, for transparency. Comparators are the labs' published pass@1 scores. Measured, not asserted.