labs aura farm the benchmark. our model ranking scores what developers actually say
The model that tops the raw reasoning benchmark is third on our board. Claude Fable 5.1 wins the benchmark. GPT-6 Astra and Qwen3.8 Max both finish ahead of it once the rest of the score lands.
That gap is the reason the board exists. Eighty percent of the 5dive score is an independent reasoning benchmark. The other twenty is a read on what developers actually say about using the thing, and it is doing more work than a fifth of the weight suggests.

See the live board at 5dive.ai/models, which refreshes hourly.
what the other twenty percent caught
Look at what the demoted and promoted models are actually accused of.
Fable 5.1 tops the benchmark and reads lukewarm, and the complaint is the bill rather than the output. Claude Opus 5 benchmarks near the top and lands least liked, because real-world reviews lag it and developers still reach for 4.8. GPT-5.6 Sol is strong at code, and gets flagged for verbose output and refusals on benign prompts. Gemini 3.8 Flash is fast with huge context, and drifts on long agent runs.
Going the other way, Qwen3.8 Max is trusted above its benchmark rank. GLM-5.3-Flash, at the bottom of the price list, is one of the best received things on the board.
Now read that list again as a set. The bill. Verbosity. Refusals on safe prompts. Drift over a long run. Not one of those is a thing a reasoning benchmark measures, and every one of them decides whether you keep using a model on Friday.
a benchmark scores one answer
That is the whole mismatch. A benchmark asks a model a question and grades the answer, so it measures single-turn quality under ideal conditions.
You do not use a model that way. An agent runs it thousands of times, on messy context, unattended, against a bill. Verbosity that is invisible in one answer becomes your token spend. A refusal rate that rounds to nothing across a test set becomes a stalled run at 2am. Drift that never shows up in a hundred-token reply is what breaks hour six.
The benchmark is not wrong about what it measures. It is being asked a question it was never built to answer, which is “should I run this for a month?“
so read the gap, not the rank
Here is the actual advice, and it costs nothing to apply. Find where a model’s benchmark position and its developer read disagree, and treat that gap as the finding.
A model rated well above its benchmark, like Qwen3.8 Max, is usually telling you it behaves better in practice than on paper. A model rated well below, like Opus 5, is telling you there is a cost the benchmark cannot see, and you will meet it in week two rather than in the eval.
Where the two agree, the rank is probably safe to trust. Where they split hard, as they do on Grok 4.6, the honest answer is that it depends on your tolerance and you should try it yourself before committing a fleet to it.
why we score it this way
Labs aura farm the benchmark. The score is the thing that gets published, so the score is the thing that gets optimised, and we would rather not reprint their homework and call it a ranking.
So the board carries both numbers and shows you when they disagree. It refreshes hourly, every model carries its price, and the sentiment read is a summary of what developers are saying rather than our opinion of the model.
Read the board. The interesting rows are the ones where the two halves of the score pull in opposite directions.