A colleague asked me, without irony, which model was winning. Grok 4.6 had just launched. GPT-5.6 Sol had been the coding default for a month. Anthropic was selling two flagships at once: Claude Fable 5, the expensive one, and Claude Opus 5, the one they want you to use every day. The tables all looked like a race. Someone was supposed to be ahead.

I used to treat "ahead on the index" as a useful default for a regular user. I no longer do, at least not at a gap this small. The surprise was not that the numbers disagreed. It was that the honest ones still were not a reason to switch tools this week.

I read the public boards dated 12 to 15 August 2026, including Artificial Analysis, Anthropic's Opus 5 post, xAI's Grok 4.6 launch table as checked by Digital Applied, and the VentureBeat write-up of the Grok launch. Then I gave Grok 4.6, GPT-5.6 Sol, and Claude Opus 5 the same two jobs: explain to a non-specialist whether a two-point index gap is a reason to switch, and compute a month of token cost from a fixed volume. Claude Fable 5 was not in my harness, so every Fable claim below is someone else's measurement, labelled as such.

Three things this piece does not do. It does not rank models for all time. It does not run a coding agent, a terminal benchmark, or a private LMSYS. And it is not a review of Grok's brand history, even though that history is a real input for some buyers and I will not pretend otherwise.

The short version

  • The job: pick a default model for everyday knowledge work this month, without becoming a benchmark analyst.
  • The trap: a launch table can be factually accurate and still leave out the model that leads the independent board.
  • What I compared: Grok 4.6, GPT-5.6 Sol, Claude Opus 5, and Claude Fable 5, on boards from 12 to 15 August 2026.
  • What I ran: the same explainer and the same arithmetic on the first three. All three said do not switch. All three got $312.
  • Where it breaks: coding agents, long terminal jobs, and regulated brand-safety reviews are different jobs. The everyday-writing result does not transfer to them.
  • Decision: split. Opus 5 for daily work, Grok 4.6 when the bill is the constraint, Sol for long terminal agents, Fable 5 only when you can name the remaining gap.

Start with the job

You already have a chatbot. You draft, summarise, analyse, and occasionally ask it to write code. Once a month a launch post tells you the old default is now losing. The implied action is to switch, today, because two points on a composite index look like a generation.

They are not. A composite index averages many tests into one number. Artificial Analysis's Intelligence Index is the one most of this month's coverage quotes. On the 15 August public snapshot it had Claude Opus 5 first at 63, Claude Fable 5 at 62, and Grok 4.6 around 61, with GPT-5.6 Sol in the same neighbourhood depending on the effort rung. A two-point gap on a 100-point composite is inside the sort of movement you get when the grader model changes or a lab posts a new effort tier.

Four labelled checks to apply before quoting a model score: effort rung, missing column, cost per task, and snapshot date, with examples from the August 2026 boards.
Fig. 1The four checks I used on every number in this piece. Skip any one of them and a two-point lead can reverse without the underlying models changing.

The job, then, is not "who is number one." It is "what should I type into on Monday, and what would make me change my mind." Does a two-point gap show up in a brief you would actually send? The rest of this piece is my attempt to answer that.

How to read a leaderboard

1. Read the rung before the score

Grok 4.6 has four documented reasoning rungs: low, medium, high, and xhigh, with high as the default and no max. GPT-5.6 Sol has six, ending at max. Claude Opus 5 and Fable 5 have five, also ending at max. Digital Applied's 12 August fact-check of xAI's launch table is the piece that made this load-bearing for me: xAI printed Grok 4.6 at high against Sol and Fable 5 at max. A High row against two Max rows is a default against two ceilings, not a race at the same effort.

I had assumed vendor tables inflate their own row. Here the opposite happened. On CursorBench v3.2, Grok 4.6 at Extra High is reported at 70.8%, above Fable 5 Max at 70.5%, and xAI printed its 69.9% High row instead. Accurate numbers, conservative rung, still a framed comparison.

2. Count the missing columns

The same xAI table has four columns: Grok 4.6 High, Grok 4.5 High, GPT-5.6 Sol Max, Claude Fable 5 Max. Claude Opus 5 is not in it. Digital Applied checked the eight independently verifiable rows and found Opus 5 leading seven of them, including Artificial Analysis's index at 63.05. Bolding the best score per row, inside a field that excludes the leader, is how an honest table still points at the wrong default.

3. Price the task, not the million tokens

Headline API rates as of the Grok 4.6 launch write-up: Grok 4.6 at $2 input and $6 output per million tokens under 200K, Claude Opus 5 at $5 and $25, GPT-5.6 Sol standard at $5 and $30, Claude Fable 5 at $10 and $50. Artificial Analysis's cost per Intelligence Index task is the more useful unit, because models spend different numbers of tokens to finish the same work. Digital Applied reports Grok 4.6 high at $0.84, Sol max at $1.23, Opus 5 max at $2.34, and Fable 5 with fallback at $3.14. Grok tying Sol on the index at a third of Fable's task cost is a real fact. It is also not the same fact as "Grok is more capable."

What the boards actually show

Reported

On Artificial Analysis's Intelligence Index, the 15 August public snapshot ranks Claude Opus 5 first at 63, Claude Fable 5 at 62, and Grok 4.6 at about 61. Digital Applied's 12 August table, using AA's own labels, has Opus 5 max at 63.05, Fable 5 with fallback at 62.07, GPT-5.6 Sol max at 60.93, and Grok 4.6 high at 60.92. The Sol and Grok high/max pair sit inside AA's stated plus-or-minus 1% interval.

Those four numbers are the headline. They are also a July-to-August moving object. Artificial Analysis's own 24 July Opus 5 note had Opus 5 max at 61, tied with Fable 5 at 60, ahead of Sol max at 59. The index version changed. The grader models changed. If your switching rule is "whoever is first today," you will switch more often than your work changes.

Reported

On knowledge-work boards, Opus 5 is the model the independent write-ups keep handing the lead. Artificial Analysis reported Opus 5 max at 1861 Elo on GDPval-AA v2 and 1720 on AA-Briefcase at launch. VentureBeat, citing AA on the Grok launch, has Grok 4.6 at 1753 Elo on GDPval-AA v2, Fable 5 Max at 1741, and Sol Max at 1728. Those two snapshots are not the same run. They are the same story: knowledge work is not where Sol pulls away.

Reported

On software-agent boards, the lead moves. VentureBeat's Grok 4.6 table, transcribed from xAI and checked against public boards, has GPT-5.6 Sol Max at 73% on DeepSWE v1.1 and 34.6% on Terminal-Bench v3.0, Fable 5 Max at 70% and 34.1%, Grok 4.6 High at 65.9% and 26%. CursorBench v3.2 is the close one: Fable 5 Max 70.5%, Grok 4.6 High 69.9%. Anthropic's Opus 5 post says Opus 5 at max is within 0.5 points of Fable 5's CursorBench peak at half the cost per task.

I expected Fable 5 to own the independent index because it is the expensive flagship. It does not. Opus 5 does, on the August snapshot, at half Fable's token price. That was the first surprise, and it is the one Anthropic has been trying to say out loud since 24 July. My read is that "expensive flagship" is no longer a useful proxy for "best on the independent board," at least not this month.

Why an honest table can still mislead

xAI's Grok 4.6 launch table is the cleanest example I have seen all month, because Digital Applied reproduced 31 of 31 publicly checkable cells and still ended up warning you not to route budget off the bolding.

The table does not invent scores. It selects them. Each cell is the best number available from a public board or a system card, produced under whatever harness that board used. Fable 5 still wins five of the ten rows xAI printed. Grok 4.6 wins three. Two of those three, on AA's own confidence intervals, sit inside the error bars. The model that leads the independent index is not in the picture.

I am not accusing xAI of cheating. I am saying this is how almost every launch table works, and this one is better than most because the footnote admits the cells were not run under one harness. "Third-party model scores are the best of self-reported or publicly available results." Once you can see that sentence, you can read the rest of the month the same way.

Inferred

A regular user who switches defaults because Grok 4.6 "beat" Fable 5 on xAI's table has not compared those two models at matched effort, and has not compared either of them to Opus 5 at all.

What I asked

I wanted a test a regular reader could recognise as work, not as a benchmark.

The question

If I give Grok 4.6, GPT-5.6 Sol, and Claude Opus 5 the same explainer and the same arithmetic, does the two-point index gap show up in a way I would notice without a rubric?

Baseline
All three give the same advice and the same numbers.
Continue if
One model is clearly better on the explainer, or one of them flubs the arithmetic.
Stop if
They agree on the advice and match on the math. The board did not show up in the work.

Task A: 160 to 200 words for a colleague who uses ChatGPT at work and does not follow AI news. If Opus 5 scores 63 and Grok 4.6 scores 61, should they switch this week? One concrete analogy. No em dash.

Task B: 40 million input tokens and 8 million output tokens. Sol at $5 / $30. Grok at $2 / $6. Cost on Sol, cost on Grok, saving. Three labelled lines, then stop.

I ran Grok 4.6 at high-fast, GPT-5.6 Sol at xhigh, and Claude Opus 5 at thinking-high, on 15 August 2026. Those are not matched rungs. I am doing the thing I just criticised, in the only harness I had. Fable 5 was not available, so it is absent on purpose. The prompts, the three outputs, and the arithmetic live in the prototype repository.

What I expected: Opus 5 would be the clearer explainer, Grok would be looser, and someone would drop a line of arithmetic. I also expected the index gap to be faintly visible in the prose, the way a 63 and a 61 on a hiring screen are faintly visible in an interview. That was the view I walked in with.

What happened

Table of the 15 August 2026 trial. Grok 4.6, GPT-5.6 Sol, and Claude Opus 5 all advised against switching on a two-point gap and all computed Sol 440 dollars, Grok 128 dollars, saving 312 dollars. Claude Fable 5 was not run.
Fig. 2My run, not a public board. A win would have been a difference I noticed without a scoring rubric. I did not get one.

Reproduced

All three models advised against switching on a two-point composite gap. Grok and Opus reached for a hiring analogy. Sol reached for a company car. All three computed Sol at $440, Grok at $128, and a saving of $312. None flubbed the arithmetic.

The stop threshold fired. I would not have known, from these two jobs, which model led the index.

That was the second surprise, and it is the one that changed the verdict. I had assumed "best on the board" would show up immediately in ordinary prose. On this pair of jobs it did not. The third surprise was smaller and more irritating: I had also assumed I would catch at least one model rounding the money. I did not. The $312 saving is the only number from the trial I trust as a fact, because I can redo it on paper.

Unknown

I am not sure whether a coding agent, a long terminal session, or a messy internal document would have separated them. I did not run those. The public boards suggest they might. My two jobs do not.

Where a regular user should not trust the number

The index is a blend, and your week is not

GDPval and AA-Briefcase look like knowledge work. Terminal-Bench and DeepSWE look like agents in a repo. CursorBench looks like in-editor coding. Averaging them produces a 63 and a 61 that can reverse on the one job you actually pay for. If your week is briefs and analysis, Opus 5's knowledge-work lead is the relevant column. If your week is a long autonomous terminal job, Sol's DeepSWE and Terminal-Bench rows are. Using the blend to pick either is how you buy the wrong default.

Effort rungs move the score more than the logo does

Opus 5's own GDPval spread across effort levels is more than 400 Elo, with output tokens moving about eightfold from low to max, per Artificial Analysis. Sol has the same shape. Grok's default is already "high." Comparing Grok high to Sol max is how you get a photogenic tie that disappears when both sit on high, or when both sit on max. I ran unmatched rungs myself. I am not going to pretend that is a fully fair bake-off.

Cost per million tokens is not cost per job

Grok 4.6's $2 / $6 headline applies below 200K input. Above that it doubles. Sol Fast mode is $10 / $60. Opus 5 Fast is twice the base. Fable 5 is $10 / $50 before you count retries. Artificial Analysis also reported Grok 4.6 finishing AA-Briefcase in about 53 turns against roughly 103 for Opus 5 max, which would flip a "Grok is cheaper per token" story if your harness is turn-heavy. I did not measure turns. I measured a hypothetical token volume, which is the version a regular user can check.

Brand-safety is a different scoreboard

VentureBeat's Grok 4.6 piece is careful on this, and I will be too. Earlier Grok deployments produced a string of public failures that procurement teams in banks, governments, and consumer brands will not evaluate as "the old model." I have no evidence that Grok 4.6 repeats those incidents. I also have no evidence a risk committee will treat that distinction as decisive. If your constraint is vendor governance rather than Elo, the index will not save you. I'd say skip Grok as a customer-facing default until your own review says otherwise, and do not ask a leaderboard to do that review.

A working default

Four-row decision table. Everyday writing starts at Claude Opus 5. A constrained bill starts at Grok 4.6. Long terminal work starts at GPT-5.6 Sol. A named remaining gap at twice the price starts at Claude Fable 5. Revert after a week if the difference is not felt.
Fig. 3A starting point for this month, with the kill switch in the last line. The index is allowed to disagree.

Everyday writing

4/54 out of 5

High confidence

All three models I ran gave the same advice and usable prose. A regular user would not have picked the index leader from the page.

Hard reasoning

3/53 out of 5

Low confidence

Public boards give Opus 5 a loud ARC-AGI-3 lead in Anthropic's telling. I did not reproduce it. Arithmetic, which I did run, was a three-way tie.

Coding agents

3/53 out of 5

Medium confidence

Sol and Fable lead several terminal and DeepSWE rows. Opus 5 is close on CursorBench at half Fable's cost. The right pick depends on the harness.

Cost per task

5/55 out of 5

High confidence

AA's index-task costs, the API rate cards, and my $312 example all point the same way. Grok is the cheap frontier option. Fable is twice Opus.

Speed

3/53 out of 5

Low confidence

Reported first-token and Fast-mode gaps exist. I did not time the trial. Do not pick a default from a speed claim I did not measure.

Trust and brand risk

3/53 out of 5

Medium confidence

Opus 5 and Sol are the boring procurement answers. Grok 4.6 may be fine internally and still fail a customer-facing review for reasons the Elo will not show.

Average
3.5 / 5
Binding constraint
The job and the bill, not the two-point index gap. Everyday writing did not separate the three models I ran. Cost per task and the coding-agent boards did.
Override applied
None. Split is already the conservative reading of a 3.5 average.

The average says the landscape is usable. The binding constraint says do not let the blend pick the tool.

How to try this yourself

You can replay the recorded trial without an API key. You still need three of your own files if you want to choose a default.

bash
git clone https://github.com/shravan1996/generative-ai-leaderboard-is-not-the-job
cd generative-ai-leaderboard-is-not-the-job
python3 -m unittest discover -s tests -t .
python3 eval.py

No network, no GPU, no third-party package. The full implementation is in the prototype repository.

  1. 01

    Name one job you did last week

    Owner
    You
    Artifact
    One sentence: the task, the input, the output you needed
    Signal
    If you cannot name it, you are shopping, not choosing
  2. 02

    Freeze three real examples

    Owner
    You
    Artifact
    Three anonymised prompts from actual work, not a toy
    Signal
    They should embarrass you slightly if they went wrong
  3. 03

    Run two models on the same three, same day

    Owner
    You
    Artifact
    Side-by-side outputs, no editing before you score
    Signal
    A difference you can feel without a rubric
  4. 04

    Price the month

    Owner
    You
    Artifact
    Last month's token volume at both rate cards, or the subscription you actually pay
    Signal
    A number large enough to care about, or not
  5. 05

    Keep, split, or revert after a week

    Owner
    You
    Artifact
    A one-line default and a date to look again
    Signal
    If the week did not show a difference, the index does not get a vote

Ship gates

  • The new default wins on your three examples, not on a screenshot
  • The monthly bill at realistic volume is acceptable
  • Anyone else on the team can find the same model without a private nickname

Kill criteria

  • A week of real work shows no difference you can feel
  • The bill doubles because the "cheaper" model used more turns
  • A customer-facing or regulated workflow fails a review the index never saw

Should you switch this week?

No, not because of the index. Switch if a named job is failing, or if the bill is the actual constraint, and then switch to the model that owns that job, not to whoever is bolded in this morning's table.

My verdict

Split for everyday knowledge work in August 2026.

As far as I can tell, the two-point index gap is not a reason to switch a team's default this week. Claude Opus 5 is the daily default I would hand a colleague who does not want to think about this. I expect it to stay there while it leads the independent index and costs half of Fable 5. Grok 4.6 is the model I would trial when the bill is the point, on the same three tasks, internally first. GPT-5.6 Sol is the model I would trial for long terminal and repo agents, because that is where the public rows keep handing it the lead. Claude Fable 5 is a 2x token premium. Pay it when you can point at a remaining gap. Do not pay it because the logo is heavier.

I would change this verdict if:

  1. A matched-rung independent board, dated after this piece, opens a gap on everyday writing large enough that a regular user notices it on three real files.
  2. Grok 4.6's cost-per-task advantage disappears once you count turns on the actual harness you use.
  3. Your own three-example test, not mine, cleanly prefers one model on the job you pay for.

What to remember

If you ran three real files and the ranking flipped, write to me. I would rather correct the default than defend the table.

Acknowledgements

This article builds on public boards and vendor posts. Digital Applied's 12 August fact-check of the Grok 4.6 launch table is the reason I trusted that table's cells and still refused its framing. Artificial Analysis, Anthropic, and VentureBeat are cited below for specific numbers.

References

Independent boards and fact-checks

Vendor posts

  • Introducing Claude Opus 5. Anthropic, 24 July 2026. Pricing, everyday-default positioning, CursorBench and Frontier-Bench claims, Fable 5 at twice the token price.
  • SpaceXAI debuts Grok 4.6. VentureBeat, 12 August 2026. API rates, GDPval / CursorBench / DeepSWE / Terminal-Bench rows, cost-per-task and turn counts, procurement caveats.

This trial

Sources last checked: 2026-08-15

A leaderboard is a photograph of someone else's test. Your week is the test that gets to vote.