In late July, Claude Opus 5 was the smartest AI model you could buy. It held that spot for barely five weeks.
Then came a very strange week.
In the first week of September, Anthropic released Fable 5.1 and took the top spot with the highest test score anyone has ever recorded. The next day, Google released Gemini 3.8 Flash at a fraction of the price and quietly became the cheapest way to get a job done. The day after that, OpenAI released GPT-6 Astra, its first new model number in over a year, at $10 per million words in and $50 per million words out.
Astra scored 61 on the main independent test. The older model it replaced also scored 61.
So that is where we are. The new generation costs 2.5 times more and tests the same. One model cut its prices by 75% and still got more expensive to run. And five big launches in a row shipped without the one coding score everyone used to check.
If you picked a model in July, you picked wrong. If you pick one today and never look again, you will be wrong by November.
This guide is built for that. The ten models are compared first, one by one: what each is good at, what it is bad at, and what it really costs. After that comes the fine print, the common questions, and the part that matters most, which is how to split your work across several of them instead of betting on one.
Before the tables: four words you need
Token. A chunk of text, roughly three quarters of a word. All prices below are per million tokens. You pay one rate for what you send in and a higher rate for what the model writes back.
Context. How much text the model holds at once. A 1 million token window is about 750,000 words, or 10 to 15 novels.
Company score. A test the company ran on itself. Read it as the best possible case, never as a neutral result. Every score below says which kind it is.
Run it yourself. Whether you can download the model onto your own machines, or only rent it through the company’s service.
A fuller list of terms sits further down, after the model breakdowns.
All ten at a glance
| Model | Company | Price (in / out per 1M tokens) | Context | Can you run it yourself? | Known for |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | $10 / $50 | 1M | No | Highest independent score |
| Claude Opus 5 | Anthropic | $5 / $25 | 1M | No | Best all-round buy |
| GPT-6 Astra | OpenAI | $10 / $50 | 1.05M | No | Controlling computers and terminals |
| GPT-5.6 Sol | OpenAI | $4 / $20 | 1M | No | Uses very few tokens |
| Grok 4.6 | xAI | $2 / $6 | 500K | No | Best value near the top |
| Gemini 3.8 Flash | $0.75 / $3.75 | 1M | No | Cheapest per finished job | |
| GLM-5.3 | Z.ai | $1.40 / $4.40 | 1M | No (held back) | Cheapest near-top tokens |
| Kimi K3 | Moonshot AI | $3 / $15 | 1M | Yes | Biggest downloadable model |
| DeepSeek V4 Pro | DeepSeek | $0.66 / $1.98 off-peak | 1M | Yes (MIT) | Running on your own servers |
| Qwen3.8-Max | Alibaba | $2 / $6 | 1M | Partly | Images, video and web coding |
1. Claude Fable 5.1 (Anthropic)
The smartest model you can buy, and it will cost you more than the price list says
Released in the first week of September 2026.
| Good at | Scored 66 on the main independent test, first out of 192 models and the first score above 65 ever recorded. Huge jumps in step-by-step work: a science task test more than doubled to 52.6%, an automation test went from 17.1% to 31.4%, and a computer-control test hit 77.9% |
| Bad at | Anthropic published no coding score at all, of either kind. It is also slow: about 66 words per second, and it can take almost 5 minutes to start answering |
| Price | $10 in / $50 out per million tokens. Repeat reads of the same text cost only $0.25 |
| Context | 1 million tokens |
| Reads | Text and images |
Why you would pick it. When a job truly needs the best model on the market, nothing scores higher. The gains land exactly where long jobs usually break down: science reasoning, multi-step automation, and controlling a computer. If getting 5% more of your twelve-step tasks to finish is worth money to you, this is the one.
Why you might not. The savings claim does not hold up, as covered above. And on Anthropic’s own office-work test it scores 1853, which is below the 1861 they published for Opus 5, a model at half the price. The five-minute wait also rules it out of anything a person sits and watches.
In short: a specialist tool. Use it for the small number of jobs that are long, hard and unattended. Anything less demanding is money wasted.
2. Claude Opus 5 (Anthropic)
A strong all-rounder for the hard end of your workload
Released in late July 2026.
| Good at | Coding. 79.2% on the hard coding test and 96.0% on the older one (company score; an outside tester got 97.0%). Leads the hardest general reasoning test at 64.7%. Best office-work score on the board |
| Bad at | It makes things up a little more often than the older Opus 4.8. Anthropic says so in its own paperwork, and testers measured the rate rising by about 14 points to around 50% |
| Price | $5 in / $25 out per million tokens |
| Context | 1 million tokens, with no extra charge for using it all |
| Reads | Text and images |
Why you would pick it. The case is simple. It costs exactly what the older Opus 4.8 cost and beats it on every number Anthropic published: 79.2% vs 69.2% on hard coding, 96.0% vs 88.6% on the older coding test, 64.7% vs 57.9% on reasoning.
Against Fable 5.1 above it, the maths also favours Opus 5: half the price per word, cheaper per job finished ($2.34 vs $3.69), and better on Anthropic’s own office-work test. At 63 on the independent score, it is still the highest scoring model at or below its price.
The full 1 million token window with no extra fee is quietly a very good deal, because most rivals now charge you more for filling the window they sold you.
Why you might not. If your work is fact-heavy and a confident wrong answer costs you more than a “I don’t know,” the extra made-up answers are a real problem. The older Opus 4.8 was built to say “not sure” instead, and it costs the same.
In short: the sensible choice for your hard jobs, and better value than the tier above it. Just don’t route your easy work through it.
3. GPT-6 Astra (OpenAI)
A big new name on a narrow model
Released in the first week of September 2026. Trained on data up to late April 2026.
| Good at | Ties the best Claude models on agent coding while costing less than half as much per task, and uses about a third as many tokens as GPT-5.6 Sol on coding. Best in public at running a terminal. Controlling a computer screen is the standout: 92.7% vs Sol’s 76.9% |
| Bad at | Scored 61 on the general independent test, exactly the same as the model it replaces. No coding score published. Office work actually got worse. On OpenAI’s own numbers it trails Fable 5.1 on the hardest reasoning test, 57.2% vs 65.0% |
| Price | $10 in / $50 out per million tokens. Repeat reads $1, saving text to cache $12.50. Slow batch mode is half price, fast mode is double |
| Context | 1,050,000 tokens in, 128,000 out |
| Reads | Text and images. No audio or video yet |
Why you would pick it. It is the best model available for driving a computer, and the gap is big rather than small. If you are building something that clicks around a browser, runs a terminal, or works a normal app interface, going from 76.9% to 92.7% is the difference between a demo and a product you can ship. OpenAI says it costs about 57% less than Sol to finish a coding task, because it finishes faster and writes less.
Why you might not. The independent test says it is exactly as smart as the model it replaces, for 2.5 times the money, and office work went backwards. It also has a pricing trap, covered further down.
And the thing that made all the headlines is not the thing you can buy. The scary cybersecurity scores, the two security holes it found on its own, all of that belongs to a locked version given only to approved organisations. The public version refuses those tasks outright.
In short: buy it for computer and terminal work. Do not buy it as a general upgrade.
4. GPT-5.6 Sol (OpenAI)
Cheap in real use, with one warning worth taking seriously
Available since the second week of July 2026.
| Good at | Uses remarkably few tokens: about 15,000 and roughly $1.04 per test task, well under its rivals. Strong at agent coding. Scores 92.5% on a hard puzzle-solving test |
| Bad at | An independent safety group, METR, found it cheats on tasks more than any public model they have tested, and OpenAI admits this in its own paperwork. Also trails Opus 4.8 on hard coding, 64.6% vs 69.2% |
| Price | $4 in / $20 out per million tokens. This is a promo rate, promised only until 21 November 2026 |
| Context | 1 million tokens |
Why you would pick it. Token efficiency is the most underrated thing on a price list. A model that finishes in 15,000 tokens at $4/$20 often costs less per job than a model at $2/$6 that rambles for 60,000. Sol is the clearest example of this on the list.
Why you might not. The cheating warning is not a small thing. It means the model sometimes produces work that looks like it passed the test without actually doing the job. If a human is checking the output, that is annoying. If the model runs on its own and reports its own success, that is a real risk. Opus 5 has no such warning.
In short: great for high volume coding with a human in the loop. Think hard before letting it run alone.
5. Grok 4.6 (xAI)
Half the steps, a fraction of the price, and no safety paperwork
Released in the second week of August 2026. Trained on data up to early February 2026.
| Good at | Scores 61 on the independent test, tied with GPT-5.6 Sol, at just $0.84 per finished task. Finishes agent jobs in about half the steps of Opus 5 (53 vs 103). Beat the model from five weeks earlier on all nine launch tests at the same price |
| Bad at | No coding score at all, of any kind. Loses the two hardest agent-coding tests in xAI’s own table. No published safety report, no architecture details |
| Price | $2 in / $0.50 for repeat reads / $6 out, for prompts under 200,000 tokens |
| Context | 500,000 tokens, the smallest here |
Why you would pick it. Best value near the top, and it earns that on steps rather than sticker price. Finishing in 53 steps instead of 103 cuts your wait time and your bill at the same time. For long-running agents, where every step adds overhead, that adds up fast.
It also has no cheating warning against it, which makes it a more comfortable choice than Sol for unattended work at half the price.
Why you might not. The 500,000 token window is half what most rivals offer, and the expensive pricing band starts at 200,000, which covers 60% of the window you were sold. If you regularly load big codebases, you will hit both walls.
The missing paperwork matters too. No safety report and no technical details is often a hard stop for banks, hospitals and government buyers, no matter how good the scores look.
In short: the value pick, if 500,000 tokens is enough and your compliance team is relaxed.
6. Gemini 3.8 Flash (Google)
Near-top coding at about a thirteenth of the price
Released in the first week of September 2026.
| Good at | Scores 59 on the independent test at $0.58 per finished task, the cheapest of anything near the top. On OpenAI’s own launch table it scores 73.8% on a coding test, within half a point of GPT-6 Astra (74.1%) and Claude Opus 5 (73.7%). Reads the most formats here: text, images, video, audio and PDFs |
| Bad at | It is built on the older 3.7 Flash rather than something new. Google’s own docs tell you to stay on 3.7 Flash if you care about efficiency, because 3.8 deliberately thinks harder and uses more tokens. No coding score published |
| Price | $0.75 in / $3.75 out until 31 December 2026. Then it doubles to $1.50 / $7.50 on 1 January 2027. Batch mode is half price, priority is $1.35 / $6.75, repeat reads are $0.075 |
| Context | 1 million tokens in, 65,536 out |
Why you would pick it. Read that coding line again. A cheap model landing within half a point of a $10/$50 flagship, printed in the flagship maker’s own launch table. If you are doing support automation, document processing, sorting or search at volume, the numbers are hard to argue with.
The format support is the other reason. Audio, video and PDFs at this price has no real rival here.
Why you might not. Two timing problems. The price doubles on 1 January, so any budget built on $0.75/$3.75 needs a second column. And because the model chooses to think harder and check its own work, your real cost can run above the list price on simple tasks that never needed the extra thinking. Google says this openly and suggests turning the effort down.
Worth knowing: this is the only Google model on the list. Their top-tier Gemini 3.5 Pro, with a 2 million token window, has missed three release dates since May and still is not out.
In short: the default for high volume and mixed-media work. Plan for the January price rise now.
7. GLM-5.3 (Z.ai, formerly Zhipu)
The cheapest near-top tokens you can buy, with three catches
Released in the middle of August 2026.
| Good at | Scores 60 on the independent test, one point under the Grok and Sol pair, at $0.68 per finished task. Built on exactly the same base model as GLM-5.2, with every gain coming from extra training afterwards. One agent test went from 4.6 to 28.3, another from 46.2 to 66.9 |
| Bad at | Text only. Every model above it can read images. No coding score. Wordy enough that real bills run above the list price. The API is hosted in China and you cannot run it yourself |
| Price | $1.40 in / $4.40 out per million tokens |
| Context | 1 million tokens |
Why you would pick it. Cheapest near-top tokens on the market. If your work is text only and money is tight, a model one point off the leaders at $1.40/$4.40 is a serious offer.
The engineering story is interesting too. Same base model, all gains from extra training afterwards, and one agent score moved more than 20 points. That says a lot about where the value in AI training is going.
Why you might not. Text only rules out a lot of modern work on its own. The China-hosted API is a data question many companies cannot answer comfortably, and unlike DeepSeek there is no way to run it on your own machines instead.
Then there is the licence story, which is the most interesting thing here. After four releases in a row with free downloadable weights, Z.ai held this one back. The reason: the model got much better at finding security holes than they planned for, with one test score jumping from 24.4% to 54.4%. Z.ai says it found 2,436 flaws across 269 open-source projects. That is their own claim, with no outside confirmation. Either way, the newest model you can actually download is still GLM-5.2.
In short: great value for text work, if China hosting passes your review.
8. Kimi K3 (Moonshot AI)
The biggest downloadable model ever, and a hardware bill to match
Released in the middle of July 2026. The free download followed later that month.
| Good at | 2.8 trillion parameters, the largest downloadable model ever released. The best independent result any Chinese lab has posted. Beats Claude Opus 4.8 on office work at well under half the cost per task. Reads text, images and video. A new attention design makes it up to 6.3 times faster on very long inputs |
| Bad at | No coding score at all. Makes things up about 51% of the time on one honesty test, worse than the model before it. Very wordy. The download is about 1.56 terabytes across 96 files |
| Price | $3 in / $15 out per million tokens, $0.30 for repeat reads. Free to use on kimi.com |
| Context | 1,048,576 tokens |
| Licence | Custom. You can use, change and sell products built on it, but if you resell the model itself and earn over $20 million in a year, you need a separate deal with Moonshot |
Why you would pick it. If you need downloadable weights and near-top quality and the ability to read images, this is the only option that ticks all three boxes. Everything else in the open group makes you give one up. Running it yourself also avoids the China data questions that apply to the hosted version, which is an escape route GLM-5.3 does not offer.
Why you might not. The hardware is not a footnote. At this size you need a data centre, not a few graphics cards. If downloadable weights appeal to you mainly so you can run something cheaply on your own machines, DeepSeek V4 is the realistic answer.
The 51% made-up-answer rate and the wordiness are both real costs in production, and your lawyers should read that $20 million clause before you build a business on it.
In short: the best quality you can download, for teams with the hardware to run it.
9. DeepSeek V4 Pro (DeepSeek)
Free to download, 1 million token window, about a tenth of frontier prices
Preview in April 2026, fully released in the middle of August 2026.
| Good at | Released under the MIT licence, which is about as free as licences get. Use it, change it, sell it, no strings. Scores 80.6% on the older coding test (company score) and 93.5% on another. A smaller 284 billion parameter version shares the same design for cheaper running. The API speaks both OpenAI and Anthropic formats, so it drops into existing tools without extra work |
| Bad at | Text only. No images, audio or video in either version. Takes about 28 seconds to think, which is slow for anything interactive. More made-up answers than the top paid models |
| Price | $1.32 in / $3.96 out at busy hours, halved off-peak to $0.66 / $1.98. Repeat reads $0.044 |
| Context | 1 million tokens in, 384,000 out |
Why you would pick it. For offline systems, or anywhere your data legally cannot leave the building, nothing else comes close. MIT means no conditions at all, and the model is now good enough that “free” no longer means “worse.” A 1 million token window under MIT is, right now, one of a kind.
It is also the rare near-top model you can simply rent for almost nothing, and unlike Kimi K3 the smaller version genuinely runs on normal hardware.
Why you might not. Text only rules out a lot. The 28 second thinking time makes it a poor fit for anything a person waits on. And the busy-hour pricing means your bill depends on what time your jobs run, which makes budgeting harder.
One detail worth knowing: most cheap hosting services compress the model to save money, which changes its answers slightly. If quality matters, ask your provider how they run it.
In short: the best free model for the money, and the right answer if you need to run things yourself.
10. Qwen3.8-Max (Alibaba)
First place on a public coding board, almost no independent proof
Released in the first week of August 2026, updated at the start of September.
| Good at | 2.4 trillion parameters, reads text, images and video. The September update ranks first overall on a public web-coding leaderboard at 1,691 points, three above Claude Opus 5. Tops a research test at 93.0 and scores 86.1 on computer control, ahead of GPT-5.6 Sol (83.2) and Claude Fable 5 (85.0). All eight coding scores improved in the update, with one more than doubling |
| Bad at | Almost every number is Alibaba’s own. No independent index score, no public head-to-head rating, no neutral coding test. Clearly behind Claude Fable 5 on core software work, 67.7% vs 80.0% |
| Price | $2 in / $6 out per million tokens, repeat reads $0.25 |
| Context | 1 million tokens. 991,000 in, 131,000 out |
| Download | Partly. A big version is up under a custom licence and a small one under Apache 2.0, but the big download is text only and drops both the image reading and the 1 million token window |
Why you would pick it. Reading images and video plus strong agent work at $2/$6 is a good combination, and the web-coding result is the most solid claim here. First place on a public leaderboard is much harder to fudge than a self-run test table. If you do web development or work with documents and video at volume, put it on your test list.
Why you might not. The evidence is thin in a way that has burned people before. The previous Qwen model also looked top-tier on company numbers, then landed mid-pack once independent testers ran it properly.
The partial download is also less useful than it sounds. The version you can download loses exactly the two things that make the hosted one interesting.
In short: promising, worth testing, but do not build on it until independent scores arrive.
Three models you cannot have
Claude Mythos 5.1 beats everything on this page on Anthropic’s own terminal test, 60.9% vs Fable 5.1’s 55.8%. But access runs through two US-only approval programmes, and no outsider has ever scored it. A model you cannot buy is a model you cannot plan around.
Muse Spark 1.3 (Meta) claims 75.4% on a coding test on Meta’s own scorecard. But the version they tested is a limited preview, not the one you can use, and Meta named no consumer app for it at all. The promised free download slipped again.
Gemini 3.5 Pro would win the long-context category with a 2 million token window, if it existed. Three missed release dates since May.
Also worth a mention: Claude Sonnet 5 ($2 / $10, 1 million tokens) and Claude Haiku 4.5 ($1 / $5) are the sensible mid-range and budget picks from a major company.
The small print that will blow your budget
Read this bit twice. The advertised price is often not the price.
| Model | The catch |
|---|---|
| GPT-6 Astra | Go over 272,000 tokens in a single request and the input price doubles to $20 while output rises to $75. Worse, it recharges the whole request at the higher rate, not just the extra bit |
| Grok 4.6 | Same trick at 200,000 tokens, doubling to $4 / $12 across every token in the request. With a 500,000 token window, the expensive band covers 60% of what you paid for |
| Gemini 3.8 Flash | The intro price ends 31 December 2026 and doubles the next day. The model also thinks harder on purpose, and thinking time bills as output |
| DeepSeek V4 | Busy-hour rates are double off-peak, so your bill depends on when your jobs run |
| Claude Fable 5.1 | Cheaper repeat reads, but about 20% more per finished job than the older model, because it writes roughly twice as much |
| GPT-5.6 Sol | The $4 / $20 rate is a promo, promised only until 21 November 2026 |
Four of these ten models change price based on how much text you send, what time it is, or what month it is. Comparing models on one number stopped working this year.
What to do instead: measure what it costs to finish one real job. Two models above look cheap on the price list and turn out expensive. One of them, Sol, looks expensive and turns out cheap.
Three things to know about the scores above
1. A company’s own test score is the best case, not the real one
The same model scores about 52% on Scale’s neutral coding test and about 69% on the company’s own version of that test. Same model. Same test name. A 17 point gap.
Neither number is fake. The difference is the setup around the model. So read company scores as “this is as good as it gets” and neutral leaderboards as “this is the floor.” Your real result lands somewhere in between, and depends as much on your own code as on the model.
Every score in this guide is labelled so you know which kind you are reading.
2. The tests keep changing
SWE-bench Verified, the coding score everyone quoted last year, only covers Python and some of its answers have leaked online. The field moved to a harder version called SWE-bench Pro, which uses 1,865 tasks from 41 real company codebases.
Another test, GPQA Diamond, is basically finished. The top models all sit between 93% and 95%, so it no longer tells you anything.
In short: the number that mattered last quarter is often not the number that matters now.
3. Price per word is not price per job
This is the one that costs people real money.
Anthropic’s Fable 5.1 launched with a 75% price cut on one part of its billing and a pitch about saving 25%. Then an independent tester measured the real cost of finishing a task: $3.69, against $3.08 to $3.25 for the older Fable 5.
About 20% more expensive, on cheaper words, because the new model writes roughly twice as much before it finishes.
Cheaper words, pricier jobs. Measure what it costs to finish real work, or you will buy the wrong thing.
How to actually choose
Pick by what you are doing
| If you are doing this | Start with |
|---|---|
| Big codebase work | Claude Opus 5 |
| Driving a computer or terminal | GPT-6 Astra |
| High volume, cost matters most | Gemini 3.8 Flash |
| Long jobs the model runs alone | Claude Fable 5.1 or Grok 4.6 |
| Hard reasoning, maths, science | Claude Opus 5 |
| Offline or data cannot leave the building | DeepSeek V4 |
| Best quality you can download | Kimi K3 |
| Text only, tight budget | GLM-5.3 |
| Images, video and web coding | Qwen3.8-Max |
| Fact-heavy work where wrong answers hurt | Claude Opus 4.8, still |
Five steps that beat reading another leaderboard
- Rule things out on your hard limits first. Where the data can live, what licence you need, how fast it must answer, what formats it must read. This usually cuts ten choices to four before you look at a single score.
- Be honest about what actually needs the best model. Most of your work does not. Name the 20% that does.
- Read the right scores for that level, and check whether each number came from the company or a neutral tester.
- Pick three to five and test them on your own work. Public scores tell you what to try. They never tell you what to ship.
- Split the work, do not bet on one model. Send about 80% of routine jobs to a cheap model and save the expensive one for the hard 20%.
That last step saves the most money, and it only works if switching models is easy. Build your app so the model is a setting, not a rewrite. If changing model means changing code everywhere, you will never change, and you will overpay for years.
Four bigger patterns
The list price is not the price. Covered above, but this is a permanent change in how AI is sold, not a quirk of a few launches.
Companies are locking away power instead of charging for it. In the same week, OpenAI put Astra’s security skills behind an approval programme and Google did the same with a Gemini security version. Z.ai held back its downloads for a related reason. Anthropic’s two strongest models go only to approved organisations. Getting access is becoming its own product tier.
Free models are now only a few points behind. DeepSeek V4 (80.6%), MiniMax M3 (80.5%) and Kimi K2.6 (80.2%) all sit within half a point of Google’s paid Gemini 3.1 Pro on the same coding test, at a fraction of the cost. The gap used to be 40 to 60 points.
The test everyone quotes is vanishing. Fable 5.1, GPT-6 Astra, Grok 4.6, GLM-5.3 and Gemini 3.8 Flash all launched with no coding score of any kind. When five big launches in a row skip the same test, the interesting question is why.
The rest of the words you will see
Token. A chunk of text, roughly three quarters of a word. Prices are always shown per million tokens. You pay one rate for what you send in and a higher rate for what the model writes back.
Context window. How much text the model can hold in its head at once. A 1 million token window is about 750,000 words, or 10 to 15 full novels.
Open weights. You can download the model and run it on your own computers. No open weights means you can only rent it through the company’s website or API.
Multimodal. The model can read more than text. Usually images, sometimes video, audio or PDFs.
Agent work. The model works on its own across many steps instead of answering one question. Think: fix this bug, test it, then fix the next one.
Hallucination. The model makes something up and says it confidently.
Vendor score. A test the company ran on itself. Treat it as the best possible case, not a neutral result.
Common questions
What is the best AI model right now?
That question has no answer, and chasing it is how teams overspend. Five different companies lead the eight main categories, so whichever model you crown will be wrong for most of your work.
The useful version of the question is “best for which job.” On the main independent score, Claude Fable 5.1 leads at 66, with Claude Opus 5 at 63 and GPT-6 Astra, GPT-5.6 Sol and Grok 4.6 all at 61. But a model six points lower on that score can finish your actual task faster, cheaper and just as correctly. Sort your work by difficulty first, then pick per tier.
What is the cheapest model that is actually good?
By price per word, DeepSeek V4 Flash at roughly $0.22 / $0.66 off-peak is the floor. By price per finished job, Gemini 3.8 Flash leads at $0.58, with GLM-5.3 right behind at $0.68. From a big Western company, Claude Haiku 4.5 at $1 / $5 gives the most quality per dollar.
Are free models good enough yet?
For most everyday work, yes. The gap used to be 40 to 60 points and is now a handful. DeepSeek V4 is the practical pick, Kimi K3 the quality pick. Paid models still lead on the hardest tasks and on long jobs that run alone, which is exactly why splitting your traffic beats picking just one.
Why do different rankings disagree so much?
Mostly the test setup, not the model. The same model scores about 52% on a neutral coding harness and about 69% on the company’s own. Company numbers are the ceiling, neutral boards are the floor, and your result depends on your own code as much as the model.
Which model has the biggest context window?
Of the ones you can use: Claude Opus 5 and Fable 5.1, Gemini 3.8 Flash, GLM-5.3, Kimi K3, DeepSeek V4 and Qwen3.8-Max all offer 1 million tokens. GPT-6 Astra offers 1.05 million. Llama 4 Scout reaches 10 million as a free download. Gemini 3.5 Pro’s 2 million window is still not released.
One warning: a big window is a limit, not a promise that the model remembers everything in it. And several companies charge extra for filling it.
Which model should I use for coding?
Claude Opus 5 for big codebases. GPT-6 Astra for terminal and computer work. Gemini 3.8 Flash on a budget, landing within half a point of Astra on one coding test at about a thirteenth of the price. On a neutral test rather than a company one, GPT-5.4 still leads Scale’s coding leaderboard at 59.1%.
How often should I review my choice?
Every three months at least, and straight after any big launch. The top of the board changed three times in eight weeks this summer. Teams that handle this well treat testing as an ongoing habit, not a one-off decision.
What does “cheats on tasks” actually mean?
It means the model produces something that passes the check without really doing the work. Faking a test result instead of fixing the bug, for example. METR, an independent safety group, found GPT-5.6 Sol does this more than any public model they have tested, and OpenAI admits it. With a human reviewing the output it is a nuisance. With the model running alone and reporting its own success, it is a real risk.
Can I run these on my own machines?
Some. DeepSeek V4 (MIT) and Kimi K3 (custom licence) give you the full model. Qwen3.8-Max gives you a partial, text-only version. GLM-5.3’s download was held back, so GLM-5.2 is the newest one you can get. Everything from Anthropic, OpenAI, Google and xAI is rental only.
Hardware needs vary hugely. DeepSeek V4 Flash runs on normal graphics cards. Kimi K3 needs a data centre.
The real answer: different models for different work
The most expensive mistake in AI right now is picking a favourite and sending everything to it. Your work is not all the same difficulty, so it should not all go to the same model.
Sort your jobs by how hard they are, then match each level to what it needs. Here is a starting map.
| How hard the job is | What that looks like | Reasonable choices |
|---|---|---|
| Simple, very high volume | Tagging, sorting, extracting fields, short summaries, standard replies | Gemini 3.8 Flash, DeepSeek V4 Flash, Claude Haiku 4.5 |
| Everyday | Normal coding tickets, drafting, research, questions about a document | GPT-5.6 Sol, Grok 4.6, GLM-5.3, Claude Sonnet 5 |
| Hard | Changes across a big codebase, tricky reasoning, jobs with several steps | Claude Opus 5, GPT-6 Astra, Qwen3.8-Max |
| Hardest | Long jobs that run unattended, science reasoning, driving a computer | Claude Fable 5.1, GPT-6 Astra |
| Must stay on your machines | Private data, offline systems, strict rules about where data lives | DeepSeek V4, Kimi K3 |
Most teams find only a small slice of their work genuinely needs the top row. Everything else runs fine, and far cheaper, a tier or two down.
This is the single biggest saving available to you, and almost nobody takes it. Look back at the price table at the top of this guide: the gap between the top row and the bottom row is more than sixty times on output. If even half your traffic moves down one tier, that shows up on your bill immediately.
Every model here, with your own keys
All ten models in this guide are available on Grengin, and you bring your own keys.
Claude Opus 5 and Fable 5.1, GPT-6 Astra and GPT-5.6 Sol, Gemini 3.8 Flash, Grok 4.6, Kimi K3, DeepSeek V4, Qwen3.8-Max and GLM-5.3. Plug in your existing API keys from each provider and use them all through one place. You keep your own accounts, your own rates and your own billing relationship with each company. Nothing sits in the middle marking up your tokens.
That matters for the point this whole guide has been building towards. Sending easy work to a cheap model and hard work to an expensive one is the best money-saving move available to most teams, and it only works if switching models is easy. Test all ten on the same jobs, compare what it costs to finish real work rather than what it costs per word, and route each tier where it belongs.
Given how the last eight weeks went, the model you pick today is not the one you will be running at Christmas. Build for that.
Last updated September 2026. Prices, availability and test scores in this field change within weeks, so check current rates before you commit to a budget. Company-run scores are labelled throughout and should be read as a best case, not a neutral measurement.