15 AI models ranked for real work
A practical ranking of current AI models using effectiveness, ease of use, and cost/value. It includes the reasoning behind every placement and the prices available on 23 August 2026.
This is a ranking for choosing a model to do real work. It covers fifteen current models from Anthropic, OpenAI, Moonshot AI, Z.ai, Google, SpaceXAI, DeepSeek, Cursor, and Meta.
The table is not a benchmark leaderboard. Each developer tests its models with different harnesses, budgets, effort settings, and versions. Those results help us understand what a model is built to do. They do not produce one fair score that settles the order.
The placements combine published facts with our experience using models on long, practical tasks. That distinction matters most for Claude Opus 5. Anthropic presents it as close to Fable capability at half Fable's token price. Our experience has been much worse, so its F tier is our verdict rather than a claim that the wider evidence proves.
How we ranked them
Effectiveness carries 50%
The first question is whether the model completes the task well enough to use. We looked at first-pass quality, reasoning, coding ability, reliability across long tasks, and how often the result survives checking. A clever answer that needs rebuilding does not score well here.
Ease of use carries 30%
Ease includes availability, speed, integration breadth, product maturity, and steering burden. A model loses ground when access is narrow, the product is still a preview, or repeated corrections become part of every task.
Cost/value carries 20%
We used current standard API prices where they were public, with temporary prices dated in the source record. Price is judged against the nearest useful alternative. Saving a few dollars on tokens does not help if a person spends an extra hour repairing the answer.
The weights guide the judgement. They are not fed into a formula that pretends subjective experience has decimal-point precision.
S tier
Claude Fable 5
Fable 5 is the model we trust most when the task is difficult, long, and expensive to get wrong. It holds context, follows the intended direction, and reaches usable work with fewer interventions than the rest of the field. Its $10 per million input token price and $50 per million output token price are the weakness. It is the most expensive generally available model in this list. It still takes S because effectiveness carries half the decision and because correction time costs more than tokens on serious work.
A tier
GPT-5.6 Sol
Sol is the strongest all-round alternative to Fable 5 in our judgement. It is widely available across OpenAI's products and API, handles complex professional work, and brings a broad tool set with a 1.05 million token context window. OpenAI announced a temporary price reduction on 21 August, with its model guide showing $4 input and $20 output at the time of this review. In our use, the hardest tasks still need more checking than they do with Fable, which keeps Sol out of S.
Kimi K3
Kimi K3 combines frontier-level capability, open weights, native vision, and a 1 million token context window. It is available through Kimi's apps, coding product, and API at $3 input and $15 output. It does not yet match the product maturity or broad business adoption of the two leaders, but the capability-to-price ratio earns A.
GLM-5.3
GLM-5.3 is the placement most likely to surprise people. Z.ai's published results put it among the strongest current models for agentic coding and knowledge work, while its listed API rate of $1.40 input and $4.40 output is far below the other frontier models. Access and independent operating history are thinner, which keeps it out of S. The value case is too strong to leave it in the lower half.
Gemini 3.7 Flash
Google's newest Flash model is generally available rather than preview-only. It is built for coding, agents, and multi-step work, includes a 1 million token context window, and costs $0.75 input and $3.75 output under the introductory rate running through 2026. Google now recommends it as the migration target from Gemini 3.1 Pro. It belongs in A because it combines an easy route into the product with unusually strong value.
B tier
Grok 4.6
Grok 4.6 is strong across coding and knowledge work and is available through Grok, the API, Cursor, and partner platforms. The base API rate of $2 input and $6 output is competitive. Our B-tier judgement reflects less consistent results across the tasks we use and a surrounding product that feels less settled for controlled business work than the models above it.
Claude Sonnet 5
Sonnet 5 is fast, broadly available, and much cheaper than Fable. Its price is $2 input and $10 output. Anthropic made those rates permanent on 10 August. It is a capable default for everyday work. B reflects the gap between being a good default and being the model we choose when the result has to be right the first time.
GPT-5.6 Luna
Luna is the value specialist in OpenAI's current family. It costs $0.20 input and $1.20 output, supports the same broad tool set as the larger variants, and is easy to access at scale. It earns B by making routine and high-volume work economical. In our use, the capability ceiling arrives sooner on difficult tasks, which stops it moving higher.
DeepSeek V4 Pro
DeepSeek presents V4 Pro as its stronger reasoning and agentic coding model. It charges $0.66 input and $1.98 output off-peak, or $1.32 and $3.96 during its weekday peak windows. Those rates remain low beside the expensive frontier group. Our B-tier judgement accounts for the value while leaving room for questions around product maturity, operational fit, and whether a business is comfortable placing sensitive work with the provider.
C tier
GPT-5.6 Terra
Terra is OpenAI's balanced model at $2 input and $12 output. It is straightforward to use and sits inside the same mature tool ecosystem as Sol and Luna. OpenAI positions Sol for frontier capability and Luna for cost-sensitive volume. In our use, that leaves Terra as a sensible model that is rarely the obvious selection.
DeepSeek V4 Flash
V4 Flash is cheap even at DeepSeek's weekday peak rate. It costs $0.22 input and $0.66 output off-peak, or $0.44 and $1.32 at peak, with a 1 million token context window and familiar API formats. It is an excellent high-volume option when tasks are bounded and easy to verify. In our tests, longer reasoning work creates enough correction burden to hold it in C despite the price.
Composer 2.5
Composer 2.5 is Cursor's own coding model rather than a general model from SpaceXAI. It is tuned tightly for Cursor's tools and costs $0.50 input and $2.50 output in standard mode. That integration makes it effective inside its home product and awkward outside it. C is a portability judgement more than a criticism of the model.
D tier
Muse Spark 1.2
Meta positions Muse Spark 1.2 as a coding model for long agentic tasks, paired with Muse Code beta and expanded global access through the Meta Model API. It may rise quickly. Today the product is very new and a stable public token price was not clear in the official material we reviewed. D reflects low confidence and product maturity rather than a settled view of the underlying capability.
Gemini 3.1 Pro Preview
The word Preview belongs in the name. Gemini 3.1 Pro is capable, particularly on multimodal reasoning, but Google now points developers towards the newer generally available 3.7 Flash. The older preview costs more, with rates rising again above 200,000 input tokens. There is little practical reason to choose it now, which puts it in D.
F tier
Claude Opus 5
Opus 5 is the controversial placement and the one most dependent on our own use. Anthropic prices it at $5 input and $25 output and presents it as close to Fable capability at half Fable's price. Its published benchmark case is strong.
Our experience does not match that case. Opus 5 has been harder to steer, more prone to returning work that looks complete before it is checked, and more expensive in correction time than its token price suggests. A model can be powerful on a test and poor value in a workflow. For us, the repeated failure to produce dependable usable work puts it in F.
How to use the table
Start with the job rather than the tier. Use Fable or Sol when the task is difficult and checking mistakes is expensive. Use Gemini 3.7 Flash, Luna, or DeepSeek when volume and cost matter more. Keep preview products away from critical workflows until their access, pricing, and behaviour settle.
This ranking is a snapshot dated 23 August 2026. Model prices and product access move quickly. The durable method is to test representative work, record the correction burden, and judge the total cost of reaching a usable result.
Primary sources
Prices and availability were checked against the developers' own material on 23 August 2026. These are the pages used for the publishable claims in this article.
- Anthropic model overview and pricing
- Claude Fable 5 release
- Claude Opus 5 release and Fable comparison
- Claude Sonnet 5 release
- Claude Platform release notes
- OpenAI GPT-5.6 model guide
- OpenAI GPT-5.6 launch and 21 August price update
- OpenAI GPT-5.6 Luna model page
- Moonshot AI Kimi K3 release
- Z.ai GLM-5.3 release
- Z.ai model pricing
- Google Gemini model guide
- Google Gemini 3.7 Flash and migration guidance
- Google Gemini pricing
- DeepSeek V4 pricing
- DeepSeek V4 model positioning
- SpaceXAI Grok 4.6 release
- SpaceXAI Grok pricing
- Cursor Composer 2.5 model guide
- Meta Muse Code and Muse Spark 1.2 release