Every month I update this page with the best AI right now, plus the cheapest one for high-volume work and the open model in the lead. The scores come from the independent Artificial Analysis ranking. The judgement is mine, the same one I use to close the group’s daily message, which goes out in Spanish.
One thing before the table: first place alone does not tell you which model suits your shop or office. Compare the price, usage limits, where your data goes and whether the AI sits inside an app you already use.
The best AI right now: my October 2026 ranking
This is how it stood on 1 October 2026:
| For | Model | Why |
|---|---|---|
| Best on the market | Claude Opus 5.5, by Anthropic | 58 points on Artificial Analysis, the highest score on their current tests |
| Best value for medium tasks | GPT-6.1 Sol, by OpenAI | 52 points at $0.72 per task; Opus 5.5 at max costs $5.98 |
| Best value for simple, high-volume tasks | GPT-6 Luna, by OpenAI | $0.10 per million input tokens and $0.50 per million output tokens |
| Best open model | MiMo-V2.6-Pro, by Xiaomi | 46 points, the best of the models you can download |
| Favourite harness (the program you work with agents in) | Claude Code | With Sonnet 5.5 at max effort, first on that index |
This month the second line changes. GPT-6.1 Sol, which OpenAI released on 29 September 2026, scores 52, six points below Opus 5.5, but in Artificial Analysis’s tests it costs $0.72 per task at max, against $5.98 for Opus 5.5 at max. Those are API costs in those tests: they do not translate into tasks included in a subscription. You get it on ChatGPT’s paid plans, Plus included. Opus 5.5 still leads this index and comes with Claude Pro, at €18 a month plus VAT (about €22). If you already pay for Claude, try Sonnet 5.5 for everyday work, on medium effort. Anthropic recommends lower effort for a lower cost per task; Opus remains stronger at complex work. Check usage in your account.
If you don’t want to pay. Free Claude uses Sonnet and Haiku, not Opus 5.5. For me, the most complete free option today is ChatGPT, which since 22 September 2026 lets you try GPT-6 Luna in Work and Codex, inside its desktop app. If you work in Gmail and Drive, try Gemini. And whichever you use, switch off training on your conversations.
Scores from the Artificial Analysis main leaderboard, on index v4.3.2: max effort where available, high for both Gemini models and xhigh for Grok. Claude has safety fallback enabled, which routes some requests to another model. These are specific configurations, not a score for every answer from the app:
| Model | Score | Price per million tokens |
|---|---|---|
| Claude Opus 5.5 (Anthropic) | 58 | $4 / $20 |
| Claude Sonnet 5.5 (Anthropic) | 56 | $2 / $10 |
| Claude Fable 5.1 (Anthropic) | 53 | $10 / $50 |
| GPT-6 Astra (OpenAI) | 53 | $10 / $50 |
| Gemini 4 Argon (Google) | 53 | $2 / $10 |
| GPT-6.1 Sol (OpenAI) | 52 | $2 / $10 |
| Muse Spark 1.3 (Meta) | 48 | $1.25 / $4.25 |
| Grok 4.7 (SpaceXAI) | 46 | $2 / $6 |
| MiMo-V2.6-Pro (Xiaomi) | 46 | $0.435 / $0.87 |
| GLM-5.3 (Zhipu, Z.ai) | 45 | $1.40 / $4.40 |
| Kimi K3 (Moonshot) | 44 | $3 / $15 |
| Gemini 3.8 Flash (Google) | 41 | $0.75 / $3.75 |
| GPT-6 Luna (OpenAI) | 37 | $0.10 / $0.50 |
API prices in dollars, before VAT: input / output, standard rates excluding caching, batches and tool surcharges. OpenAI uses the short-context tier; long context costs more. Meta is the Standard tier and Xiaomi the international rate. Google prices are promotional: Gemini 3.8 Flash until 31 December 2026; Argon will later cost $4 / $20. Argon remains limited to trusted testers and cyber defenders, with no public date for wider access.
What changed this month
Since the previous ranking, on 24 September 2026:
- 28 September 2026. Anthropic releases Claude Sonnet 5.5, which comes second with 56 points at the same price per token as Sonnet 5. Careful: at max it produces more text per task than any model they have measured, so in the tests it ends up costing more than Opus 5.5.
- 29 September 2026. At its DevDay, OpenAI presents GPT-6.1 Sol, which lands one point behind GPT-6 Astra at a fifth of the price per token. Pro 500 also arrives, and new Pro 200 subscriptions have a lower allowance. Some subscribers keep the previous allowance through 29 October.
- 30 September 2026. Google presents Gemini 4 Argon, which ties on 53 with GPT-6 Astra and Fable 5.1. Cyber security experts try it first; after that it will reach paid API customers and Google AI Ultra.
There is a discrepancy in the source: GPT-6 Luna has 37 on the main leaderboard and 38 on its individual profile. I use 37 to keep one reference table; the difference does not prove the model improved. Compare the same index version and configuration. Announced next: Haiku 5.5, which Anthropic expects in the coming weeks, and wider access to Argon.
The smartest AI in the world, and when you don’t need it
The first line is the model with the highest score on this index. Use it for anything with money or a client at stake, like a lease before you sign it, and even then check every figure. To answer emails, summarise a meeting or add up a sales spreadsheet you don’t need it: any of the good ones will do, and many are free.
Also, the smartest models let you choose how much they think before answering, what they call effort. At max they do a bit better, but they take much longer and eat into your usage limit. For everyday work, medium or high is plenty.
At work, other things matter more than two points on a score:
- Price and limits. Start with the plan you already have. Upgrade when you have tested a need for a feature or more capacity; compare full cost, quality and net time, including review. Running an agent for hours does not guarantee an expensive plan is worth it.
- Your data. Before using clients’ personal data, check the tool, contract and permissions. Switching off training is not enough.
- The app you already use. If you live in Gmail and Drive, Gemini is already there; if you live in Word and Excel, Copilot.
Prices in Spain are in which AI subscription is worth it.
The one I’d use for everyday work
The second line is for normal work: a report, a long email to a client, the quarter’s sales spreadsheet. I look for the one that gives the most for what it costs, in money and in usage limits.
My advice: set it as your default model, on medium or high effort, and go up to max when something gets stuck. If you already pay for ChatGPT Plus or Google AI Pro, don’t switch for a few points; only move if your plan falls short.
The cheapest AI for high-volume tasks
The third line is for the things you do in bulk: sorting emails, pulling data out of invoices, answering the same customer questions. That rarely goes through the chat. You pay per use, what they call the API, from an automation or an agent, at so much per million tokens, the chunks of text they charge by.
And the price varies enormously: in Artificial Analysis’s tests, as of 1 October 2026, each task cost GPT-6 Luna about 7 cents and Opus 5.5 at max almost $6. Your email classification task may cost a different amount. Before you build anything, ask the AI to do the maths:
Every month I get about 2,000 customer emails, half a page each, and I want an AI to sort them into urgent, quotes and the rest. Work out roughly how many tokens that is and how much it would cost per month with a model that charges $[input price] per million input tokens and $[output price] per million output tokens. Show me the calculation.
And test it first on twenty or thirty fictional cases resembling your work.
The best open model and the Chinese models
The fourth line is the best open-weights model: you can download it and run it on your own servers, but check its licence; open weights do not always mean MIT or Apache. MiMo-V2.6-Pro does use MIT and has over a trillion parameters. Keeping data in-house also requires control over connections, tools and logs. Large models need powerful machines; you can try smaller ones on a laptop with LM Studio or Ollama.
For months now, the best open models have been Chinese. Before you use them:
- Be careful with their apps. On DeepSeek’s website or app, for example, what you type is processed on servers in China. In January 2025 Italy’s data protection authority blocked it (in Italian), and in Spain the consumer group OCU asked the Spanish Data Protection Agency to investigate (in Spanish). My rule: no client data in Chinese apps.
- The app is not the model. You can install the weights yourself or use another hosting provider. Check where it processes data, what it logs and what its contract permits.
- Check for bias and omissions. For history or politics, check answers against sources; for work, also test cases where it might leave out important information.
In its September 2026 report, Anthropic accused seven Chinese labs, including Xiaomi, Zhipu, Moonshot, Alibaba and DeepSeek, of harvesting Claude’s answers at scale to train their own models. For now it’s an accusation, but it’s worth knowing.
How to read AI rankings without being fooled
Artificial Analysis uses ten tests in index v4.3.2, mainly English text. Agents and long documents account for 45%; it also tests coding, science and knowledge. Do not assume a small lead transfers to your work: it estimates a 95% confidence interval below ±1 point from repeated runs on certain models, not a guarantee for every case. The tests also change: Fable 5.1 scored 66 at release and 53 on the later index; that does not prove a loss of capability.
Arena, called LMArena until January 2026, compares anonymous answers and collects votes. Preferring an answer does not guarantee it is correct. Since July 2026 it also offers a factuality adjustment on its text and search leaderboards. Check the view, filters, models and date you are comparing; do not mix a preference ranking with one that includes that adjustment.
Companies’ own charts, treat them as advertising: each one picks the tests that flatter it.
What counts most is trying them on your own work: in which AI to use for what I explain how to do it in twenty minutes.
History: who was top each month
| Month | Smartest on Artificial Analysis | Best open model |
|---|---|---|
| June 2026 | Claude Fable 5, from the 9th | GLM-5.2, by Zhipu, from the 16th |
| July 2026 | Claude Fable 5; from the 24th, Claude Opus 5 | GLM-5.2; at the end of the month, Kimi K3, by Moonshot |
| August 2026 | Claude Opus 5 | Kimi K3 and GLM-5.3 (retrospective reconstruction) |
| September 2026 | Claude Fable 5.1, later level with GPT-6 Astra; from the 22nd, Claude Opus 5.5 | GLM-5.3 and Kimi K3; by month end, MiMo-V2.6-Pro, by Xiaomi |
This history summarises documented milestones, not every day of each month: Fable 5 on 9 June, GLM-5.2 on the 16th, Opus 5 on 24 July and Fable 5.1 on 1 September. August is reconstructed from the September account, not a saved month-end snapshot. Xiaomi dates the MiMo release to 22 September; that does not establish its exact leaderboard entry date. I do not compare scores across different test versions.
If you don’t want to follow all this yourself, in how to keep up with AI I explain how I filter the news.
Review on 1 October 2026: scores, configurations, rates and the limits of the comparisons. Confirm the current offer before subscribing.
Sources, checked on 1 October 2026: Artificial Analysis model ranking, its coding agent comparison, its Sonnet 5.5 analysis, GPT-6.1 Sol on Artificial Analysis, its Opus 5.5 analysis, its methodology, Google’s Gemini 4 Argon announcement and its Artificial Analysis profile, OpenAI’s DevDay recap, API prices from Anthropic, OpenAI, Google, Meta, SpaceXAI, Xiaomi, Z.ai and Moonshot, Claude prices in Spain (in Spanish), ChatGPT Work and Codex models and the Arena leaderboard.


