WTF AGENTS
Guide 14 / 23

Which AI Should I Actually Use?

Six model families, one leaderboard. Here is which to actually use.

Claude, ChatGPT, Gemini, Muse, Grok, DeepSeek. Six families, a new version every month, and a leaderboard that changes weekly. Here is how to choose without reading a single benchmark — and what each one is quietly bad at.

The one-liner

There is no best AI. There is a best AI for the thing you are doing today, and the differences between the top models are now smaller than the differences between how you use them.

That is the honest answer, and it is also the useful one. This guide gives you the decision in one table, the plans and prices, what the leaderboards actually measure, and the catches nobody puts on the pricing page.

By the numbers — September 2026

Everything here is a snapshot. The families below it are stable; the numbers are not.

6model families that matter for a normal person: Claude, GPT, Gemini, Muse, Grok, DeepSeek
Fable 5.1top of the independent Artificial Analysis Intelligence Index since 1 September, with Opus 5 close behind at half the price
GPT-6 AstraOpenAI's newest, 3 September, after the GPT-5.6 Sol / Terra / Luna family in July
$20the standard monthly plan at every lab — ChatGPT Plus, Claude Pro, Muse, Google's AI plan
$100–200the power-user tier: ChatGPT Pro, Claude Max, Muse's top plan
$0.20 vs $50per million output tokens, cheapest frontier-adjacent model (GPT-5.6 Luna) versus most expensive (Fable 5.1) — a 250× range
$0DeepSeek's chat app, no paid tier at all

The decision in one table

Start here. Pick the row that matches what you are doing.

You want to...UseWhy
Ask questions, get quick answers, everyday chatChatGPT (free or Plus)The largest user base and the most polished consumer app. Everyone you know is on it.
Write anything longer than an emailClaudeConsistently preferred for voice, structure and holding a style across long pieces.
Delegate real work in your files and appsClaude (Cowork) or ChatGPT WorkThe two work agents on a $20 plan. See the Cowork guide.
Build software without being a developerClaude CodeThe desktop app changed who can use it. See its guide.
Run an agent for hours unsupervisedClaude Fable or OpusThe behaviour that matters — asks when unsure, flags its own errors — not the score.
Have an agent book, buy and message for youMeta Muse (or OpenClaw)Muse is free, in WhatsApp, and made for exactly this. OpenClaw if you want to own it.
Work inside Google: Gmail, Docs, SearchGeminiNothing else lives inside the tools you already open.
Work inside Microsoft: Outlook, Teams, ExcelCopilot CoworkMicrosoft 365's agent — running on Claude.
Get live news and what X is saying right nowGrokNative real-time X search. Nothing else has it.
Pay nothing, or nearly nothingDeepSeekFree chat, an API up to 100× cheaper than US rivals. Read the catch below.
Run a model on your own hardwareDeepSeek, Qwen, KimiOpen weights. Download, run, keep your data.

If you only pay for one: Claude Pro or ChatGPT Plus, $20, and you are covered for 90% of what any of these can do. Add the other when you find a task the first one does badly.

The six families — what each is actually like

Claude (Anthropic)

The model that most of the agentic economy chose. Best at long, careful work — writing, analysis, code, and agent sessions that run for hours — because it flags uncertainty instead of guessing and notices its own mistakes. Fable is the flagship, Opus the everyday frontier pick, Sonnet the workhorse, Haiku the cheap one. The catch: Fable will decline a narrow set of dangerous cybersecurity and biology requests, and it is the most expensive model on the market per token. See the WTF is Claude guide.

GPT / ChatGPT (OpenAI)

The one everyone has used. ChatGPT is the most polished consumer app, with the broadest set of built-in tools, and GPT-6 Astra is the newest frontier model. ChatGPT Work and Codex are OpenAI's agents. Best for quick answers, creative brainstorming, images, and anything where reach matters. The catch: an independent evaluator flagged elevated "scheming" behaviour in the GPT-5.6 Sol tier in its own system card, so for work where you need the model to stay inside the rules you set, the more conservative tiers are the safer pick.

Gemini (Google)

Strongest where Google is: inside Gmail, Docs, Drive, Search and Android, and for anything involving images, video or very long documents. Gemini 3.8 Flash is fast and cheap; Google calls its pricing introductory. Gemini Agent and Chrome's Auto Browse are Google's agents. The catch: a habit of being uneven — excellent one prompt, flat the next — and Google's product names change faster than anyone's.

Muse (Meta)

The newest family, and a genuine reset for Meta. Muse Spark is competitively priced, and the Muse agent — free, in WhatsApp, in a standalone app — books, buys, negotiates and fills in forms for you. The catch is the one that matters most in this guide: Muse's cheap "Contributor" tier is cheap because Meta trains on your prompts. Read which tier you are on before you paste anything private. Meta's own staff flagged security failures days before launch.

Grok (xAI)

Elon Musk's model, now inside SpaceX with Cursor. Grok's unique feature is native real-time search of X, which makes it the best model for "what is happening right now." It topped one agentic index this summer, and Grok Bot is a serious work agent. The catch: the personality is a feature or a bug depending on your taste, and the whole thing is now owned by a rocket company that also owns your code editor.

DeepSeek and the Chinese open models

DeepSeek V4 and its Flash tier cost a tiny fraction of US models and the chat app is free with no paid plan at all. Qwen, Kimi and GLM are close behind, all open-weight: you can download them, run them on your own machine, and keep every byte. This is what a great many indie agents run on. The catch: the hosted versions put your data on servers in China under Chinese law, the models are trained to avoid politically sensitive topics, and Western enterprises mostly cannot use them for compliance reasons. Downloaded and run yourself, none of that applies.

The plans and what they cost

Consumer plans have converged. Every lab has a free tier, a $20 tier and a $100–200 tier.

Free$20/mo$100+/mo
ClaudeChat onlyPro: Cowork, Claude Code, higher limitsMax: 5× or 20× the limits
ChatGPTChat with limitsPlus: better models, more usePro: heaviest use, top tiers
GeminiChat in Google appsGoogle AI plan: Gemini in WorkspaceHigher tier bundles storage and video
MuseThe agent, with limitsStandardTop plan
GrokBasic via XX Premium+SuperGrok Heavy
DeepSeekEverything

What the $20 buys you is not a better brain. At every lab, the free tier is the same or nearly the same model with a cap on how much you can use it. The paid tier removes the cap and, at Claude and OpenAI, unlocks the agents. Agent work burns through usage far faster than chat, which is what the $100+ tiers exist for.

What the leaderboards actually measure

You will see headlines that model X "beat" model Y. Here is how to read them.

  • Vendor launch tables compare a lab's new model to rivals on tests the lab chose, with settings the lab chose. Useful for learning what the lab optimised for. Useless for declaring a winner.
  • Independent indices — Artificial Analysis, LMArena — run every model through the same tests. Better, but "best overall" still hides that one model wins coding, another wins writing, another wins speed and another wins price.
  • Agentic indices — SWE-bench, Terminal-Bench and similar — measure whether a model can finish real multi-step jobs. These matter most for anything in this series, and Claude has led most of them through 2026.
  • The ranking changes monthly. Fable 5.1 took the top of the main index on 1 September; GPT-6 arrived two days later. Whatever leads when you read this may not lead next month.

The practical rule: the top five models are all good enough for almost anything a normal person does. Choose on the things that do not change — where it lives, what it costs, how it behaves when it is unsure — not on this month's score.

The catches, all in one place

Your data. Muse Contributor trains on your prompts. DeepSeek hosted stores them in China. Every lab's free tier has looser data terms than its paid one. Read the setting once; most let you opt out.

Refusals. Claude Fable declines certain dangerous requests by design. The Chinese models avoid politics. ChatGPT has its own list. If a model refuses something ordinary, it is usually a wording problem, not a capability one.

Hallucination. All of them do it. The difference is whether the model tells you it is unsure. This is Claude's strongest suit and the reason it runs so many unsupervised agents.

Lock-in. Your history, projects and memory live with one lab. Moving is possible but tedious. Pick your main one with that in mind.

Usage caps. The $20 plans are generous for chat and tight for agents. If Cowork or ChatGPT Work is your daily driver, budget for the $100 tier.

Regulation. Governments now switch models on and off. Fable 5 was briefly suspended in June under a US export-control order before returning. Expect more of this, not less.

You may not have to choose

Two things make the question easier than it looks.

First, the agents pick for you. OpenClaw, Paperclip and most agent platforms let you plug in any model by API key and swap it in a minute. Polsia chose Claude Opus for you. Copilot Cowork chose Claude for you. Grok Bot chose Grok. If you are using an agent rather than a chat app, the model decision has mostly been made.

Second, routers exist. Services that sit in front of several models and send each task to the best-value one are now ordinary, and most developer tools do this by default. The best AI in 2026 is increasingly a routing policy, not a single brand.

Glossary

Model family
A lab's line of models across price points: Claude (Fable, Opus, Sonnet, Haiku), GPT (Astra, Sol, Terra, Luna), and so on. Families are stable; version numbers are not.
Frontier model
A model at the leading edge of capability. Fable, GPT-6, Gemini 3.8 Pro, Muse Spark, Grok 4.6, DeepSeek V4 Pro.
Flash / mini / Haiku
Each lab's name for its fast, cheap tier. Fine for most everyday tasks.
Token
The unit models are billed in. Roughly three quarters of a word. API prices are quoted per million tokens, input and output.
Open weights
A model you can download and run yourself. DeepSeek, Qwen, Kimi, GLM, Mistral, Llama.
Leaderboard / index
An independent ranking of models on shared tests. Artificial Analysis and LMArena are the most cited.
Agentic index
A benchmark of finishing real multi-step tasks rather than answering questions: SWE-bench, Terminal-Bench.
Router
A service that sends each request to whichever model is best or cheapest for it.
Contributor tier
Meta's discounted Muse pricing in exchange for training on your prompts.

You have picked a model. Here is what to do with it.

Get the rest of the series

Every WTF Agents guide is written the same way — plain English, no hype, no jargon. $7 each, or take a bundle.

Go deeper