WORKFLOWS
The AI Model Cheat Sheet 🚀
The 8 AI models worth knowing right now, what each is genuinely best at, the honest weakness of each, and what it costs. Plus the switch-between-models.
Everyone keeps asking me the same thing. Which AI model is the best one? I get why people want a single answer. But I think it is the wrong question to ask.
The better question is not which model is best. It is which model is best for this task. The people getting the most out of AI stopped picking a favorite. They switch between models based on the job in front of them, and once I saw the pattern I could not unsee it.
So here is the actual cheat sheet. The eight models worth knowing right now, what each one is genuinely good at, and the honest weakness of each. Save this one.
The 8 Models, At A Glance
Scroll sideways on your phone to see the whole table. Cost is per million tokens, shown as input and output, which is how these are actually billed.
| Model | Reach for it when | Real weakness | Cost in / out |
|---|---|---|---|
| Claude Fable 5 | The hardest problem you have. Best coding scores of anything you can use right now. | Most expensive on this list. Overkill for everyday work. | $10 / $50 |
| Claude Opus 4.8 | Long jobs that run for hours. Big migrations, deep research, agents that hold context. Great natural writing. | Pricey, and slower than you want for quick questions. | $5 / $25 |
| Claude Sonnet 5 | Your daily driver. Plans, uses a browser and terminal, runs on its own. | Uses more tokens per task, so at max effort it can cost more than Opus. | $2 / $10 |
| GPT-5.6 Sol | Hard multi-step problems and coding that spans files, tests, and follow-up fixes. | Priciest of the GPT tiers, and you rarely need this ceiling. | $5 / $30 |
| GPT-5.6 Terra | The broad middle of real work. Support replies, document analysis, drafting at scale. | Not the top ceiling. Hard reasoning goes to Sol. | $2.50 / $15 |
| GPT-5.6 Luna | Summarizing, labeling, extracting, first drafts. One fifth the cost of Sol. | Long-document recall falls off a cliff. Do not hand it a huge file. | $1 / $6 |
| Gemini 3.5 Flash | Speed, and anything needing live Google Search grounding. | Gemini writing can read a little flat next to Claude. | $1.50 / $9 |
| Grok 4.3 | Breaking news and live X, plus video you want it to actually watch. | Trails the others on coding, and takes about 12 seconds to start answering. | $1.25 / $2.50 |
Prices checked July 2026. Sonnet 5's intro pricing rises about 50% on September 1, 2026.
Everyday work: Claude Sonnet 5 or GPT-5.6 Terra.
Cheap and high volume: GPT-5.6 Luna or Gemini 3.5 Flash.
What happened in the last hour: Grok 4.3.
Look Up Your Job, Get The Model
This is the one to screenshot. Find what you are actually doing, use the first pick, and fall back to the second if you do not have access to it.
| What you are doing | Use this | Or this | Why |
|---|---|---|---|
| Writing people will read blogs, newsletters, emails in your voice | Claude Opus 4.8 | Claude Sonnet 5 | Most natural prose, and it holds your tone over a long piece. |
| Short punchy copy hooks, subject lines, ads, captions | GPT-5.6 Terra | GPT-5.6 Sol | GPT is stronger at short business copy than long-form voice. |
| Hardest coding problem | Claude Fable 5 | Claude Opus 4.8 | Top coding score of any model you can use today. |
| Everyday coding and code review | Claude Sonnet 5 | GPT-5.6 Terra | Handles most real work without the flagship price. |
| Agents and terminal work | GPT-5.6 Sol | Claude Sonnet 5 | Sol genuinely leads on agentic and terminal operation. |
| A job that runs for hours migrations, deep research | Claude Opus 4.8 | Claude Fable 5 | Built to hold context across long autonomous work. |
| Work you will not double-check | Claude Opus 4.8 | Grok 4.3 | Both are unusually willing to say they are unsure instead of inventing an answer. |
| Research that must be current | Gemini 3.5 Flash | Grok 4.3 | Most reliable live Google Search grounding. |
| Breaking news, live sentiment | Grok 4.3 | Gemini 3.5 Flash | Reads X live rather than a stale index. |
| Big documents and spreadsheets | Gemini 3.5 Flash | Claude Opus 4.8 | Chews through large files fast without slowing down. |
| Video you want understood | Grok 4.3 | Gemini 3.5 Flash | Takes video directly, up to five minutes, and follows what happens. |
| Bulk simple tasks summaries, labeling, extracting | GPT-5.6 Luna | Gemini 3.5 Flash | Cheapest way to run volume. Just never on long documents. |
Pros And Cons Of Each
The table gets you moving. This part is the why, so you can make the call yourself when a new model shows up next month.
Claude Fable 5
Anthropic · the ceiling · $10 / $50
Good At
- The hardest coding job you have. It holds the top coding score of any model you can currently use.
- Problems that beat everything else. When another model has already stalled twice, this is the escalation.
- Planning and architecture where getting the shape wrong costs you weeks.
Not Good At
- Everyday questions. It costs double Opus 4.8, which is already the expensive one.
- High volume work. Running your bulk tasks through this will wreck your bill fast.
- Anything you need back in seconds. It thinks before it answers.
Buy it if: you hit a wall the other models cannot get past, and being right matters more than the cost.
Claude Opus 4.8
Anthropic · the long-haul worker · $5 / $25
Good At
- Jobs that run for hours without you. Big migrations, deep research, agents that hold context across a long task.
- Writing people will actually read. The most natural long-form prose of anything here, and the best at holding your voice.
- Being honest. It is the strongest of these at saying it is unsure instead of confidently making something up, which matters enormously when you cannot check the work.
- Production quality code where correctness beats speed.
Not Good At
- Speed. It is roughly half as fast as GPT-5.6 Sol on the same task.
- Cheap bulk output. It produces close to three times more output per task than Sol, so output-heavy work gets expensive.
- Quick back and forth. Use a smaller model for chatting.
Buy it if: you are writing something real, or running long jobs where a confident wrong answer would cost you.
Claude Sonnet 5
Anthropic · the daily default · $2 / $10
Good At
- Almost everything, at a fair price. This is the one to default to.
- Working on its own. The most capable Sonnet yet at planning a task and using a browser or terminal to finish it.
- Daily coding, code review, and pulling information out of documents.
Not Good At
- The absolute hardest problems. That is what Fable 5 and Opus are for.
- Maximum effort settings. It emits more tokens per task, so pushed all the way it can actually cost more than Opus 4.8.
Buy it if: you want one model for most of your week and you do not want to think about it.
GPT-5.6 Sol
OpenAI · the flagship · $5 / $30
Good At
- Speed on hard problems. It finishes complex tasks in roughly half the time Opus takes.
- Agentic and terminal work. This is where it genuinely leads.
- Coding that spans many files, tests, and follow-up fixes without losing the plot.
Not Good At
- Being trusted without checking. It hallucinates noticeably more than Claude Opus 4.8, and at medium effort it has been caught skipping its own verification step, producing wrong answers that still pass the tests.
- Cheap work. It is the priciest of the three GPT tiers.
- Natural long-form writing. Claude reads better.
Buy it if: you need hard problems solved fast and you are going to review the output anyway.
GPT-5.6 Terra
OpenAI · the middle · $2.50 / $15
Good At
- The broad middle of real work. Customer replies, reading documents, internal knowledge search, drafting at volume.
- Short business copy. Hooks, subject lines, ad copy, calls to action.
- Getting a lot done for the money. Roughly half the cost of the previous generation flagship.
Not Good At
- Genuinely hard reasoning. That belongs with Sol.
- Long-form voice work. It drifts off your tone over a long piece.
Buy it if: most of your work is business-as-usual and you want the best value per task.
GPT-5.6 Luna
OpenAI · the cheap one · $1 / $6
Good At
- Simple, repetitive, high-volume tasks. Summaries, labeling, tagging, pulling fields out.
- Rough first drafts you are going to rewrite anyway.
- Cost. About a fifth of what Sol costs.
Not Good At
- Long documents. This is the big one. Its ability to remember detail across a long file drops off sharply compared to Sol and Terra. Give it a big document and it will quietly miss things.
- Anything with real stakes. Do not put it on work you will not check.
- Nuanced writing. It is a workhorse, not a writer.
Buy it if: you have a pile of small boring tasks and you are paying too much to run them elsewhere.
Gemini 3.5 Flash
Google · the fast one · $1.50 / $9
Good At
- Speed. It beats Google's own bigger model on coding and agent work while producing output several times faster.
- Live search. The most reliable connection to current Google results of anything here, so it is the one for questions where being current matters.
- Large documents and spreadsheets. It chews through big files without slowing down.
- Google Workspace if your work already lives in Docs, Sheets, and Gmail.
Not Good At
- Writing with personality. Gemini output tends to read flatter than Claude.
- Being the top of the range. Gemini 3.5 Pro has been delayed repeatedly and is still not out, so Flash is your 3.5 option right now.
Buy it if: you want speed and current information, or your whole team already lives in Google.
Grok 4.3
xAI · the live one · $1.25 / $2.50
Good At
- What is happening right now. It reads X live instead of a stale index, so on breaking events it is genuinely ahead.
- Video. It takes video directly, up to five minutes, and can actually follow what happens in it.
- Admitting it does not know. It scores unusually well at saying so instead of inventing an answer.
- Cost. The cheapest output on this entire list.
Not Good At
- Coding. It trails the others here, and xAI has not published the usual coding benchmarks for it.
- Fast conversation. It takes about 12 seconds before it starts answering, which you feel immediately.
- Very long inputs. Anything over 200,000 tokens of input is billed at double rate.
Buy it if: you need live information, you work with video, or you want frontier-adjacent quality for a fraction of the price.
The Three-Model Workflow
The workflow I keep seeing has three steps, and you do not use one model for all of them. You hand each part of the job to the model that is best at that part. Think first, then execute, then review.
Start with your smartest model. This is the one you use to plan the strategy, break the problem into pieces, and map out the approach. You want deep reasoning here, not speed. Give it room to think before anything gets built. In my flow this is a strong reasoning model like Claude's Fable 5.
Now switch to your fastest model. The hard thinking is done, so this part is about doing the work quickly. Building, writing, coding, churning through the plan you just made. You do not need your deepest thinker for this, you need speed. A quick model like GPT-5.6 Sol does the heavy lifting fast.
Here is the step most people skip. Take the finished work and hand it to a different model to look for mistakes. Not the one that built it. A fresh model reads it cold and catches what the others missed. This single step has saved me more times than I can count.
Why A Different Model To Review
Every model has different strengths. Every model also has different blind spots. That is the whole point of switching.
When one model builds something, it tends to be blind to its own mistakes. It made the same assumptions on the way in and on the way out. Ask it to check its own work and it will often tell you it looks great, because it is reading the problem the same way it did the first time.
A second, different model does not carry those same blind spots. It reads the work cold, with no attachment to how it got made, so it notices the gaps the first one talked itself past. It is the same reason a friend catches the typo you read past ten times. Fresh eyes see what tired eyes cannot.
My Lineup, By Job
Here is how the eight above map onto real jobs. This is a starting point, not a rulebook.
- Plan and strategy: Claude Fable 5. The one I trust to think a problem all the way through before anything gets built.
- Build and code: GPT-5.6 Sol, or Claude Sonnet 5 if you want the cheaper lane. Once the plan is set I want it made, not another deep-thinking session.
- Review and critique: any third model that did not do the work. Its only job is to read the result cold and tell me what is broken.
- Writing anything a person will read: Claude Opus 4.8. It has the most natural voice of the eight.
- Research with live sources: Gemini 3.5 Flash for search grounding, Grok 4.3 when the thing happened in the last few hours.
- Bulk and boring: GPT-5.6 Luna or Gemini 3.5 Flash. Short tasks, high volume, low cost. Just do not hand Luna a long document.
Think Streaming Services
Here is how I explain it to people. AI models are like streaming services. Netflix, Hulu, all of them. If you only ever open one, you miss all the good stuff sitting on the others.
Nobody watches only Netflix and thinks they have seen everything worth watching. The good shows are spread across all of them. The skill is knowing which app to open for what you are in the mood for.
AI works the same way. The advantage is not owning the one perfect model. It is knowing which one to open for the job in front of you, and being able to switch in seconds.
In outbound marketing and career positioning, generic applications have near-zero conversion. Treat yourself as a high-ticket solution: identify the company's pressing operational pain points, build a mini-audit or work sample using AI before you apply, and bypass crowded channels by reaching out directly to the decision-maker with structured value.