Claude Model Comparison: opus 4.8 is Smart but it still takes a village
What content marketers need to know about Anthropic’s latest model Opus 4.8
by Lisa Peyton and her team of AI superheroes
Last week Anthropic released Opus 4.8 just weeks after releasing Opus 4.7. I can barely keep up these days – I was JUST starting to get to know Opus 4.7. I have been using Opus 4.8 all week and have noticed that I prefer working with this model on strategic planning and anything that requires strong reasoning and synthesis and analysis of disparate data and concepts. This model is incredibly intelligent and currently tops the intelligence leaderboards on Artificial Analysis.
But is this the best model for copywriting and creative writing? In order to determine that, I ran my newsletter prompt through four Claude models and did a head-to-head comparison. The TL:DR version of this post is that there’s not one model pulling ahead but instead each has it’s own strengths. We’re still in a period of testing and learning, but below I share my findings with a comparison chart for quick reference.
A quick word on the effort slider
One thing changed with 4.8 that’s easy to miss and worth knowing, because I leaned on it all through this test. There’s now an effort control sitting right next to the model picker in the Claude app and in Cowork, on every plan, free included. In the API it’s a setting called effort. Think of it as a dial for how hard you want the model to think before it answers.
The dial runs from low to max and sits at high by default. Below high, the model replies faster and lighter. Above high sit two settings, usually shown as extra and max, where it slows down and reasons more deeply before it writes a word.
The trade is straightforward. More effort means more thinking, which means stronger results on genuinely hard problems, but also more time and more tokens, so it costs more or eats through your plan’s limits faster. Less effort means quick and cheap, with less depth.
For content work, here’s how I use it. High, the default, handles most drafting. I bump to extra for the pieces that matter, the voice-sensitive or reasoning-heavy ones, and that’s the setting that produced the strongest sections in this test. Max is usually overkill for writing. It burns a lot of tokens for very little extra polish on prose, so I keep it for heavy analysis. For quick or high-volume jobs, low or medium does the work without the wait.
So when you see me mention running a model “at higher effort” below, that’s this dial, turned up a notch from the default.
How I tested it
My weekly newsletter is a good stress test, because it asks a model to do a little of everything in one pass. Summarize a dozen articles accurately, capture my style and tone, pull out what matters to my readers, and format it cleanly. So I ran the same prompt and the same sources through four Claude models, Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6, and ran the two newer ones (Opus 4.8 and 4.7) at the higher effort setting ‘extra’, which tends to help on content work.
What I’m actually grading
When I look at AI-written copy, I’m weighing four things, in this order. Accuracy first, because reader trust is on the line and ESSENTIAL in the AI era. Then voice, so it captures my style and tone and is engaging. Then strategic judgment, whether it surfaces what my audience needs. Then formatting. Here’s how the four did.
Accuracy
Good news first. On the facts I could check, all four held up nicely. I verified about a dozen specifics against primary sources, dates, stats, quotes, and they stood up. That’s genuinely reassuring.
The most interesting thing I learned came from a single slip. Surprisingly the most intelligent model, Opus 4.8, writing a summary for a Forbes article it couldn’t actually open, filled the gap by assuming I’d written it. I hadn’t. It was upfront about the part it knew it was guessing. That’s worth knowing about even the smartest models. They’ll flag what they know they don’t know but can still include incorrect data. None of the other, ‘less intelligent’ models included the incorrect information and instead just let me know that it couldn’t fetch the article and it needed a manual summary.
So I did the sensible thing and ran the drafts through a second AI to fact-check, then a third to check that. What came back was fascinating. The three didn’t agree. One had invented the byline. One read the blocked article correctly and caught the mistake. One couldn’t reach the source and doubted the model that got it right. Three capable tools, three different answers, and not one of them could tell me which was true. I could, because I knew what I’d written and could open the page they were debating. That’s not a knock on any of them. It’s a clear picture of where you sit in this workflow. You’re not down in the weeds checking the AI’s arithmetic. You’re the one who decides what’s true when the tools disagree.
So if you’re wondering whether a second AI makes fact-checking bulletproof, it helps, and it’s worth doing for checkable facts. Just keep the final call on anything about you, and anything from a source the model couldn’t open, for yourself.
Voice and tone
This is where my real question lived. Is the newest, smartest model also the best writer? Not quite, and that’s the fun part. The warmest, most on-brand writing came from Opus 4.6, an older model, which also wrote the headline I liked best. Sonnet 4.6 wrote fluently and quickly, and for my particular voice rules it needed the most editing, mostly because it loves the constructions I tend to cut. None of the four landed publish-ready in my voice, which is normal. Every draft gets a pass from me no matter which model wrote it. The freeing part: the model on top of the intelligence charts isn’t automatically your best copywriter, so you get to choose by the job instead of the ranking.
Analysis and strategic thinking
This is 4.8’s home turf, and you can feel it. It’s the model I reach for now on strategy and on making sense of messy, unrelated inputs, and the leaderboards agree. The happy surprise in this test: the sharpest strategic synthesis, the section that ties the week’s stories into one argument, came together best under Opus 4.7 at higher effort. Even the intelligence champ doesn’t win every single job. That’s the whole point of this post – it takes a village of models.
Formatting
Happily, the least to worry about. All four kept formatting clean, links intact, headers tight, structure on template. A few tiny slips, nothing a quick read won’t catch.
The quick-reference chart
| What you care about | What I found across the four | Strongest for this in my test | The part that stays yours |
|---|---|---|---|
| Accuracy | All four held up on checkable facts. The only real miss was one model assuming I’d authored a source it couldn’t read | Close across the board; the older models were safest on what they couldn’t verify | Confirm anything about you, and any source the model couldn’t open. When the tools disagree, you decide |
| Voice and tone | Warmest, most on-brand writing from Opus 4.6; Sonnet 4.6 was fluent but needed the most editing; all needed a voice pass | Opus 4.6 | Your voice pass, every time |
| Analysis and strategy | 4.8 shines here and tops the intelligence index; this newsletter’s sharpest synthesis came together under 4.7 at higher effort | Opus 4.8 generally, Opus 4.7 for this synthesis | Set the angle going in, judge the takeaway coming out |
| Formatting | Clean across all four, only tiny slips | A wash, all solid | A quick scan of links and headers |
How the issue actually came together
| Section of my newsletter | The version I picked |
|---|---|
| Hot Takes and tool writeups | Opus 4.6 |
| Title and the Number | Opus 4.8, at higher effort |
| My Take, the strategic argument | Opus 4.7, at higher effort |
| Dependable all-rounder | Sonnet 4.6 |
I didn’t publish from one model. My final issue borrowed its best parts from across them, which is the most honest picture of where we are.
What it costs to run them
Cost belongs in this decision, and most comparisons skip it. Here are the current rates, billed per million tokens of input and output.
| Model | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Opus 4.8 | $5 | $25 |
| Opus 4.7 | $5 | $25 |
| Opus 4.6 | $5 | $25 |
| Sonnet 4.6 | $3 | $15 |
Two things stand out. The three Opus models cost exactly the same to run, so moving up to 4.8 doesn’t cost you a cent more per token than 4.6 or 4.7 did. And Sonnet 4.6 is the value pick, roughly 40% cheaper per token than any Opus model.
The sticker price isn’t your real bill, though. What you actually pay tracks how many tokens the model burns, and that comes down to two things: how wordy the model is, and the effort level you pick. A max-effort Opus run does a lot more thinking, which means a lot more tokens, which means a much bigger bill than the same task at the default. The newer Opus models also tend to produce more tokens per request than 4.6, so two runs at the same rate can still land at different costs.
If you work inside the Claude app on a subscription instead of the API, you won’t see a per-token charge at all. There the cost shows up as how fast you burn through your plan’s usage limits, and Opus at high effort burns them quickest. (Opus 4.8 also offers a Fast mode at $10 and $50 per million for roughly 2.5 times the speed. Useful when latency genuinely matters, which for most content work it doesn’t.)
This gives the village idea a budget angle. You don’t need to run everything through the priciest setup. Draft on Sonnet 4.6, or at a lower effort, when that’s good enough, and save high-effort Opus for the work where the reasoning or the accuracy actually earns the spend. Matching the model to the job protects your budget, not just your quality.
So which model should you use?
If you take one thing from this, take this: there isn’t a single winner right now, and you don’t need there to be. Opus 4.8 is brilliant, and I’d reach for it on reasoning and strategy in a heartbeat. For warm, on-brand copy, an older model often serves better. The smart play is to keep a couple of models in rotation and match them to the task. We’re early, the tools shift weekly, and testing on your own work beats taking any ranking on faith, including mine.
The bigger picture
A new model lands every few weeks now, and it’s easy to feel behind. You’re not. You don’t have to adopt every release or crown a single favorite. You just need a simple way to decide what earns a place in your work, and lately the answer is rarely one model. It takes a village of models to make the world go round. The job quietly becoming the most valuable one isn’t picking the smartest tool. It’s knowing which to reach for, and catching what they miss. AI brings the knowledge. We bring the wisdom.
Lisa Peyton is an AI practitioner, professor, and pioneer helping marketers navigate the evolving AI landscape. Find more resources at lisapeyton.com/ai-marketing-resources or connect with her at linktr.ee/lisapeyton.
