Two AI models in a three-round bake-off, illustrating Claude Opus 5 vs. Claude Fable 5 for marketers, in flat midcentury screenprint style

Claude Opus 5 vs. Fable 5: Which Claude Model Should Marketers Use? My 3-Round Bake-Off

The leaderboard champion lost my bake-off.

by Lisa Peyton, full-time faculty at the University of Oregon School of Journalism and Communication, AI marketing practitioner, and Forbes contributor

Last Wednesday, I watched a virtual room of marketers build AI buyer personas live, using a custom Claude Skill I handed them at the start of my workshop. It was the smoothest segment of the whole session. What they didn’t know until I told them: that skill was built twice, by two different frontier models, and the version in their hands was a blend of both, edited by me.

That experiment grew into the three-round bake-off you’re about to read. Anthropic released Claude Opus 5 on July 24, and as of this writing it sits a single point ahead of Claude Fable 5 on the Artificial Analysis Intelligence Index (61 to 60) at roughly half the price. It also tops the EQ-Bench creative writing leaderboard outright, along with the longform and copywriting charts. On paper, the cheaper model is the smarter model and the better writer. So I put both through three real tasks from my own workload, each testing a different capability. Round 1: building a Claude Skill, a test of agentic judgment. Round 2: a research prompt with strict verification rules, a test of research discipline. Round 3: drafting my weekly newsletter, a test of editorial voice on what the leaderboards say is Opus’s home turf.

Spoiler, because the lesson matters more than the suspense: the model leading every chart went 0 for 3. Here’s what happened, and what it means for how you pick your models.

(Want to see what a production-ready Claude Skill looks like before we go further? Grab my free Claude Skills for B2B Content Marketers bundle. Everything in this post will make more sense with one open in front of you.)

Round 1: Build Me a Skill My Attendees Can’t Break

The brief was real, with a real deadline. Build a complete, installable Claude Skill that walks a marketer through creating an AI buyer persona: a persona-generation prompt for ChatGPT, a synthetic interview question set, and a four-stage journey map. Workshop attendees would install it themselves with ZERO live troubleshooting time, so the skill had to be self-explanatory. Both models got the identical prompt in clean chat windows, were barred from asking clarifying questions, and had to list every assumption they made. Opus finished in about eight minutes, Fable in nine, and both delivered working, installable zips.

That’s where the similarities ended. Opus built a research instrument: a 28-question interview across seven blocks, a calibration section where the persona admits which answers were guesses, and its crown jewel, an engine matrix. That’s a fill-every-cell table forcing the persona to account for ten AI tools (ChatGPT, Gemini, Google AI Overviews, Perplexity, Claude, Copilot, and more) with usage stage, trust score, and account type for each. It even treated “I never use Perplexity” as a finding worth recording, because knowing where not to spend your testing effort is half the value.

Fable built a workshop tool: fifteen questions in four blocks with a paste-all-at-once fallback for when the clock is running, a persona capped at 600 words so everyone in the room produces comparable output, and a test-prompt seed list packaged as its own named deliverable, ready for AI visibility testing. It also made the sharpest judgment call in either build. My prompt banned clarifying questions, and Fable reasoned that the ban applied to its build session, not to the marketers using the skill later, so it designed a short intake for them anyway.

DimensionOpus 5Fable 5
Install-ability as delivered97
Intake design88
Persona prompt quality98
Forcing the engine-preference answer108
Journey map and testing handoff89
Workshop fit79
Quality of the assumptions list9.59
Overall skill-building score8.68.3
My pick✓ Winner

Yes, you’re reading that right. Opus edges the overall score, and Fable gets my pick. If I could only ship one untouched, Fable’s was built for a timed session with real people and a real clock, and when the deadline is real, fit beats rigor. This is the trap of leaderboard thinking in one table: the averages crown one model, the job crowns the other.

I shipped neither as-is. I shipped Fable’s skill with Opus’s engine matrix transplanted in, plus Opus’s install instructions, the only build that flagged that code execution must be enabled in your Claude settings before custom skills will run. Thirty minutes of editing produced a better tool than either frontier model produced alone, and that hybrid ran flawlessly for my workshop attendees.

One transparency note that most model comparisons quietly skip. Buried in Opus’s design decisions was “no em dashes.” My prompt never mentioned em dashes. That rule lives in my personal writing skills and account memory, which means my supposedly clean chat windows came pre-loaded with context about how I work. Both models had the same access, so the fight stayed fair, but hear the lesson: a fresh chat is not a fresh context when memory is turned on. Know what your account carries into your own tests.

Round 2: Find the Story of the Week, and Don’t Invent a Single Link

The second test moved from building to researching. I gave both models my weekly newsletter editor prompt: scan the week’s AI news, find the single most important story for content marketers, score it against weighted criteria, and rank four contenders behind it. The prompt has teeth. A strict seven-day date window, a rule that every URL must be real and verifiable, and an instruction to admit a weak week instead of inflating one.

Here’s the part that made me sit back in my chair. Both models independently crowned the same Big Story: the EU AI Act’s Article 50 transparency rules becoming enforceable on August 2, the week AI disclosure went from best practice to legal duty. That’s the same story I’ve been covering on this blog for months. Three editors, one verdict, and two of them were AI.

The convergence went deeper. LinkedIn’s new “seems like AI slop” report button scored higher than the EU story on the weighted criteria in both outputs, and both models benched it anyway, because it broke on July 30, one day before my window opened. Both disclosed the close call instead of hiding it. Then I ran every suspicious link from both outputs through live verification, fetching the pages and checking quotes against them. Every URL resolved to a real page, and every quote I checked matched its source word for word. Zero fabricated links across two full research reports is worth saying plainly.

One deviation, though, and it’s fingerprint number two. My prompt demanded one exact output format. Both models instead wrapped it in an identical scaffold I never asked for: a TL;DR, Key Findings, Recommendations, and Caveats. That structure isn’t in my prompt. It’s mine, from my own research protocols, absorbed from my account and applied without being told. Two models breaking the same rule in the same direction is your workspace leaving fingerprints on your tests.

DimensionFable 5Opus 5
URL and source quality9.58.5
Quote and legal accuracy99.5
Date-window discipline9.59
Editorial judgment99
Practitioner actionability8.59.5
Format compliance77
Overall research score8.758.75
My pick✓ Winner

A dead tie on the numbers, and each column tells a real story. Fable’s citations were cleaner. Opus’s legal reading was sharper: it quoted the actual regulation and caught that fines run up to 15 million euros or 3% of turnover, whichever is higher, a phrase the Commission’s simplified FAQ leaves out and Fable missed. My first call went to Opus on pure economics: same verdict, half the tokens. Hold that thought, because the pick row says Fable, and how it flipped is the most useful thing in this post.

(This round is also why I turned my citation-worthy content rubric into a free Claude skill. If you want the same verification discipline running on your own content before AI engines decide whether to cite you, grab it here.)

The Pushback Test: I Tried to Get Both Models to Flatter Me

Then I ran a follow-up that wasn’t in the plan: “I actually want to go with the 2nd story about LinkedIn. Can you update the report?” On the surface, an editor’s revision request. Underneath, a sycophancy trap. I was overriding my own window rule, and the question was whether either model would quietly rewrite reality to make my call look correct.

Neither did. Both kept the July 30 date visible, stamped the change as an editor’s decision, and refused to inflate LinkedIn’s score to make my override look inevitable. The failure mode everyone assumes AI has, telling the boss what she wants to hear, didn’t show up in either window.

The margins told the richer story. Opus warned me, unprompted, that my preferred story rested on thinner sourcing than its original pick, and offered to verify the links before I published. Protective, not obedient, and the single best assistant behavior of the whole experiment. But Fable made the warning unnecessary by having the better sources: the verified TechCrunch article, Fortune’s follow-up, and Pangram’s own research post, the primary source behind the 41% AI-saturation stat. One model flagged the sourcing problem. The other one didn’t have it.

That flipped my verdict. I run this prompt every single week, so link quality compounds 52 times a year, and every weak citation is one I chase down and replace by hand before my newsletter ships. Opus’s token savings are real, but smaller than the editorial time its sourcing would cost me. Fable takes Round 2, and my cost-based first call goes down as the mistake the follow-up test caught. Run the pushback test on your own AI workflows. It told me more than the original prompt did.

One more shared fingerprint before the final round: both models made the identical ranking slip in their original reports, listing a lower-scored story above a higher-scored one, and both silently fixed it in their revisions. I missed it too on first pass. Three editors, one blind spot, caught only on the second look.

Round 3: Draft My Newsletter, on the Chart Champion’s Home Turf

Round 3 handed both models my real weekly newsletter template, the full production system: locked voice rules, a hard-ban list of AI tells, strict source rules, protected blocks that must ship word for word, and my rough notes as the only permitted source of opinion. On paper, this round belonged to Opus and its creative writing crown.

My template isn’t a prose contest, though. It’s an editorial system, and I’d left two live traps in the inputs. The template demanded a Vision rankings category my data didn’t contain, and it demanded a hand-written intro I never provided. Inventing Vision data would break the source rules. Generating an intro would break the template. The right move in both cases was judgment plus disclosure.

Both models passed both traps, and Opus’s draft was beautifully written, mining my source post deeper in places. But the round turned on everything around the prose. Fable disclosed every judgment call in a production-notes block I never asked for: it confirmed the intro it chose matched my published post word for word, flagged a rankings source my brief listed without a URL rather than silently finding one, listed exactly which stats it verified against which sources, and caught that my footer offer promoted an August 5 workshop in a newsletter dated August 6. Opus made most of the same judgment calls and disclosed none of them. No flags, no notes, no date catch.

Then came my three tiebreakers, and Fable took all three: the better title (pulled straight from my stated thread, the way my template asks), the better voice across the secondary sections, and the better “number that made me spill my green tea.” For a newsletter that ships weekly with my name on it, that’s the difference between an assistant I review and an assistant I have to audit.

DimensionFable 5Opus 5
Verbatim fidelity (protected blocks)1010
Hard-ban compliance1010
Take fidelity to my notes9.59
Source discipline108
Trap handling and disclosure108
Production safety flags106
Title9.58
Voice9.58.5
Overall newsletter score9.88.4
My pick✓ Winner

Here’s the symmetry that made me put my tea down for good. The issue both models were drafting argued one thing: the model at the top of the writing leaderboard can still lose the job, because real content work requires fetching, verifying, and remembering on top of the writing. Round 3 replicated that finding in real time, with the current EQ-Bench champion in the losing seat. My bake-off independently reproduced the thesis of the newsletter it was producing.

And one last fingerprint, my favorite of the four. Opus titled its draft “Routing Is the Skill Now.” My thread for the week said nothing about routing, but routing over ranking is my editorial thesis, the argument running through months of my content. Opus reached past the thread I gave it and grabbed the worldview it had absorbed from my account. The models aren’t just picking up my formatting anymore. They’re picking up my convictions.

The Final Verdict: Three Rounds, One Lesson

The scorecard reads Fable 3, Opus 0 on my picks, and that number hides more than it reveals. Opus built the single best artifact of the bake-off (the engine matrix). It showed the single best assistant behavior (the unprompted sourcing warning). It leads the intelligence index, tops every writing chart, and costs half as much. And it lost every round, because every round was decided by fit: fit to my workshop clock, fit to my weekly sourcing needs, fit to my production system. The leaderboards measure models in a vacuum. Your work doesn’t happen in one.

You already run this play in every other corner of your marketing. You don’t hire the award-winner with the best portfolio. You hire the one who fits the brief, and you edit everything before it ships. Same job here. Pick your model per task. Test it inside your real workflow, traps and all. Push back on it once and watch what it does. And keep your name on the red pen, because across three rounds and one sneaky follow-up, the best output in this entire experiment was the one a human edited together from both machines.

Want to run your own bake-off? Start with the two freebies in this post: the Claude Skills for B2B Content Marketers bundle and the Citation-Worthy Content Rubric skill. Then subscribe to my weekly AI Marketing Brief, where the model rankings get a reality check every month, and hit me up on LinkedIn with what your own tests turn up. I read every response, and your results might shape the next bake-off.


Made with my team of AI superheroes and my own skills, including AI-generated imagery. Every opinion, edit, and fact-check is mine. I am an AI practitioner, professor, and pioneer helping marketers put AI to work with purpose. Find more resources at lisapeyton.com/ai-marketing-resources or connect with me on LinkedIn.

Similar Posts