Why another benchmark?
Here's an uncomfortable truth about AI writing: every model is now good at writing. Grammar is solved. Tone is solved. "Write me a LinkedIn post about resilience" returns something perfectly competent from every frontier lab on earth.
And that's exactly the problem. Competent doesn't get watched. Competent gets scrolled past in 0.4 seconds. The feed doesn't reward correctness. It rewards the one line that makes a thumb stop moving. And no existing benchmark measures that.
MMLU tells you whether a model knows biology. HumanEval tells you whether it can write Python. Chatbot arenas tell you whether a human prefers one answer to another when they're sitting down, paying attention, with nothing better to do. None of that resembles a person on a train with 11% battery and a thumb that's already moving.
We didn't need a benchmark for "is this good writing." We needed one for "would this survive the feed."
So we built one. SVBench™ (the Social Virality Bench) is our attempt to measure the gap between content that's fine and content that travels. It's the benchmark we use internally to decide whether a ViralML™ checkpoint ships, and starting with v0.9 we're publishing it.
What SVBench™ measures
Virality isn't one thing. A video can hook hard and then bleed viewers by second six. A post can get saved thousands of times and never shared once. So SVBench™ decomposes performance into six axes, each grounded in a behaviour we can actually observe on a platform:
Share of viewers still watching at the 3-second mark. The single strongest predictor of distribution on every short-form platform we tested.
Average percentage watched, normalized by length. Penalizes scripts that front-load and then sag.
Would the viewer send this to someone? Measured via DM-share rate and a blind panel "who would you send this to?" task.
Comments in the first hour. Rewards open loops, hot takes and questions people feel compelled to answer.
Saves per thousand views: the "I'll come back to this" signal that platforms quietly love.
Does it still sound like the brand? Viral-but-off-brand is a liability, not a win. Graded by brand-side reviewers.
Notice what's not on that list: fluency, grammar, factual trivia, "helpfulness." Those are table stakes. If a model gets them wrong it fails the brief before it ever reaches scoring.
Methodology
Benchmarks die when they're easy to game, so we built SVBench™ in three layers, each designed to catch what the previous one misses.
1 · The brief set
12,400 real creative briefs, collected from brands and agencies across 38 verticals: DTC skincare, B2B SaaS, fintech, coffee, gaming, local restaurants, and a surprising number of pet brands. Every brief includes product context, audience, platform, target length and brand voice. None of these briefs, and none of the resulting content, appear anywhere in ViralML™'s training data. The brief set was frozen before fine-tuning began and lives on a separate, access-controlled volume in our Frankfurt datacenter.
2 · The predictor panel
Each model receives the identical brief and the identical system prompt, with no model-specific prompt engineering, no cherry-picked temperature. Outputs are stripped of formatting tells and shuffled. A panel of 1,800 paid participants, recruited to match each platform's demographic mix, sees them inside a simulated feed, interleaved with real organic content, with a real swipe gesture, on their own phones. We record what they stop on, how long they stay, and what they'd share.
3 · The live deployment slice
Panels are good. Reality is better. For a rotating 5% of briefs, partner brands actually post the outputs (A/B, same account, same time slot, same creator) and we pull the real platform metrics after 72 hours. The live slice is used to continuously calibrate the panel so that panel scores stay predictive of what happens in the wild. Current calibration: Spearman ρ = 0.81 between panel and live score.
How the score works
Each axis is normalized against the distribution of the actual viral posts for that category and platform, so a score of 100 means "indistinguishable from the top 1% of organically viral content." A score of 50 means "median brand post." Here's the aggregation, in full:
# SVBench™ v0.9: per-brief score def svbench(o): axes = { "hook_hold": (norm(o.hold_3s), 0.25), "retention": (norm(o.avg_watch_pct), 0.20), "share": (norm(o.share_rate), 0.20), "comment": (norm(o.comment_vel), 0.15), "save": (norm(o.save_rate), 0.10), "brand_fit": (norm(o.brand_review), 0.10), } return sum(s * w for s, w in axes.values()) * brief_pass(o) # 0 if the brief was violated
The headline number is the mean across all 12,400 briefs, with a 95% bootstrap confidence interval of roughly ±0.6 points for every model on the board. In other words: the gaps you see in the leaderboard are not noise.
Reading the results
Three things jump out.
First, the gap is enormous. ViralML™ scores 91.4. The strongest general-purpose model, Opus 5.5, scores 68.2. That's a 23-point gap on a benchmark where most model-to-model differences are measured in single digits. Fable 5.1 and Sonnet 5.5 follow closely behind Opus, then Astra 6 and Kimi K3, and after that, the field fades out fast.
Second, general models are good, just not at this. Opus 5.5 is an extraordinary model. It will out-reason ViralML™ on almost anything else you throw at it. But reasoning is not the bottleneck in a 3-second hook. Pattern is. General models learned to write from the whole internet, and the whole internet is overwhelmingly content nobody watched. They're fluent in the median. We trained on the outliers.
Third, AI alone doesn't beat good humans, but ViralML™ does. Every general model on the board lands below the professional copywriter median (74.0). ViralML™ lands 17 points above it. That's the difference between "AI that helps you write" and "AI that writes what gets watched."
General models are fluent in the median. ViralML™ is fluent in the outliers.
Where the gap is widest
The largest single-axis gap is share intent (+25.7 over Opus 5.5). This matches what we see qualitatively: general models write content people agree with. ViralML™ writes content people send to someone. Those are very different things, and the second one is what distribution runs on.
The smallest gap is brand-fit (+21.7). General models are genuinely good at staying on-brand. They're just on-brand and invisible. ViralML™ manages to be on-brand and visible, which is harder than it sounds.
Limitations
We'd rather you hear these from us:
- We built the benchmark and the model. That's a conflict of interest, full stop. It's why the brief set was frozen before training, why the panel is blind, and why we're opening SVBench™ to third-party audits in v1.0.
- v0.9 is preliminary. Numbers will move as the live deployment slice grows. We'll version every change and publish a diff.
- Virality isn't only copy. Creator, edit, audio and timing matter enormously. SVBench™ holds all of those constant to isolate the writing, which means a 91 script can still flop with a bad edit.
- Platforms change. What works on TikTok in September 2026 won't all work in March 2027. That's why ViralML™ is paired with TrendWatch™ and retrained on a rolling window.
FAQ
Can I run SVBench™ on my own model?+
Yes. Starting with v1.0 we'll offer a hosted evaluation endpoint. Send us an API, we'll run the full panel and publish the result (with your permission). Join the waitlist below to get notified.
Why aren't prompts tuned per model?+
Because your marketing team won't tune them either. SVBench™ measures what a normal brand gets out of a model on a normal brief. We tested per-model prompt optimisation internally: it narrows the gap by about 3 points, not 23.
Is ViralML™ just overfitting to the benchmark?+
The brief set is fully held out, the panel is blind, and the live slice measures real platform performance on content posted by real brands. The thing we're "overfitting" to is people actually watching, which is the point.
Where does the evaluation run?+
Everything (brief storage, model inference and panel data) runs in our own datacenter in Frankfurt am Main. Panel data is pseudonymised and never leaves the EU.