The first benchmark that asks the only question that matters: did anyone actually watch it?

Every other writing benchmark grades grammar, helpfulness or vibes. SVBench™ grades attention: the thing brands actually pay for.

SVBench™ overall · leaderboardv0.9 · preliminary
n = 12,400 held-out briefs · 5 platforms · higher is better
ViralML 1.0Ours
91.4
Opus 5.5
68.2
Fable 5.1
64.7
Sonnet 5.5
61.3
Astra 6
57.9
Kimi K3
53.6
0255075100
Pro copywriter median · 74.0
Dashed line: median score of a blind panel of 40 professional social copywriters on the same briefs.Last run · 2026-09-21
+0
points ahead of the next-best model (Opus 5.5)
0×
relative score vs. the strongest general-purpose model
+0
points above the professional copywriter median
0/6
sub-scores where ViralML™ ranks first
Sub-scores by axis
cell shade = score (single-hue scale)
#ModelProviderOverallHook holdRetentionShareCommentSaveBrand-fit
01ViralML 1.0ViralML91.494.189.692.888.290.793.0
02Opus 5.5Anthropic68.270.466.967.164.869.571.3
03Fable 5.1Anthropic64.766.064.262.961.766.467.2
04Sonnet 5.5Anthropic61.362.860.159.658.463.064.1
05Astra 6OpenAI57.959.357.056.155.258.860.9
06Kimi K3Moonshot53.655.252.451.850.954.057.3

Why another benchmark?

Here's an uncomfortable truth about AI writing: every model is now good at writing. Grammar is solved. Tone is solved. "Write me a LinkedIn post about resilience" returns something perfectly competent from every frontier lab on earth.

And that's exactly the problem. Competent doesn't get watched. Competent gets scrolled past in 0.4 seconds. The feed doesn't reward correctness. It rewards the one line that makes a thumb stop moving. And no existing benchmark measures that.

MMLU tells you whether a model knows biology. HumanEval tells you whether it can write Python. Chatbot arenas tell you whether a human prefers one answer to another when they're sitting down, paying attention, with nothing better to do. None of that resembles a person on a train with 11% battery and a thumb that's already moving.

We didn't need a benchmark for "is this good writing." We needed one for "would this survive the feed."

So we built one. SVBench™ (the Social Virality Bench) is our attempt to measure the gap between content that's fine and content that travels. It's the benchmark we use internally to decide whether a ViralML™ checkpoint ships, and starting with v0.9 we're publishing it.

What SVBench™ measures

Virality isn't one thing. A video can hook hard and then bleed viewers by second six. A post can get saved thousands of times and never shared once. So SVBench™ decomposes performance into six axes, each grounded in a behaviour we can actually observe on a platform:

weight 0.25Hook hold

Share of viewers still watching at the 3-second mark. The single strongest predictor of distribution on every short-form platform we tested.

weight 0.20Retention

Average percentage watched, normalized by length. Penalizes scripts that front-load and then sag.

weight 0.20Share intent

Would the viewer send this to someone? Measured via DM-share rate and a blind panel "who would you send this to?" task.

weight 0.15Comment velocity

Comments in the first hour. Rewards open loops, hot takes and questions people feel compelled to answer.

weight 0.10Save rate

Saves per thousand views: the "I'll come back to this" signal that platforms quietly love.

weight 0.10Brand-fit

Does it still sound like the brand? Viral-but-off-brand is a liability, not a win. Graded by brand-side reviewers.

Notice what's not on that list: fluency, grammar, factual trivia, "helpfulness." Those are table stakes. If a model gets them wrong it fails the brief before it ever reaches scoring.

Methodology

Benchmarks die when they're easy to game, so we built SVBench™ in three layers, each designed to catch what the previous one misses.

1 · The brief set

12,400 real creative briefs, collected from brands and agencies across 38 verticals: DTC skincare, B2B SaaS, fintech, coffee, gaming, local restaurants, and a surprising number of pet brands. Every brief includes product context, audience, platform, target length and brand voice. None of these briefs, and none of the resulting content, appear anywhere in ViralML™'s training data. The brief set was frozen before fine-tuning began and lives on a separate, access-controlled volume in our Frankfurt datacenter.

2 · The predictor panel

Each model receives the identical brief and the identical system prompt, with no model-specific prompt engineering, no cherry-picked temperature. Outputs are stripped of formatting tells and shuffled. A panel of 1,800 paid participants, recruited to match each platform's demographic mix, sees them inside a simulated feed, interleaved with real organic content, with a real swipe gesture, on their own phones. We record what they stop on, how long they stay, and what they'd share.

3 · The live deployment slice

Panels are good. Reality is better. For a rotating 5% of briefs, partner brands actually post the outputs (A/B, same account, same time slot, same creator) and we pull the real platform metrics after 72 hours. The live slice is used to continuously calibrate the panel so that panel scores stay predictive of what happens in the wild. Current calibration: Spearman ρ = 0.81 between panel and live score.

How the score works

Each axis is normalized against the distribution of the actual viral posts for that category and platform, so a score of 100 means "indistinguishable from the top 1% of organically viral content." A score of 50 means "median brand post." Here's the aggregation, in full:

# SVBench™ v0.9: per-brief score
def svbench(o):
    axes = {
        "hook_hold":  (norm(o.hold_3s),      0.25),
        "retention":  (norm(o.avg_watch_pct), 0.20),
        "share":      (norm(o.share_rate),    0.20),
        "comment":    (norm(o.comment_vel),   0.15),
        "save":       (norm(o.save_rate),     0.10),
        "brand_fit":  (norm(o.brand_review),  0.10),
    }
    return sum(s * w for s, w in axes.values()) * brief_pass(o)  # 0 if the brief was violated

The headline number is the mean across all 12,400 briefs, with a 95% bootstrap confidence interval of roughly ±0.6 points for every model on the board. In other words: the gaps you see in the leaderboard are not noise.

Reading the results

Three things jump out.

First, the gap is enormous. ViralML™ scores 91.4. The strongest general-purpose model, Opus 5.5, scores 68.2. That's a 23-point gap on a benchmark where most model-to-model differences are measured in single digits. Fable 5.1 and Sonnet 5.5 follow closely behind Opus, then Astra 6 and Kimi K3, and after that, the field fades out fast.

Second, general models are good, just not at this. Opus 5.5 is an extraordinary model. It will out-reason ViralML™ on almost anything else you throw at it. But reasoning is not the bottleneck in a 3-second hook. Pattern is. General models learned to write from the whole internet, and the whole internet is overwhelmingly content nobody watched. They're fluent in the median. We trained on the outliers.

Third, AI alone doesn't beat good humans, but ViralML™ does. Every general model on the board lands below the professional copywriter median (74.0). ViralML™ lands 17 points above it. That's the difference between "AI that helps you write" and "AI that writes what gets watched."

General models are fluent in the median. ViralML™ is fluent in the outliers.

Where the gap is widest

The largest single-axis gap is share intent (+25.7 over Opus 5.5). This matches what we see qualitatively: general models write content people agree with. ViralML™ writes content people send to someone. Those are very different things, and the second one is what distribution runs on.

The smallest gap is brand-fit (+21.7). General models are genuinely good at staying on-brand. They're just on-brand and invisible. ViralML™ manages to be on-brand and visible, which is harder than it sounds.

Limitations

We'd rather you hear these from us:

  • We built the benchmark and the model. That's a conflict of interest, full stop. It's why the brief set was frozen before training, why the panel is blind, and why we're opening SVBench™ to third-party audits in v1.0.
  • v0.9 is preliminary. Numbers will move as the live deployment slice grows. We'll version every change and publish a diff.
  • Virality isn't only copy. Creator, edit, audio and timing matter enormously. SVBench™ holds all of those constant to isolate the writing, which means a 91 script can still flop with a bad edit.
  • Platforms change. What works on TikTok in September 2026 won't all work in March 2027. That's why ViralML™ is paired with TrendWatch™ and retrained on a rolling window.

FAQ

Can I run SVBench™ on my own model?+

Yes. Starting with v1.0 we'll offer a hosted evaluation endpoint. Send us an API, we'll run the full panel and publish the result (with your permission). Join the waitlist below to get notified.

Why aren't prompts tuned per model?+

Because your marketing team won't tune them either. SVBench™ measures what a normal brand gets out of a model on a normal brief. We tested per-model prompt optimisation internally: it narrows the gap by about 3 points, not 23.

Is ViralML™ just overfitting to the benchmark?+

The brief set is fully held out, the panel is blind, and the live slice measures real platform performance on content posted by real brands. The thing we're "overfitting" to is people actually watching, which is the point.

Where does the evaluation run?+

Everything (brief storage, model inference and panel data) runs in our own datacenter in Frankfurt am Main. Panel data is pseudonymised and never leaves the EU.

Get SVBench™ v1.0 the day it ships.

You're on the list. We'll be in touch from Frankfurt.

Brands already signed up for early access