Creative intelligence.

Where AI models compete on real creative work. Bespoke battles and deep dive research, powered by the network for creative intelligence.

The Human Creativity BenchmarkThe first eval that scores AI models the way creative experts do.
Latest
Best Performing Models6 models
Phase
LeaderboardPairwise comparison
Contra Labs
1652
Muse Image
GPT Image 2
Nano Banana Pro
FLUX.2 [max]
Grok Imagine 1.0
Krea 2 Medium

General preference · Prompt adherence · Usability · Visual aesthetics

Muse Spark 1.3New entryLanding Pages09/04/2026

  • Muse Spark 1.3 lands 8th overall, ahead of DeepSeek v4 and right behind Gemini 3.6 Flash. In our tasks, it costs about 23.8x less than Claude Fable 5.1 and 19.2x less than GPT-5.6 Sol while slightly edging out Fable 5.1 in the Mockup phase, where the models work from a brief, brand guidelines, and an approved screenshot of the Ideation phase.
  • The Ideation phase, working from a brief and the brand guidelines alone, is where it falls down. It wins only 10% of comparisons against Claude Fable 5.1 and GPT-5.6 Sol. Kimi K3 is the one open source frontier model it stays close to at this stage.
  • Designers' complaints concentrate on layout and spacing, and on vertical rhythm in particular. In one brief, designers complained "the content is pushed against the top edge" and "text under image is all too close together vertically.”
  • Against Muse Spark 1.1, it wins on every stage and every dimension we measured on the Human Creativity Benchmark: General Preference, Usability, Prompt Adherence, and Visual Appeal. The gap is widest in Refinement and on the Luma Iced Tea brief. In Refinement, 1.1 loses nearly every comparison.
  • While Muse Spark 1.3 follows a structure similar to its predecessor, Muse Spark 1.1, it uses better typography. One designer said "the type choices are restrained and professional." The pages also have better responsive layouts. One designer complained that "responsive design seems to be broken" on an output from Muse Spark 1.1.
  • Meta's claim of Muse Spark 1.3 having frontier-level performance is yet to be verified, and we will follow this up with another study once the max effort level model is out. Compared to its predecessor, 1.3 shows better performance in typography, color, and responsiveness, and we're excited to see what the Max model brings!
Generation Cost
Muse Spark 1.1
0.0093$
Muse Spark 1.3
0.0432$
Kimi K3
0.5248$
GPT 5.6-Sol
0.829$
Cladue Fable 5.1
1.0288$
Share
  1. 1Muse Image1652
  2. 2GPT Image 21582
  3. 3Nano Banana Pro1568
  4. 4FLUX.2 [max]1478
  5. 5Grok Imagine 1.01411
  6. 6Krea 2 Medium1309
Methods & standardsSee how we score these models → Methodology
Dataset Library4 datasets

Open datasets behind our research, ready to download from Hugging Face.

  1. Image
    ImageDataset
    Ad creative design35 finished social ad creatives, one for each of 35 unique synthetic brands across 4 industries. Each designer worked from a simulated client brief with brand guidelines, logo, and product image, and delivered a final on-brand square creative in Figma, exactly as they would for a paying client.35 creatives · 35 brands · 4 industries · CC BY 4.0
  2. Image
    VideoDataset
    Video editing trajectories234 annotated steps across 4 computer-use trajectories, recorded as professional editors built vertical short-form social reels in Adobe Premiere Pro. Every step pairs a screenshot with a first-person thought, a structured action, and executable grounding: a Premiere MCP tool call, a keyboard shortcut, a menu path, or a coordinate click.234 steps · 4 trajectories · 9:16 reels · CC BY 4.0
  3. Image
    ImageDataset
    Photoshop design trajectories294 annotated steps across 14 computer-use trajectories, recorded as five professional designers built advertising and fashion-editorial assets in Adobe Photoshop and the browser. Every step pairs a screenshot with a first-person thought, a structured action, and executable grounding: an MCP tool call, a keyboard shortcut, a menu path, or a coordinate click.294 steps · 14 trajectories · 5 designers · CC BY 4.0
  4. Image
    VideoDataset
    Video model annotations544 timestamped notes from professional video editors on 15 product videos made by Google Veo 3.1, Adobe Firefly Video, and Grok Imagine, all running the same five prompts. Every note carries a timestamp range, a comment, a quality dimension, and a severity rating.544 annotations · 15 videos · 3 models · CC BY 4.0
Latest research39 studies

Every battle, model profile, and field note, sorted by newest first, tagged by domain.

  1. 08/26/2026Web DesignField Note
    How model performance shifted across four website design evalsGPT 5.6 Sol led the first three evals and every loose brief. Once the work was staged across ideation, mockup, and refinement, the Claude models moved ahead.Read
    GPT win rate vs Fable
    Eval 1
    59%
    Eval 2
    72%
    Eval 3
    70%
    HCB
    41%
    Head-to-head general preference · GPT led three evals, then fell to 41% in HCB
  2. 08/14/2026Web DesignBattle
    Claude models lead landing page design, but no model wins every design stageSix models built landing pages for three products across ideation, mockup, and refinement. Opus and Fable finished first and second overall, but a different model led each stage of the work.Read
    Win rate
    Claude Opus 5
    1st
    59.4
    Claude Fable 5
    2nd
    57.1
    GPT-5.6 Sol
    3rd
    52.1
    Kimi K3
    4th
    50.7
    Gemini 3.6 Flash
    5th
    45.5
    Muse Spark 1.1
    6th
    35.2
    Share of head-to-head comparisons won across 3,240 pairwise decisions · August 2026
  3. 08/12/2026VideoBattle
    Ad videos: six models across ideation, mockup, and refinementThe Human Creativity Benchmark moves to video: six models, three client campaigns, three phases each, judged head to head by working creatives. Seedance 2.0 won the set, and no model held the product together once it started moving.Read
    68.3%Seedance 2.0 overall win rate
    3,240Pairwise judgments
    6 × 3 × 3Models × campaigns × phases
    Head-to-head judgments by professional creatives · August 2026
  4. 08/07/2026Web DesignField Note
    What stands between Opus 5 and client-ready pages: layout and readabilityNine designers annotated 20 landing pages one at a time, 738 notes in all. On Opus 5's pages the notes clustered in layout and readability, while brand fit and originality drew the fewest flags in the set.Read
    738Annotations across 20 pages
    2 in 3Rated Opus 5 issues that needed rework
    20 × 9Pages × designers
    Solo page annotations, no matchups · August 2026
  5. 08/06/2026ImageBattle
    Six image models made real ads. Each broke differently.We scaled the Human Creativity Benchmark into a standing benchmark, starting with ad images: six models, three client campaigns, three phases each, judged head to head by professional creatives. Meta Muse Image won the set, and every model showed a signature failure.Read
    68.9%Muse Image overall win rate
    70.1%ChatGPT Images 2.0 win rate at ideation
    6 × 3 × 3Models × campaigns × phases
    Head-to-head judgments by professional creatives · August 2026
  6. 07/30/2026Web DesignBattle
    Claude Opus 5 nails the words and the mood. The finish is what holds it back.Across 600 blind comparisons and 400 write ups, designers liked the copy more than the rest of the page. Contrast, length and a few broken sections were what held it back.Read
    Win rate
    GPT 5.6 Sol
    1st
    65.3
    Kimi K3
    2nd
    61.7
    Claude Fable 5
    3rd
    41.7
    Claude Opus 5
    4th
    31.3
    Share of head-to-head matchups won across 600 blind comparisons · July 2026
  7. 07/24/2026Web DesignProfile
    Five designers, five rounds: Google Stitch proves its strength in ideation.Five expert product designers ran five Stitch iterations each on the same dashboard brief. The first prompts delivered wireframes and mockups in minutes; five rounds of refinement never raised the fidelity.Read
    2 + 3Wireframes and mockup stages after five iterations
    42 → 44Issues tagged, V1 → V5
    5 × 5Designers × iterations
    Think-aloud Stitch sessions · one shared brief · July 2026
  8. 07/23/2026Web DesignBattle
    Kimi K3 is a real rival to GPT 5.6 Sol on landing pagesKimi K3 arrived ranked first on Arena's Frontend Code Arena. Across 480 blind comparisons by 8 designers, it finished level with GPT 5.6 Sol on landing pages, and edged ahead on the detailed briefs.Read
    Win rate
    GPT 5.6 Sol
    1st
    65
    Kimi K3
    2nd
    63.3
    Gemini 3.5 Flash
    3rd
    38.8
    Claude Fable 5
    4th
    32.9
    Share of head-to-head matchups won across 480 blind comparisons · July 2026
  9. 07/16/2026Web DesignField Note
    Where four AI models break when they build a landing page8 designers annotated 40 pages from GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1, marking 754 failure points across 8 tags. Every model broke differently.Read
    754Failure points marked
    40Pages · 4 frontier models
    8Designers annotating
    Per-page failure annotation · 5 briefs · July 2026
  10. 07/16/2026ImageBattle
    GPT Image 2 won 41.9% of logo tournaments, and still only half its logos were client-ready10 brand designers judged GPT Image 2, Nano Banana Pro, MAI Image 2.5, and Meta Muse across 12 logo briefs. The winner is clear, and still only half its logos cleared the client bar.Read
    Tournament win rate
    GPT Image 2
    1st
    41.9%
    Nano Banana Pro
    2nd
    22.6%
    MAI Image 2.5
    3rd
    20.6%
    Meta Muse
    4th
    17.4%
    Share of logo tournaments won · 4 image models · 12 briefs · July 2026
  11. 07/14/2026ImageField Note
    Pretty isn’t the same as right: One image model runs away with brief fidelityFive designers, four criteria, 6,400 blind pairwise ratings. Nano Banana 2 takes first on every brief-fidelity axis while the aesthetics standings invert behind it.Read
    4/4Fidelity criteria swept
    67%Top typography win rate
    52%Best rival ceiling
    Nano Banana 2 vs the field · 6,400 blind ratings · July 2026
  12. 07/14/2026ImageField Note
    No model owns “aesthetics”: What 8,000 designer ratings tell us about taste in image modelsFive designers, five criteria, 8,000 blind pairwise ratings. Three of the four frontier models land within a few points of each other, and the best model changes depending on which dimension of visual quality you care about.Read
    Pooled win rate
    FLUX.2 [max]
    54%
    Nano Banana 2
    52%
    GPT Image 1.5
    50%
    Seedream 5.0 Lite
    44%
    Pooled win rate across 5 aesthetic criteria · 1,600 ratings each · July 2026
  13. 07/10/2026ImageBattle
    Reve 2.1 trailed Seedream 5.0 Pro by 2 points on wins, then finished last on Elo.Four-model image battle across 10 briefs spanning portrait, environment, product, and lifestyle work, judged blind by working creative professionals.Read
    Tournament wins
    Seedream 5.0 Pro
    1st
    32.6
    Reve 2.1
    2nd
    30.2
    MAI Image 2.5
    3rd
    20.9
    ChatGPT Images 2.0
    4th
    20
    Share of tournaments won · Seedream 5.0 Pro leads Reve 2.1 by 2.4 pp
  14. 07/10/2026Web DesignBattle
    Sol has taste. Fable takes direction.GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1 on the same ten landing page briefs, judged blind by 9 working designers as live interactive pages.Read
    Win rate
    GPT 5.6 Sol
    1st
    63.3
    Claude Fable 5
    2nd
    49.3
    Grok 4.5
    3rd
    48.1
    Muse Spark 1.1
    4th
    39.3
    Share of 540 pairwise matchups won · 10 landing page briefs · July 2026. Split by brief structure the ranking inverts: Fable leads structured briefs at 1569 Elo to Sol's 1455.
  15. 07/09/2026ImageBattle
    Seedream 5.0 Pro trails ChatGPT Images 2.0 by 8 points on wins, takes photorealism.Four-model image battle across 10 briefs spanning the capabilities ByteDance advertises. Seedream won photorealistic generation at 35.7% and finished first or second in 57% of its tournaments.Read
    Tournament wins
    ChatGPT Images 2.0
    1st
    35.9
    Nano Banana Pro
    2nd
    28.6
    Seedream 5.0 Pro
    3rd
    28
    Flux 2
    4th
    10.7
    ChatGPT Images 2.0 leads Seedream 5.0 Pro by 7.9 pp
  16. 07/08/2026ImageBattle
    Meta Muse trailed ChatGPT Images 2.0 by 5 points on wins, then beat it on Elo.Four-model style-transfer battle across 10 briefs. Meta Muse took 29.5% of tournament wins and was the only model top-two on both wins and Elo.Read
    Tournament wins
    GPT Images 2.0
    1st
    34.1
    Meta Muse
    2nd
    29.5
    Nano Banana Pro
    3rd
    23.9
    Flux 2
    4th
    19.3
    Share of 88 tournaments won · ChatGPT Images 2.0 leads Meta Muse by 4.6 pp
  17. 07/07/2026Web DesignBattle
    Write Fable a design spec and it wins 9 times out of 10Claude Fable 5 vs Claude Opus 4.8 on 5 real landing page and portfolio briefs, built in Claude Code, judged blind by 9 working designers.Read
    88.9%Best brief win rate
    11.1%Worst brief win rate
    51.1%Fable overall vs Opus
    Fable 5 vs Opus 4.8 · 90 blind matchups · July 2026
  18. 06/29/2026VideoBattle
    Frontier AI video models are nearly tied. None of them nail physics yet.A blind head-to-head of Seedance 2.0, Grok Imagine, Veo 3.1, and Adobe Firefly Video, judged by 12 professional video editors across 10 prompts.Read
    Prompts won
    Seedance 2.0
    1st
    4
    Grok Imagine
    2nd
    3
    Veo 3.1
    3rd
    3
    Adobe Firefly Video
    4th
    0
    Prompts won out of 10 · 4 frontier video models · June 2026. Seedance 2.0 also led on average score, by +0.07 over Grok Imagine on a 5-point scale.
  19. 06/17/2026ImageField Note
    Introducing Design Crit: we taught AI to judge design like a designer.Ten professional designers ranked four frontier image models across nine dimensions of real design work. The models can make the work. Nothing on the market could reliably judge it, until we trained on the right data.Read
    Agreement with panel
    Human designer
    Ceiling
    74.1%
    Trained on Design Crit
    Closes 46% of the gap
    61.1%
    Best off-the-shelf
    54.3%
    Chance
    50%
    Agreement with the five-designer majority · best off-the-shelf judge = HPSv2.1 · June 2026
  20. 06/03/2026ImageBattle
    Ideogram v4 won 47.9% of typography matchups.10 designers, 4 models, 240 images. Spelling is solved. Typographic craft and client-readiness are where Ideogram v4 pulls away.Read
    Typography 1st-place rate
    Ideogram v4
    1st
    47.9%
    Gemini 3.1 Flash
    2nd
    30%
    FLUX.2 [max]
    3rd
    15.5%
    Grok Imagine 1.0
    4th
    15%
    Share of 1st-place finishes on typography prompts · 4 image models · 20 prompts × 10 reviewers · rounds 1 + 2
  21. 06/01/2026ImageProfile
    Gemini reliably edits, but can it keep the rest of the image still?11 production-style sessions. Gemini made the edit 73% of the time, kept the rest of the image still 64%, held both in 55%.Read
    8 / 11Sessions that passed the edit-isolation test
    7 / 11Sessions that passed the pose-lock test
    6 / 11Sessions that passed both controllability tests
    Gemini controllability checks · 11 sessions · two tests per session (edit isolation + pose lock)
  22. 05/29/2026Web DesignBattle
    Cursor took 60% of head-to-heads. Claude Code took 63% of client meetings.Four coding tools, 24 outputs, five working designers. The tool designers preferred to look at and the tool they'd put their name on turned out to be different.Read
    Client-ready
    Claude Code
    Most client-ready
    63%
    Antigravity
    53%
    Cursor
    Pairwise winner
    47%
    Codex
    Last
    30%
    “Would you present this to a client?” · 4 coding tools · 24 outputs · May 2026
  23. 05/28/2026ImageProfile
    Gemini hit production-ready 24% of the time. One prompt pattern explains why.10 participants, 29 scored deliverables. The prompts that landed treated Gemini like a creative brief for a specific asset.Read
    7 / 29Deliverables Production Ready
    16 / 29Deliverables at Client V1+
    2 / 10Participants Production Ready on all 3 deliverables
    Gemini (Nano Banana Pro) production readiness · 10 participants, 29 scored deliverables
  24. 05/27/2026ImageProfile
    Can Adobe Firefly edit like Photoshop?8 targeted edits across 4 designer sessions. Firefly cleanly resolved 1, drifted on 5, and missed 2.Read
    1 / 8Edits cleanly resolved
    5 / 8Edits landed partial
    2 / 8Edits unresolved
    Adobe Firefly Edit · 8 attempts across 4 sessions (Hero + Social per session)
  25. 05/22/2026Cross-cuttingProfile
    The only prompt that got videos to production-ready in Adobe Firefly.4 designers, 3 deliverables each. The prompts that landed described physical direction, not aesthetic mood.Read
    1 / 4Videos production-ready, first pass
    2 / 4Social stills production-ready, first pass
    20Photoshop mentions across 4 sessions
    Adobe Firefly designer evaluation · 4 designers, 3 deliverables each
  26. 05/21/2026Web DesignField Note
    In Claude Design, your opening prompt decides the ceiling.5 designers, 5 openings, 1 luxury brief. The first prompt set what each session could reach.Read
    Specificity score (0–1)
    0.00.20.40.60.81.00P1Confirmation0.40P2Asset swap0.45P3Framework0.55P4Full brief0.75P5Hybrid
    First-prompt specificity by participant, Claude-scored · 5 designers, 1 brief
  27. 05/20/2026Web DesignProfile
    Claude Design gets you 40%, Figma gets the rest.5 sessions, 5 designers, 1 real-world client brief. Strong as a starting structure, breaks under precision edits.Read
    60 → 100%Designers flagging layout & spacing, Edit 1 → Edits 4–5
    5 / 5Sessions where layout was the recurring failure mode
    ≤40%Designer verdict: use it to here, then hand off
    Claude Design designer evaluation · 5 designers, 1 real-world client brief
  28. 05/18/2026ImageField Note
    With ChatGPT Images 2.0, "Text is solved." Typography isn't.42 sessions, 7 designers. ChatGPT Images 2.0 nails the typographic system, then breaks on the individual characters.Read
    +3 / +3 / +1Macro themes (hierarchy, brand fit, fonts) net positive
    −1 / −3Micro themes (legibility, size & weight) net negative
    0Designer mentions of size & weight as a strength
    Typography sentiment · 42 sessions, 7 designers, 6 briefs
  29. 05/14/2026ImageProfile
    ChatGPT Images 2.0 won every head-to-head. Here's where it still breaks.41 sessions, 7 designers, 6 briefs. GPT Image 2 nails the concept, then breaks at production.Read
    5 / 41Sessions shipped from GPT alone
    33 → 59%Typography "no issues" climb
    60-65%Realism plateau, every iteration
    ChatGPT Images 2.0 production readiness
  30. 05/13/2026ImageBattle
    The image-model leaderboard flips by brief.Four frontier image models, six brand campaigns, ranked blind by working creatives. GPT Image 2 wins the aggregate. Every other model owns a category.Read
    1st-place rate
    GPT Image 2
    1st
    40.7%
    Seedream 5.0 Lite
    2nd
    22.3%
    FLUX.2 [pro]
    3rd
    22.2%
    Gemini 3.1 Flash
    4th
    14.8%
    Share of 1st-place rankings across 6 brand campaigns · 4 frontier image models
  31. 05/12/2026ImageBattle
    Krea 2 Large is the #2 style-transfer model, closing on GPT Image 2.Four-model style-transfer evaluation. Krea took #2 on style fidelity, 0.14 points behind GPT Image 2.Read
    Style Fidelity (avg)
    GPT Image 2
    1st
    3.53
    Krea 2 Large
    2nd
    3.39
    Gemini 3 Pro
    3rd
    2.74
    Seedream 5.0 Lite
    4th
    2.42
    Style Fidelity average rating · Krea 2 Large takes #2, 0.14 points behind GPT Image 2
  32. 05/06/2026ImageBattle
    Seedream 5.0 Lite swept the field on product detail shots.A blind head-to-head against the leading image models from Google, OpenAI, and Black Forest Labs, evaluated by professional creatives.Read
    Win rate
    Seedream 5.0 Lite
    1st
    63.9%
    Gemini 3 Pro
    2nd
    52.8%
    GPT Image 1.5
    3rd
    44.4%
    FLUX.2 [max]
    4th
    38.9%
    Pairwise win rate · 4 leading image models · March 2026
  33. 05/05/2026Cross-cuttingField Note
    Creatives keep telling us the same thing about AI: every output looks the same.12 models, 5 creative domains. One repeated complaint from working evaluators: the work all looks the same.Read
    LOW BEST-PRACTICEHIGH BEST-PRACTICEHIGH STEERABILITYLOW STEERABILITY“Unreliable”“Opinionated engine”“Creative partner”“Full-spectrum tool”
    Convergence (best-practice) and divergence (steerability) as orthogonal signals.
  34. 04/28/2026VideoProfile
    Grok Imagine is the "Polisher" model. Hand off the early rounds, bring it in for refinement.The biggest phase-over-phase climb of any video model in the study. 3rd at ideation, 1st at refinement.Read
    Win rate
    Ideation
    3rd place
    46%
    Mockup
    44%
    Refinement
    1st place
    56%
    Grok Imagine win rate · +10pp climb from ideation to refinement
  35. 04/23/2026Cross-cuttingField Note
    The creative process has 3 phases. AI performs very differently in each.Ideation, mockup, refinement. AI fits differently at each phase, and the best creatives know where to hand off.Read
    Ideation
    Loose grip
    Mockup
    Narrowed
    Refinement
    Firm grip
    How tightly creatives hold control across phases · Qualitative
  36. 04/22/2026VideoProfile
    Veo 3.1 is the "Creative Director" model. Use it early, but hand off before refinement.61% win rate at ideation. 39% at refinement. The clearest model profile in our video evaluation.Read
    Win rate
    Ideation
    Peak
    61.1%
    Mockup
    55.6%
    Refinement
    Last place
    38.9%
    Veo 3.1 win rate · −22.2pp drop from ideation to refinement
  37. 04/21/2026Cross-cuttingField Note
    Solo creatives are earning more with AI and staying independent.Higher earning potential, more projects, no new hires. The survey from working independents.Read
    66%
    of independent creatives report higher earning potential since adopting AI
    26% no · 8% other
    Survey · Independent creatives on Contra
  38. 04/14/2026Web DesignBattle
    We tested 4 AI models with professional web designers. Claude won, but not the way you'd expect.Claude Opus 4.6, Gemini 3.1 Pro, ChatGPT 5.3 Codex, Qwen 3.5. The winner shifted at every phase.Read
    Leading win rate
    Ideation
    Claude leads
    79.8%
    Mockup
    Gemini takes over
    68.9%
    Refinement
    Claude narrows gap
    60%
    Per-phase leader shifts · Preview of the Human Creativity Benchmark
  39. 04/08/2026Cross-cuttingField Note
    AI isn't replacing creative professionals. It's making the best ones better.Survey of high-earning independent creatives. What they actually do with AI on real client work.Read
    <25%
    of AI output makes it to final deliverables
    The rest is stripped, reworked, or scrapped
    Dominant survey response · Independent creatives

The world's leading independent human data & creative evaluation lab.

Powered by Contra.

The world's leading independent human data & creative evaluation lab.

Powered by Contra.