- Muse Spark 1.3 lands 8th overall, ahead of DeepSeek v4 and right behind Gemini 3.6 Flash. In our tasks, it costs about 23.8x less than Claude Fable 5.1 and 19.2x less than GPT-5.6 Sol while slightly edging out Fable 5.1 in the Mockup phase, where the models work from a brief, brand guidelines, and an approved screenshot of the Ideation phase.
- The Ideation phase, working from a brief and the brand guidelines alone, is where it falls down. It wins only 10% of comparisons against Claude Fable 5.1 and GPT-5.6 Sol. Kimi K3 is the one open source frontier model it stays close to at this stage.
- Designers' complaints concentrate on layout and spacing, and on vertical rhythm in particular. In one brief, designers complained "the content is pushed against the top edge" and "text under image is all too close together vertically.”
- Against Muse Spark 1.1, it wins on every stage and every dimension we measured on the Human Creativity Benchmark: General Preference, Usability, Prompt Adherence, and Visual Appeal. The gap is widest in Refinement and on the Luma Iced Tea brief. In Refinement, 1.1 loses nearly every comparison.
- While Muse Spark 1.3 follows a structure similar to its predecessor, Muse Spark 1.1, it uses better typography. One designer said "the type choices are restrained and professional." The pages also have better responsive layouts. One designer complained that "responsive design seems to be broken" on an output from Muse Spark 1.1.
- Meta's claim of Muse Spark 1.3 having frontier-level performance is yet to be verified, and we will follow this up with another study once the max effort level model is out. Compared to its predecessor, 1.3 shows better performance in typography, color, and responsiveness, and we're excited to see what the Max model brings!
Creative intelligence.
Where AI models compete on real creative work. Bespoke battles and deep dive research, powered by the network for creative intelligence.
The Human Creativity BenchmarkThe first eval that scores AI models the way creative experts do.
Latest
08/26/2026Field NoteHow model performance shifted across four website design evals08/14/2026BattleClaude models lead landing page design, but no model wins every design stage08/12/2026BattleAd videos: six models across ideation, mockup, and refinement08/07/2026Field NoteWhat stands between Opus 5 and client-ready pages: layout and readability08/06/2026BattleSix image models made real ads. Each broke differently.07/30/2026BattleClaude Opus 5 nails the words and the mood. The finish is what holds it back.08/26/2026Field NoteHow model performance shifted across four website design evals08/14/2026BattleClaude models lead landing page design, but no model wins every design stage08/12/2026BattleAd videos: six models across ideation, mockup, and refinement08/07/2026Field NoteWhat stands between Opus 5 and client-ready pages: layout and readability08/06/2026BattleSix image models made real ads. Each broke differently.07/30/2026BattleClaude Opus 5 nails the words and the mood. The finish is what holds it back.
Best Performing Models6 models
Phase
LeaderboardPairwise comparison
1652
Muse Image
GPT Image 2
Nano Banana Pro
FLUX.2 [max]
Grok Imagine 1.0
Krea 2 Medium
General preference · Prompt adherence · Usability · Visual aesthetics
Generation Cost
Share
RankModelElo rating
- 1Muse Image1652
- 2GPT Image 21582
- 3Nano Banana Pro1568
- 4FLUX.2 [max]1478
- 5Grok Imagine 1.01411
- 6Krea 2 Medium1309
Methods & standardsSee how we score these models → Methodology
Dataset Library4 datasets
Open datasets behind our research, ready to download from Hugging Face.
ImageDatasetAd creative design35 finished social ad creatives, one for each of 35 unique synthetic brands across 4 industries. Each designer worked from a simulated client brief with brand guidelines, logo, and product image, and delivered a final on-brand square creative in Figma, exactly as they would for a paying client.35 creatives · 35 brands · 4 industries · CC BY 4.0
VideoDatasetVideo editing trajectories234 annotated steps across 4 computer-use trajectories, recorded as professional editors built vertical short-form social reels in Adobe Premiere Pro. Every step pairs a screenshot with a first-person thought, a structured action, and executable grounding: a Premiere MCP tool call, a keyboard shortcut, a menu path, or a coordinate click.234 steps · 4 trajectories · 9:16 reels · CC BY 4.0
ImageDatasetPhotoshop design trajectories294 annotated steps across 14 computer-use trajectories, recorded as five professional designers built advertising and fashion-editorial assets in Adobe Photoshop and the browser. Every step pairs a screenshot with a first-person thought, a structured action, and executable grounding: an MCP tool call, a keyboard shortcut, a menu path, or a coordinate click.294 steps · 14 trajectories · 5 designers · CC BY 4.0
VideoDatasetVideo model annotations544 timestamped notes from professional video editors on 15 product videos made by Google Veo 3.1, Adobe Firefly Video, and Grok Imagine, all running the same five prompts. Every note carries a timestamp range, a comment, a quality dimension, and a severity rating.544 annotations · 15 videos · 3 models · CC BY 4.0
Latest research39 studies
Every battle, model profile, and field note, sorted by newest first, tagged by domain.
08/26/2026Web DesignField NoteHow model performance shifted across four website design evalsGPT 5.6 Sol led the first three evals and every loose brief. Once the work was staged across ideation, mockup, and refinement, the Claude models moved ahead.ReadGPT win rate vs FableEval 159%Eval 272%Eval 370%HCB41%Head-to-head general preference · GPT led three evals, then fell to 41% in HCB
08/14/2026Web DesignBattleClaude models lead landing page design, but no model wins every design stageSix models built landing pages for three products across ideation, mockup, and refinement. Opus and Fable finished first and second overall, but a different model led each stage of the work.ReadWin rateClaude Opus 51st59.4Claude Fable 52nd57.1GPT-5.6 Sol3rd52.1Kimi K34th50.7Gemini 3.6 Flash5th45.5Muse Spark 1.16th35.2Share of head-to-head comparisons won across 3,240 pairwise decisions · August 2026
08/12/2026VideoBattleAd videos: six models across ideation, mockup, and refinementThe Human Creativity Benchmark moves to video: six models, three client campaigns, three phases each, judged head to head by working creatives. Seedance 2.0 won the set, and no model held the product together once it started moving.Read68.3%Seedance 2.0 overall win rate3,240Pairwise judgments6 × 3 × 3Models × campaigns × phasesHead-to-head judgments by professional creatives · August 2026
08/07/2026Web DesignField NoteWhat stands between Opus 5 and client-ready pages: layout and readabilityNine designers annotated 20 landing pages one at a time, 738 notes in all. On Opus 5's pages the notes clustered in layout and readability, while brand fit and originality drew the fewest flags in the set.Read738Annotations across 20 pages2 in 3Rated Opus 5 issues that needed rework20 × 9Pages × designersSolo page annotations, no matchups · August 2026
08/06/2026ImageBattleSix image models made real ads. Each broke differently.We scaled the Human Creativity Benchmark into a standing benchmark, starting with ad images: six models, three client campaigns, three phases each, judged head to head by professional creatives. Meta Muse Image won the set, and every model showed a signature failure.Read68.9%Muse Image overall win rate70.1%ChatGPT Images 2.0 win rate at ideation6 × 3 × 3Models × campaigns × phasesHead-to-head judgments by professional creatives · August 2026
07/30/2026Web DesignBattleClaude Opus 5 nails the words and the mood. The finish is what holds it back.Across 600 blind comparisons and 400 write ups, designers liked the copy more than the rest of the page. Contrast, length and a few broken sections were what held it back.ReadWin rateGPT 5.6 Sol1st65.3Kimi K32nd61.7Claude Fable 53rd41.7Claude Opus 54th31.3Share of head-to-head matchups won across 600 blind comparisons · July 2026
07/24/2026Web DesignProfileFive designers, five rounds: Google Stitch proves its strength in ideation.Five expert product designers ran five Stitch iterations each on the same dashboard brief. The first prompts delivered wireframes and mockups in minutes; five rounds of refinement never raised the fidelity.Read2 + 3Wireframes and mockup stages after five iterations42 → 44Issues tagged, V1 → V55 × 5Designers × iterationsThink-aloud Stitch sessions · one shared brief · July 2026
07/23/2026Web DesignBattleKimi K3 is a real rival to GPT 5.6 Sol on landing pagesKimi K3 arrived ranked first on Arena's Frontend Code Arena. Across 480 blind comparisons by 8 designers, it finished level with GPT 5.6 Sol on landing pages, and edged ahead on the detailed briefs.ReadWin rateGPT 5.6 Sol1st65Kimi K32nd63.3Gemini 3.5 Flash3rd38.8Claude Fable 54th32.9Share of head-to-head matchups won across 480 blind comparisons · July 2026
07/16/2026Web DesignField NoteWhere four AI models break when they build a landing page8 designers annotated 40 pages from GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1, marking 754 failure points across 8 tags. Every model broke differently.Read754Failure points marked40Pages · 4 frontier models8Designers annotatingPer-page failure annotation · 5 briefs · July 2026
07/16/2026ImageBattleGPT Image 2 won 41.9% of logo tournaments, and still only half its logos were client-ready10 brand designers judged GPT Image 2, Nano Banana Pro, MAI Image 2.5, and Meta Muse across 12 logo briefs. The winner is clear, and still only half its logos cleared the client bar.ReadTournament win rateGPT Image 21st41.9%Nano Banana Pro2nd22.6%MAI Image 2.53rd20.6%Meta Muse4th17.4%Share of logo tournaments won · 4 image models · 12 briefs · July 2026
07/14/2026ImageField NotePretty isn’t the same as right: One image model runs away with brief fidelityFive designers, four criteria, 6,400 blind pairwise ratings. Nano Banana 2 takes first on every brief-fidelity axis while the aesthetics standings invert behind it.Read4/4Fidelity criteria swept67%Top typography win rate52%Best rival ceilingNano Banana 2 vs the field · 6,400 blind ratings · July 2026
07/14/2026ImageField NoteNo model owns “aesthetics”: What 8,000 designer ratings tell us about taste in image modelsFive designers, five criteria, 8,000 blind pairwise ratings. Three of the four frontier models land within a few points of each other, and the best model changes depending on which dimension of visual quality you care about.ReadPooled win rateFLUX.2 [max]54%Nano Banana 252%GPT Image 1.550%Seedream 5.0 Lite44%Pooled win rate across 5 aesthetic criteria · 1,600 ratings each · July 2026
07/10/2026ImageBattleReve 2.1 trailed Seedream 5.0 Pro by 2 points on wins, then finished last on Elo.Four-model image battle across 10 briefs spanning portrait, environment, product, and lifestyle work, judged blind by working creative professionals.ReadTournament winsSeedream 5.0 Pro1st32.6Reve 2.12nd30.2MAI Image 2.53rd20.9ChatGPT Images 2.04th20Share of tournaments won · Seedream 5.0 Pro leads Reve 2.1 by 2.4 pp
07/10/2026Web DesignBattleSol has taste. Fable takes direction.GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1 on the same ten landing page briefs, judged blind by 9 working designers as live interactive pages.ReadWin rateGPT 5.6 Sol1st63.3Claude Fable 52nd49.3Grok 4.53rd48.1Muse Spark 1.14th39.3Share of 540 pairwise matchups won · 10 landing page briefs · July 2026. Split by brief structure the ranking inverts: Fable leads structured briefs at 1569 Elo to Sol's 1455.
07/09/2026ImageBattleSeedream 5.0 Pro trails ChatGPT Images 2.0 by 8 points on wins, takes photorealism.Four-model image battle across 10 briefs spanning the capabilities ByteDance advertises. Seedream won photorealistic generation at 35.7% and finished first or second in 57% of its tournaments.ReadTournament winsChatGPT Images 2.01st35.9Nano Banana Pro2nd28.6Seedream 5.0 Pro3rd28Flux 24th10.7ChatGPT Images 2.0 leads Seedream 5.0 Pro by 7.9 pp
07/08/2026ImageBattleMeta Muse trailed ChatGPT Images 2.0 by 5 points on wins, then beat it on Elo.Four-model style-transfer battle across 10 briefs. Meta Muse took 29.5% of tournament wins and was the only model top-two on both wins and Elo.ReadTournament winsGPT Images 2.01st34.1Meta Muse2nd29.5Nano Banana Pro3rd23.9Flux 24th19.3Share of 88 tournaments won · ChatGPT Images 2.0 leads Meta Muse by 4.6 pp
07/07/2026Web DesignBattleWrite Fable a design spec and it wins 9 times out of 10Claude Fable 5 vs Claude Opus 4.8 on 5 real landing page and portfolio briefs, built in Claude Code, judged blind by 9 working designers.Read88.9%Best brief win rate11.1%Worst brief win rate51.1%Fable overall vs OpusFable 5 vs Opus 4.8 · 90 blind matchups · July 2026
06/29/2026VideoBattleFrontier AI video models are nearly tied. None of them nail physics yet.A blind head-to-head of Seedance 2.0, Grok Imagine, Veo 3.1, and Adobe Firefly Video, judged by 12 professional video editors across 10 prompts.ReadPrompts wonSeedance 2.01st4Grok Imagine2nd3Veo 3.13rd3Adobe Firefly Video4th0Prompts won out of 10 · 4 frontier video models · June 2026. Seedance 2.0 also led on average score, by +0.07 over Grok Imagine on a 5-point scale.
06/17/2026ImageField NoteIntroducing Design Crit: we taught AI to judge design like a designer.Ten professional designers ranked four frontier image models across nine dimensions of real design work. The models can make the work. Nothing on the market could reliably judge it, until we trained on the right data.ReadAgreement with panelHuman designerCeiling74.1%Trained on Design CritCloses 46% of the gap61.1%Best off-the-shelf54.3%Chance50%Agreement with the five-designer majority · best off-the-shelf judge = HPSv2.1 · June 2026
06/03/2026ImageBattleIdeogram v4 won 47.9% of typography matchups.10 designers, 4 models, 240 images. Spelling is solved. Typographic craft and client-readiness are where Ideogram v4 pulls away.ReadTypography 1st-place rateIdeogram v41st47.9%Gemini 3.1 Flash2nd30%FLUX.2 [max]3rd15.5%Grok Imagine 1.04th15%Share of 1st-place finishes on typography prompts · 4 image models · 20 prompts × 10 reviewers · rounds 1 + 2
06/01/2026ImageProfileGemini reliably edits, but can it keep the rest of the image still?11 production-style sessions. Gemini made the edit 73% of the time, kept the rest of the image still 64%, held both in 55%.Read8 / 11Sessions that passed the edit-isolation test7 / 11Sessions that passed the pose-lock test6 / 11Sessions that passed both controllability testsGemini controllability checks · 11 sessions · two tests per session (edit isolation + pose lock)
05/29/2026Web DesignBattleCursor took 60% of head-to-heads. Claude Code took 63% of client meetings.Four coding tools, 24 outputs, five working designers. The tool designers preferred to look at and the tool they'd put their name on turned out to be different.ReadClient-readyClaude CodeMost client-ready63%Antigravity53%CursorPairwise winner47%CodexLast30%“Would you present this to a client?” · 4 coding tools · 24 outputs · May 2026
05/28/2026ImageProfileGemini hit production-ready 24% of the time. One prompt pattern explains why.10 participants, 29 scored deliverables. The prompts that landed treated Gemini like a creative brief for a specific asset.Read7 / 29Deliverables Production Ready16 / 29Deliverables at Client V1+2 / 10Participants Production Ready on all 3 deliverablesGemini (Nano Banana Pro) production readiness · 10 participants, 29 scored deliverables
05/27/2026ImageProfileCan Adobe Firefly edit like Photoshop?8 targeted edits across 4 designer sessions. Firefly cleanly resolved 1, drifted on 5, and missed 2.Read1 / 8Edits cleanly resolved5 / 8Edits landed partial2 / 8Edits unresolvedAdobe Firefly Edit · 8 attempts across 4 sessions (Hero + Social per session)
05/22/2026Cross-cuttingProfileThe only prompt that got videos to production-ready in Adobe Firefly.4 designers, 3 deliverables each. The prompts that landed described physical direction, not aesthetic mood.Read1 / 4Videos production-ready, first pass2 / 4Social stills production-ready, first pass20Photoshop mentions across 4 sessionsAdobe Firefly designer evaluation · 4 designers, 3 deliverables each
05/21/2026Web DesignField NoteIn Claude Design, your opening prompt decides the ceiling.5 designers, 5 openings, 1 luxury brief. The first prompt set what each session could reach.ReadSpecificity score (0–1)First-prompt specificity by participant, Claude-scored · 5 designers, 1 brief
05/20/2026Web DesignProfileClaude Design gets you 40%, Figma gets the rest.5 sessions, 5 designers, 1 real-world client brief. Strong as a starting structure, breaks under precision edits.Read60 → 100%Designers flagging layout & spacing, Edit 1 → Edits 4–55 / 5Sessions where layout was the recurring failure mode≤40%Designer verdict: use it to here, then hand offClaude Design designer evaluation · 5 designers, 1 real-world client brief
05/18/2026ImageField NoteWith ChatGPT Images 2.0, "Text is solved." Typography isn't.42 sessions, 7 designers. ChatGPT Images 2.0 nails the typographic system, then breaks on the individual characters.Read+3 / +3 / +1Macro themes (hierarchy, brand fit, fonts) net positive−1 / −3Micro themes (legibility, size & weight) net negative0Designer mentions of size & weight as a strengthTypography sentiment · 42 sessions, 7 designers, 6 briefs
05/14/2026ImageProfileChatGPT Images 2.0 won every head-to-head. Here's where it still breaks.41 sessions, 7 designers, 6 briefs. GPT Image 2 nails the concept, then breaks at production.Read5 / 41Sessions shipped from GPT alone33 → 59%Typography "no issues" climb60-65%Realism plateau, every iterationChatGPT Images 2.0 production readiness
05/13/2026ImageBattleThe image-model leaderboard flips by brief.Four frontier image models, six brand campaigns, ranked blind by working creatives. GPT Image 2 wins the aggregate. Every other model owns a category.Read1st-place rateGPT Image 21st40.7%Seedream 5.0 Lite2nd22.3%FLUX.2 [pro]3rd22.2%Gemini 3.1 Flash4th14.8%Share of 1st-place rankings across 6 brand campaigns · 4 frontier image models
05/12/2026ImageBattleKrea 2 Large is the #2 style-transfer model, closing on GPT Image 2.Four-model style-transfer evaluation. Krea took #2 on style fidelity, 0.14 points behind GPT Image 2.ReadStyle Fidelity (avg)GPT Image 21st3.53Krea 2 Large2nd3.39Gemini 3 Pro3rd2.74Seedream 5.0 Lite4th2.42Style Fidelity average rating · Krea 2 Large takes #2, 0.14 points behind GPT Image 2
05/06/2026ImageBattleSeedream 5.0 Lite swept the field on product detail shots.A blind head-to-head against the leading image models from Google, OpenAI, and Black Forest Labs, evaluated by professional creatives.ReadWin rateSeedream 5.0 Lite1st63.9%Gemini 3 Pro2nd52.8%GPT Image 1.53rd44.4%FLUX.2 [max]4th38.9%Pairwise win rate · 4 leading image models · March 2026
05/05/2026Cross-cuttingField NoteCreatives keep telling us the same thing about AI: every output looks the same.12 models, 5 creative domains. One repeated complaint from working evaluators: the work all looks the same.ReadConvergence (best-practice) and divergence (steerability) as orthogonal signals.
04/28/2026VideoProfileGrok Imagine is the "Polisher" model. Hand off the early rounds, bring it in for refinement.The biggest phase-over-phase climb of any video model in the study. 3rd at ideation, 1st at refinement.ReadWin rateIdeation3rd place46%Mockup44%Refinement1st place56%Grok Imagine win rate · +10pp climb from ideation to refinement
04/23/2026Cross-cuttingField NoteThe creative process has 3 phases. AI performs very differently in each.Ideation, mockup, refinement. AI fits differently at each phase, and the best creatives know where to hand off.ReadIdeationLoose gripMockupNarrowedRefinementFirm gripHow tightly creatives hold control across phases · Qualitative
04/22/2026VideoProfileVeo 3.1 is the "Creative Director" model. Use it early, but hand off before refinement.61% win rate at ideation. 39% at refinement. The clearest model profile in our video evaluation.ReadWin rateIdeationPeak61.1%Mockup55.6%RefinementLast place38.9%Veo 3.1 win rate · −22.2pp drop from ideation to refinement
04/21/2026Cross-cuttingField NoteSolo creatives are earning more with AI and staying independent.Higher earning potential, more projects, no new hires. The survey from working independents.Readof independent creatives report higher earning potential since adopting AI26% no · 8% otherSurvey · Independent creatives on Contra
04/14/2026Web DesignBattleWe tested 4 AI models with professional web designers. Claude won, but not the way you'd expect.Claude Opus 4.6, Gemini 3.1 Pro, ChatGPT 5.3 Codex, Qwen 3.5. The winner shifted at every phase.ReadLeading win rateIdeationClaude leads79.8%MockupGemini takes over68.9%RefinementClaude narrows gap60%Per-phase leader shifts · Preview of the Human Creativity Benchmark
04/08/2026Cross-cuttingField NoteAI isn't replacing creative professionals. It's making the best ones better.Survey of high-earning independent creatives. What they actually do with AI on real client work.Readof AI output makes it to final deliverablesThe rest is stripped, reworked, or scrappedDominant survey response · Independent creatives