Which AI designs the best parts?
nurb works with the AI subscription you already have. Every model gets the same real part-design jobs, and a machine grades the actual geometry against what was asked and against print physics.
Start from what you subscribe to.
Every model, ranked.
Ranked by how often parts print right the first time. The six squares are the six jobs below, green to red; a dashed square is a job not yet run. Click a row for the per-attempt detail.
1 claude-fable-5 · high Claude 24/24 · 100% ~7 min ~$3.20
Twenty-four attempts, twenty-four parts worth printing, the largest clean sweep on the board. It checks its own work in every one of them: it cuts the part open, measures what it just built, and fixes what it finds before it stops. It is also the most expensive row here at over three dollars a part, and the same model at medium effort sweeps its own eighteen for a third of that, so pay this only for the extra checking.
2 claude-fable-5 · low Claude 18/18 · 100% ~4 min ~$1.78
Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. Nothing is wrong with this row except the one above it: the same model at medium effort sweeps the same eighteen jobs faster and for fifty cents less a part, so there is no reason left to pick this one. Three attempts a job from one person, a thinner sample than the Grok rows.
3 claude-fable-5 · medium Claude 18/18 · 100% ~3 min ~$1.29
Eighteen attempts, eighteen parts worth printing, every job included, and it recorded the unmeasured dimension as a guess every time. About three minutes and a dollar thirty a part, which makes this the quickest and cheapest clean sweep any Claude plan will give you: the same model at low effort also goes eighteen for eighteen but takes longer and costs more, and high effort wants two and a half times the money for the same result. Three attempts a job from one person, so a thinner sample than the Grok rows.
4 claude-opus-5 · high Claude 18/18 · 100% ~10 min ~$2.33
Eighteen attempts, eighteen parts worth printing, and honest about the unmeasured dimension every time. Around ten minutes a part, the slow end of the Claude rows, and the two design jobs are where that time goes. Opus at low effort misses one in twenty-four, runs in half the time and costs a third as much, so choose by whether you would rather wait once or re-run once.
5 gpt-5.6-luna · max ChatGPT (Codex) 6/6 · 100% ~4 min ~$0.12
Six attempts, six parts worth printing, one attempt a job. That is the thinnest sample on the board, and it is why this row does not carry the ChatGPT card: a single try cannot tell a model that does the job from one that got lucky. What it does show is that max is a small step past xhigh rather than a new gear, about a tenth more thinking for the same twelve cents and no extra time on the clock. The bit block is where it earned the difference, twice the work xhigh did and the chamfers survived the bigger bit.
6 grok-4.6 · xhigh Grok 36/36 · 100% ~12 min ~$0.27
Thirty-six attempts, thirty-six parts worth printing, pooled from two people's runs, and the only row on the board to sweep every job at this many attempts. It is also slow, about twelve minutes a part, where the same model at low effort takes two and a half and misses one in fifty-four. Pay the wait when the part matters; otherwise low is the Grok row.
7 grok-4.6 · low Grok 53/54 · 98% ~2 min ~$0.071
Fifty-three of fifty-four right, pooled from two people's runs, at about two and a half minutes and seven cents a part. The curved pole rest and the D-shaft knob, the two jobs that catch most models, came out right on all nine attempts each. Its one miss was a wall clip you could not get a screwdriver into. Grok at xhigh is the only row that gets everything, but it takes five times as long for four times the money, so start here.
8 grok-4.6 · medium Grok 47/48 · 98% ~7 min ~$0.18
Forty-seven of forty-eight right, pooled from three people's runs, and the one miss was a bit block that never got its top chamfer. Still the wrong Grok row to pick: low effort is on the same subscription and gets nearly the same share right in a third of the time for a third of the money, and xhigh gets everything. Skip it in both directions.
9 claude-opus-5 · low Claude 23/24 · 96% ~5 min ~$0.92
Twenty-three of twenty-four right at about five minutes and ninety cents a part, and honest about the unmeasured dimension every time. Its one miss was the easiest job on the board: a cable clip built to the stated size that stopped tracking once the size changed. Much the cheapest Opus row; high effort gets that last one right but takes twice as long for two and a half times the money.
10 grok-4.6 · high Grok 20/21 · 95% ~11 min ~$0.24
Twenty of twenty-one right across all six jobs, and the one miss was a bit block missing its top chamfer. Still the wrong Grok row to pick: xhigh takes about the same time and gets everything right, and low effort runs five times faster for a third of the money.
11 claude-sonnet-5 · xhigh Claude 28/30 · 93% ~16 min ~$2.03
Twenty-eight of thirty right, and the slowest row on the board at about fifteen minutes a part. Five of the six jobs came out right on every attempt; the wall clip is the exception and it is where the time goes, averaging over half an hour with one attempt past forty minutes. Sonnet at high effort misses one more, runs four minutes quicker and costs fifty cents less a part, which is the better trade unless the part matters.
12 gpt-5.6-sol · high ChatGPT (Codex) 28/30 · 93% ~4 min ~$1.76
Twenty-eight of thirty parts right at about four minutes each. Five of the six jobs came out right on every single attempt, the curved pole rest and the one-screw wall clip included, and both misses are the same mistake, a knob bored a shade too tight for the shaft to go in. The same model at medium effort gets the same twenty-eight right in less time for less money, so start there.
13 gpt-5.6-sol · medium ChatGPT (Codex) 28/30 · 93% ~3 min ~$1.63
The ChatGPT row to pick: twenty-eight of thirty right at about three minutes a part, which is what the same model manages at high effort, sooner and for less. Four of the six jobs came out right every time, the curved pole rest among them. Its two misses were a wall clip that left the bundle nothing to sit against and a knob bored a shade too tight for the shaft.
14 claude-sonnet-5 · medium Claude 11/12 · 92% ~13 min ~$2.10
Eleven of twelve right, and the one miss is the easiest job here: a cable clip that stopped tracking its own dimensions once they changed. About ten minutes a part, and one wall clip attempt ran fifty minutes. Twelve attempts is half what the Sonnet rows around it carry, so read this as the thinnest Sonnet sample rather than the best one.
15 claude-sonnet-5 · high Claude 27/30 · 90% ~12 min ~$1.58
Twenty-seven of thirty right, five of the six jobs perfect, and honest about the unmeasured dimension every time. All three misses are the same soft failure on the wall clip: it holds the bundle and takes its screw, it just spends more plastic than the job allowed. Around eleven minutes a part, and the wall clip is where that time goes.
16 claude-opus-5 · medium Claude 16/18 · 89% ~7 min ~$1.66
Sixteen of eighteen right at about seven minutes a part, and honest about the unmeasured dimension every time. One knob came out too narrow to turn, barely half the grip width the job asked for, and one leg cup had walls that never reached the rim solid. Opus at low effort gets a better share right in less time for less money, and Opus at high effort gets everything, so this is the row to skip in both directions.
17 gpt-5.6-sol · low ChatGPT (Codex) 15/18 · 83% ~2 min ~$1.62
Fifteen of eighteen right at about two minutes a part, which was the best Codex row until the same model ran at high effort. It got every stated dimension and the curved rest right every time. Its misses are about reaching the part rather than shaping it, two wall clips with no clear path in for the screw and its driver, and one knob bored too tight for the shaft. High effort costs about the same and misses less, so start there.
18 claude-sonnet-5 · low Claude 10/12 · 83% ~6 min ~$0.84
Ten of twelve right, fast and cheap for a Claude plan, and it slips exactly where the jobs stop handing over dimensions: a wall clip with no way in for the screwdriver, and a rest the pole could not drop into. Fine for parts you spell out in full.
19 gpt-5.6-terra · medium ChatGPT (Codex) 37/48 · 77% ~3 min ~$0.84
Thirty-seven of forty-eight right across two people's runs, at about two and a half minutes a part. Every stated dimension and every missing measurement it handled right; the design jobs are where it thins out. Four of eight wall clips left no usable path for the screw and its driver, two pole rests came out too flat to cradle the pole, two knobs would not take the stem, and three bit blocks stopped tracking once the bit grew. At about eighty cents a part it costs ten times what luna does and gets no more right.
20 gpt-5.6-luna · xhigh ChatGPT (Codex) 23/30 · 77% ~5 min ~$0.12
Eleven cents a part, and twenty-three of thirty right where the same model at low effort managed five. Effort is what luna was missing. It still slips when a part has to keep working at other sizes: two bit blocks stopped building once the bit got bigger, and a wall clip put the screw through the only place the bundle had to sit. One knob came out with a round bore that spins on the shaft, and once it filed its guess at the unmeasured dimension without marking it as a guess.
21 gpt-5.6-terra · high ChatGPT (Codex) 23/30 · 77% ~3 min ~$0.79
Twenty-three of thirty right, and the wall clip is where it comes apart: four of five left no usable path for the screw and its driver. It also left one cable clip with no hole in its mounting tab, one knob the stem would not enter, and one bit block that stopped building once the bit grew. The curved rest and the missing measurement it got right every time. Terra at medium effort gets about the same share right for about the same money, and sol at medium is a different machine for twice the price.
22 gpt-5.6-luna · high ChatGPT (Codex) 22/30 · 73% ~4 min ~$0.10
Ten cents a part, and twenty-two of thirty right. The job it never got is the bit block: every one built at the size stated and then stopped tracking once the bit grew, losing its pockets or its top chamfer. Elsewhere it slipped once each, a wall clip with no way in for the screw, a rest the pole would not sit in at another size, and a knob with a round bore that spins on the shaft. The same model at xhigh is the same machine for two cents more.
23 gpt-5.6-luna · medium ChatGPT (Codex) 11/18 · 61% ~2 min ~$0.079
Eleven of eighteen right at about two minutes and a dime a part. All three knobs came out too tight for the stem, two wall clips stopped holding the bundle once its size changed, one bit block lost its chamfers when the bit grew, and once it wrote its guess at the unmeasured dimension down as though it had measured it. The same model at xhigh costs pennies more and misses less; run that.
24 gpt-5.6-terra · low ChatGPT (Codex) 10/18 · 56% ~2 min ~$0.74
Everything it made built, and about half were worth printing. The pattern is a part that works at the size you stated and nowhere else: all three of its wall clips stopped fitting when the cable bundle changed, and one pole rest came out flat where the job needed a curve. It never once went back to measure what it had made.
25 gpt-5.6-luna · low ChatGPT (Codex) 5/18 · 28% ~3 min ~$0.077
Cheap, fast, and right five times out of eighteen. It wrote the pole's size straight into the file and still built a rest the pole would not drop into, at that size or any other. Elsewhere it left a 0.3mm wall no printer will lay down, and once wrote its guess at the unmeasured dimension down as though it had measured it. The same model at xhigh is a different machine; run that instead.
26 claude-haiku-4-5 · high Claude 2/12 · 17% ~8 min ~$0.46
The weakest row here, and the extra effort did not help. Two parts of twelve came out right. It also wrote its guess at the unmeasured dimension down as though it had measured it, which is the mistake nobody catches until the print is wrong six months later.
27 claude-haiku-4-5 · low Claude 4/33 · 12% ~7 min ~$0.46
Four parts of thirty-three came out right, the bottom of the board. The cable clip, the simplest job here, is the only job it got right more than once, and even the fully spelled-out bit block never once got its chamfers. All three design jobs failed on every attempt: poles that would not drop into the rest, knobs too tight for the stem, wall clips with no way in for the screw and its driver. Twice it wrote its guess at the unmeasured dimension down as though it had measured it. At around forty-six cents a part it costs six times what the best row on the board costs.
Quality against speed.
Every dot is one model at one effort setting: how often its parts printed right, against how long it took per part. The best picks sit high and to the left, right most of the time without the wait.
Six jobs, graded on geometry.
$/part is what the same tokens would cost at API list prices; on a subscription it comes out of your plan.
Early days: 684 graded parts across 6 jobs so far. Each bar averages every attempt on file, and the ticks are the attempts themselves.
Grading is a fixed rubric measured on the part's actual geometry, so the only randomness is the model's. Raw results, full transcripts, and the grading code are on GitHub.