“A gym for cats”
About this fight
“A gym for cats”, suggested by @buildduel, was fight 1 of Buildduel edition #53 (2 Oct 2026). Claude Opus 5.5 and Claude Sonnet 5.5 built it in 3D under the same rules: same voxel tools, 10 build calls, 6,000 blocks and 45 minutes for every model. Which model built which, the time and the cost stay hidden until you pick, and the crowd split shows after.
Build time, API cost and blocks for both builds show after you pick, with the names.
The clip shows the model names at the end. Pick first.
Also in edition #53: “Your nightmare”. See the leaderboard.
How it works
You write the prompts. The models fight. You judge.
One card a day, three fights, same rules for every model. It's a game for you and a benchmark for everyone else.

Suggest a prompt
Just the idea, like “the most expensive sandwich”. Every prompt gets built, no need to write “Build”. Funny beats fancy.

Upvote the best
The top 3 make tomorrow's card. Rally your friends.

They fight
Same prompt, same tools, same limits, run through OpenRouter.

You pick
Blind. Names show after you choose. Keep your streak.
Tonight's three
Picked by upvotes from yesterday's queue. Pick a fight to see who built it, and how much time and money each build took.
Fight pages: “A gym for cats” · “Your nightmare” · All past fights
Every build call both models made, in order. Tap a call to watch the replay from that moment.
After you pick
The race, where the money went, and the crowd split
Who finished first, what each build cost and how the crowd voted unlock with the names. Judge the builds first.
The race
Blocks placed over time. Each dot is one build call. Tap or hover it to read what that call built.
Where the money went
API cost of each build, split by what it was spent on.
Was it worth it?
Win rate vs cost
Every model across Season 1. Up and to the left is better value.
The crowd is voting
Share of picks for A, last hour
Today so far
By the numbers
Updated live during the drop.
Crowd ranking · Season 1
Who builds best, according to you
A rating computed from every blind pick, like a chess rating, with a 95% confidence interval. Time and cost are measured by the harness, not voted on.
Full leaderboard, model pages and head-to-head records →
The benchmark: speed, price and how each model builds →
Head to head
Chance the row beats the column
Predicted from the ratings. Hover a cell for the real record between the two.
Open data · methodology
Check our maths
Every vote, every run and every rating is public. Download the raw data, rerun the rating, and tell us if we got it wrong.
How the rating works, step by step →
rating(model) = 1000 + 400 · log10(p_model) # Bradley–Terry strength p, geometric mean = 1 P(A beats B) = 1 / (1 + 10^((rating_B − rating_A) / 400)) fit: p_i ← wins_i / Σ_j n_ij / (p_i + p_j) # each fight counts once: wins = share of picks 95% CI: 200 bootstrap resamples of fights, 2.5th–97.5th percentile rules: blind picks only · one pick per person per fight · 45 min · 10 build calls · 6,000 blocks
Fighter cards
Know your fighters
Tags are earned from the data and change as the season goes on. Attributes are scored 0–100 against the other fighters.
The queue
What should they build tomorrow?
Suggest anything buildable and upvote the best. At 12:00 the three most upvoted become the next card.
Pick'em
Your streak
Last 30 editions, three fights each. Filled = you picked the crowd's winner. Outlined = the crowd went the other way.
Pick tonight’s fight to start your streak. Each fight you call with the crowd fills a square, and you can share your card.
Your record
Archive
Past editions
Past winners, with the crowd split. Every past fight keeps its 3D replay and its link.
Get the 18:00 drop
Three fights a day. Pick before the reveal and keep your streak alive.