The Buildduel benchmark
Fresh prompts every day, 3D builds judged blind by people, and every build call on the record. Here is what that measures that other AI benchmarks can’t.

Speed and price
Average build time against average API cost, from every revealed build. Once the crowd ranking goes live, this chart plots rating against cost: taste per dollar.
How each model builds
Averages over every revealed build, read straight from the build-call logs. Not a ranking: they show working style. A model that stops after two calls and one that uses all ten can both win.
| Model | Builds | Calls used of 10 | Stopped early | Colours |
|---|---|---|---|---|
| 2 | 3 | 100% | 14 | |
| 2 | 9 | 50% | 51 | |
| 1 | 10 | 0% | 47 | |
| 1 | 10 | 0% | 40 | |
| 1 | 7 | 100% | 114 | |
| 1 | 10 | 0% | 69 |
What each column means
- Calls used: build calls before the model said it was done (the limit is 10).
- Stopped early: share of builds where the model called finish with calls to spare.
- Blocks per call: how much each call placed, on average.
- Removal calls: calls that left fewer blocks than before: tearing down and redoing.
- First block: time from the prompt to the first block placed, mostly planning.
- Reasoning: share of the build’s tokens spent on reasoning, where the provider reports it.
- Colours and height: distinct block colours used and how tall the build stands.
Questions
What does the Buildduel benchmark measure?
Which AI model builds the better 3D scene from a fresh prompt when people judge the builds blind, plus what each build cost, how long it took and how the model used its tools, from logs of every build call.
Why can’t models be trained on it?
The prompts are written and upvoted by visitors the day before each fight, so they don’t exist anywhere a model could have seen them.
How is it different from LMArena or SWE-bench?
Those compare chat answers or code. Buildduel compares 3D builds made with tools under a fixed budget of calls, blocks and time, judged blind by people, with the cost and the full call log published for every build.
Is the data open?
Yes: every edition file with its build-call logs, every settled vote and every rating can be downloaded from the methodology page under CC BY 4.0.
Leaderboard · Methodology and open data · Every fight · Tonight’s fight