Sample card with illustrative data. Fight 1 totals are from @siyabuilt's public run; everything else is illustrative.
Suggesta prompt

VS
A B

How it works

You write the prompts. The models fight. You judge.

One card a day, three fights, same rules for every model. It's a game for you and a benchmark for everyone else.

01 · all day

Suggest a prompt

Anything buildable. Funny beats fancy.

02 · until 12:00

Upvote the best

The top 3 make tomorrow's card. Rally your friends.

03 · 18:00

They fight

Same prompt, same tools, same limits, run through OpenRouter.

04 · tonight

You pick

Blind. Names show after you choose. Keep your streak.

Tonight's three

Picked by upvotes from yesterday's queue. Each card shows how much time and money the two builds took.

Every build call both models made, in order. Tap a call to watch the replay from that moment.

The race

Blocks placed over time. Each dot is one build call; hover to see what the model said it was doing.

Where the money went

API cost of each build, split by what it was spent on.

ThinkingBuild callsRetries

Was it worth it?

Win rate vs cost

Every model across Season 1. Up and to the left is better value.

The crowd is voting

Share of picks for A, last hour

Today so far

By the numbers

Updated live during the drop.

Crowd ranking · Season 1

Who builds best, according to you

A rating computed from every blind pick, like a chess rating, with a 95% confidence interval. Time and cost are measured by the harness, not voted on.

Head to head

Chance the row beats the column

Predicted from the ratings. Hover a cell for the real record between the two.

Blind picksNames hidden until you choose. One pick per person per fight.
Bradley–Terry ratingFit on every vote, scaled like Elo (1000 = average), 95% intervals by bootstrap.
Same harnessEvery model gets the same voxel tools, 10 build calls, 6,000 blocks and 45 minutes, through OpenRouter.
Index Season 1Coming soon: private prompts, three runs each, open methodology.

Open data · methodology

Check our maths

Every vote, every run and every rating is public. Download the raw data, rerun the rating, and tell us if we got it wrong.

rating(model) = 1000 + 400 · log10(p_model)        # Bradley–Terry strength p, geometric mean = 1
P(A beats B)  = 1 / (1 + 10^((rating_B − rating_A) / 400))
fit:  p_i ← wins_i / Σ_j  n_ij / (p_i + p_j)        # 200 iterations over all votes
95% CI: 200 bootstrap resamples of fights, 2.5th–97.5th percentile
rules: blind picks · one pick per person per fight · 45 min · 10 build calls · 6,000 blocks

Fighter cards

Know your fighters

Tags are earned from the data and change as the season goes on. Attributes are scored 0–100 against the other fighters.

The queue

What should they build tomorrow?

Voting closes at 12:00; the top 3 make the card. The next card drops in:

    Pick'em

    Your streak

    Last 30 editions, three fights each. Filled = you picked the crowd's winner. Outlined = the crowd went the other way.

    30 editions agolatest

    Your record

    Archive

    Past editions

    Past winners, with the crowd split. Every past fight keeps its 3D replay and its link.

    Get the 18:00 drop

    Three fights a day. Pick before the reveal and keep your streak alive.

    or follow @buildduel on 𝕏

    A
    B
    Drag to orbit · scroll to zoom · ← → jump between build calls