how the record is kept

Methodology

Blind pairs. Same prompt. You pick. We keep score. That is the whole method, and none of it is a scientific benchmark.

01 / input

Single-attempt prompt

Every model gets the same prompt and one attempt. No retries, no edits. Change a word of the prompt and it is a new version; pairs never cross versions.

Current version: v0.

02 / viewports

Tracks stay separate

Desktop (1440×900) and mobile (390×844) are separate contests with separate scores. A desktop vote never touches mobile.

03 / scoring

Preference Elo

One Elo rating per model per track. A newly published run starts fresh at 1500.

Start
1500
K
24
Scale
400
Tie
0.5 / 0.5
Skip
no rating change

A pick scores 1 or 0, a tie 0.5, a skip nothing. Rows are provisional until 20 comparisons.

Why not win rate? Because 2-0 would outrank 80-20.

04 / identity

Who may vote

Anyone. No account needed. A username keeps your picks; no email, ever.

An account is not a person. Guests and accounts count the same (K = 24); we record which is which and make no uniqueness claim.

05 / limits

Preference is not correctness

The board ranks what people liked. Whether a page actually works is a separate reviewer label and never enters the score. Voters are self-selected. This measures taste, not intelligence.

06 / record

What you get

Picks are saved to your account. An account is not proof of a unique person. Nothing here is a promise of a reward.

Signing in keeps your picks. That is the only promise.