how the record is kept
Methodology
Blind pairs. Same prompt. You pick. We keep score. That is the whole method, and none of it is a scientific benchmark.
01 / input
Single-attempt prompt
Every model gets the same prompt and one attempt. No retries, no edits. Change a word of the prompt and it is a new version; pairs never cross versions.
Current version: v0.
02 / viewports
Tracks stay separate
Desktop (1440×900) and mobile (390×844) are separate contests with separate scores. A desktop vote never touches mobile.
03 / scoring
Preference Elo
One Elo rating per model per track. A newly published run starts fresh at 1500.
- Start
- 1500
- K
- 24
- Scale
- 400
- Tie
- 0.5 / 0.5
- Skip
- no rating change
A pick scores 1 or 0, a tie 0.5, a skip nothing. Rows are provisional until 20 comparisons.
Why not win rate? Because 2-0 would outrank 80-20.
04 / identity
Who may vote
Anyone. No account needed. A username keeps your picks; no email, ever.
An account is not a person. Guests and accounts count the same (K = 24); we record which is which and make no uniqueness claim.
05 / limits
Preference is not correctness
The board ranks what people liked. Whether a page actually works is a separate reviewer label and never enters the score. Voters are self-selected. This measures taste, not intelligence.
06 / record
What you get
Picks are saved to your account. An account is not proof of a unique person. Nothing here is a promise of a reward.
Signing in keeps your picks. That is the only promise.