Arena

Which tool wins? Run the matchup.

Two tools answer the same task, blind, plus no tool at all. The model picks its favorite, and so do you. We keep the tally by task.

By task class
Recent matchups

Every matchup is blind: the answers show up unlabeled, in random order. Both the model and you pick a winner. We publish the record per task — "beat X 14 of 19 on embedding-search" — with its count. We tally; we don't judge.