Two tools answer the same task, blind, plus no tool at all. The model picks its favorite, and so do you. We keep the tally by task.
Every matchup is blind: the answers show up unlabeled, in random order. Both the model and you pick a winner. We publish the record per task — "beat X 14 of 19 on embedding-search" — with its count. We tally; we don't judge.