Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, scored by a fixed LLM judge.
The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.
One night cannot tell you whether a gap is real. Each line is a provider's composite at the weights above, recomputed for every night still in the served window.
Axes are kept separable on purpose: agents weight them by their own constraints at query time.
Judge score for every (task, provider) pair. Hover a cell for the judge's reasoning; click to jump to the full text below.
Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.