Assessed quality, not popularity.
The right provider depends on what you are asking it to do. SharpBench measures the services agents consume on four separable axes, nightly, with the reasoning behind every score.
No provider pays to be listed, ranked, or benchmarked · every score traces to archived raw runs · how this stays neutral
Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, with quality scored by a fixed LLM judge.
The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.
Weights, not scores: each slider sets how much its axis counts relative to the others; the share beside each label is the slider's effective weight, and the four always total 100%.
Axes are kept separable on purpose: agents weight them by their own constraints at query time.
Each point is one provider on this batch: judged quality against measured cost, raw. The dashed steps trace the Pareto frontier — no provider is both better and cheaper than a point on it; anything below the steps is beaten on both axes at once. The sliders never move this chart.
Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.