SharpBench
Benchmarked ratings for the services agents consume

Web-search providers, benchmarked for agents

Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, scored by a fixed LLM judge.

Reading this ranking
  • Latency is measured from the benchmark host (GitHub Actions runners); the same provider measured ~10x slower from a laptop, so treat latency as comparable between providers, not as an absolute you will reproduce.
  • Reliability has no discriminating power yet — every provider is at 100% until failures accumulate across nights.

Ranking — weighted by your constraints

The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.

Presets:

The four axes, side by side

Axes are kept separable on purpose: agents weight them by their own constraints at query time.

Quality per task

Judge score for every (task, provider) pair. Hover a cell for the judge's reasoning; click to jump to the full text below.

Judge reasoning — the audit trail

Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.