SharpBench

Assessed quality, not popularity.

The right provider depends on what you are asking it to do. SharpBench measures the services agents consume on four separable axes, nightly, with the reasoning behind every score.

No provider pays to be listed, ranked, or benchmarked · every score traces to archived raw runs · how this stays neutral

Web-search providers, benchmarked for agents

Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, with quality scored by a fixed LLM judge.

Reading this rankingLatency is vantage-dependent · Reliability rarely discriminates here · Quality separation comes from the hardest task types
  • Latency is vantage-dependent — measured from the benchmark host (GitHub Actions runners); the same provider measured ~10x slower from a laptop, so treat latency as comparable between providers, not as an absolute you will reproduce.
  • Reliability rarely discriminates here — most providers run clean most nights, so a single night of real errors moves this axis more than quality's steady gaps. Honest provider failures stay in; failures caused by the benchmark's own account state are excluded as published errata, disclosed in the response.
  • Quality separation comes from the hardest task types — multi-hop above all; on plain factual lookups every provider scores near-perfect. The suite is weighted toward the discriminating types, and the per-task heatmap shows exactly where the gaps are.

Ranking — weighted by your constraints

The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.

Weights, not scores: each slider sets how much its axis counts relative to the others; the share beside each label is the slider's effective weight, and the four always total 100%.

Presets:

The four axes, side by side

Axes are kept separable on purpose: agents weight them by their own constraints at query time.

Quality per task

Judge reasoning — the audit trail

Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.