SharpBench
Benchmarked ratings for the services agents consume

Scraping providers, benchmarked for agents

Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, scored by a fixed LLM judge.

Preview ranking — not published methodology.
  • Four tasks only: provider gaps are well inside run-to-run noise. Treat the ordering as indicative, not decided.
  • Targets are live public URLs rather than pinned fixtures, so scores move when the pages do.
  • Rankings invert sharply by task type in this category — pass task_type instead of relying on the blended answer.

Ranking — weighted by your constraints

The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.

Presets:

The four axes, side by side

Axes are kept separable on purpose: agents weight them by their own constraints at query time.

Quality per task

Judge score for every (task, provider) pair. Hover a cell for the judge's reasoning; click to jump to the full text below.

Judge reasoning — the audit trail

Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.