The right provider depends on what you are asking it to do.
SharpBench measures the services agents consume on four separable
axes, nightly, with the reasoning behind every score.
No provider pays to be listed, ranked, or benchmarked ·
every score traces to archived raw runs ·
how this stays neutral
Assessed quality, latency, cost, and reliability — every number computed
from persisted benchmark runs, with quality scored by deterministic rubric checks blended with a fixed LLM judge.
Preview ranking — not published methodology.Ordering at the top is not settled · Targets are live public URLs · Each provider runs one fixed, persisted configuration · Latency is vantage-dependent · Rankings invert sharply by task type in this category · The quality instrument changed 2026-08-14
Ordering at the top is not settled — the measured noise floor leaves firecrawl and jina overlapping across the served window, and single-night ranks drift inside that overlap; the stability panel and each result's separated_from_next carry the per-pair verdicts for the batch you are looking at.
Targets are live public URLs — not pinned fixtures, so scores move when the pages do, and the suite deliberately excludes anti-bot-protected commercial sites: measuring who best defeats bot protection is a different product. Where possible the suite pins content anyway (versioned docs paths, RFCs, tag-pinned raw files).
Each provider runs one fixed, persisted configuration — no per-task mode fallback, because silently swapping modes would insert our own extraction into the comparison. A zero can therefore score a configuration, not a ceiling: scrapingbee's markdown-extraction mode returns empty for non-HTML documents, which is why it scores 0.00 on plaintext tasks while the raw fetch succeeds.
Latency is vantage-dependent — measured from the benchmark host (GitHub Actions runners); treat it as comparable between providers, not as an absolute you will reproduce.
Rankings invert sharply by task type in this category — pass task_type instead of relying on the blended answer.
The quality instrument changed 2026-08-14 — deterministic content checks (50%) blended with the LLM judge (50%); scores before and after that date are not directly comparable.
Ranking — weighted by your constraints
The composite is recomputed live from the four axis scores. Drag the
sliders (or pick a preset) and watch the ranking change — this is the question a free
popularity leaderboard can't answer.
Weights, not scores: each slider sets how much its
axis counts relative to the others; the share beside each label is the
slider's effective weight, and the four always total 100%.
Presets:
The same providers, ranked per task type
Does the ordering hold?
The four axes, side by side
Axes are kept separable on purpose: agents weight them by their own
constraints at query time.
Quality vs. cost — the trade the blend hides
Each point is one provider on this batch: judged quality against
measured cost, raw. The dashed steps trace the Pareto frontier — no provider is
both better and cheaper than a point on it; anything below the steps is beaten
on both axes at once. The sliders never move this chart.
Quality per task
Judge reasoning — the audit trail
Every quality score links back to a stored judgment with reasoning, and
every judgment links back to a raw run (request, response, latency, pricing
snapshot, timestamp) preserved in that batch's archive.