The right provider depends on what you are asking it to do.
SharpBench measures the services agents consume on four separable
axes, nightly, with the reasoning behind every score.
No provider pays to be listed, ranked, or benchmarked ·
every score traces to archived raw runs ·
how this stays neutral
Assessed quality, latency, cost, and reliability — every number computed
from persisted benchmark runs, with quality scored by deterministic rubric checks (no judge model involved).
Reading this rankingTop-pair separation is a live verdict, not a fixed order · Latency covers the full lifecycle an agent pays for · Cost is per second of measured lifetime · Most tasks pass on every provider
Top-pair separation is a live verdict, not a fixed order — where two providers both pass every check, the gap between them is latency and cost alone, small enough that the noise floor measured 2026-08-24 had the pair trading first place; the stability panel and each result's separated_from_next carry the per-pair verdict for the batch you are looking at.
Latency covers the full lifecycle an agent pays for — create, execute, teardown — measured from the benchmark host. Warm-pool or snapshot reuse is each provider's own optimization and shows up as their number.
Cost is per second of measured lifetime — priced from each provider's published per-resource rates at its default sandbox size, except daytona, whose Linux rates are not published: its price is derived from third-party parity reporting and flagged approximate in the per-run pricing snapshot. Sizes are recorded per run but are not identical across providers.
Most tasks pass on every provider — quality separation comes from the minority that probe environment limits (network egress, image contents, process primitives). The per-task heatmap shows exactly which checks separate the field.
Ranking — weighted by your constraints
The composite is recomputed live from the four axis scores. Drag the
sliders (or pick a preset) and watch the ranking change — this is the question a free
popularity leaderboard can't answer.
Weights, not scores: each slider sets how much its
axis counts relative to the others; the share beside each label is the
slider's effective weight, and the four always total 100%.
Presets:
The same providers, ranked per task type
Does the ordering hold?
The four axes, side by side
Axes are kept separable on purpose: agents weight them by their own
constraints at query time.
Quality vs. cost — the trade the blend hides
Each point is one provider on this batch: judged quality against
measured cost, raw. The dashed steps trace the Pareto frontier — no provider is
both better and cheaper than a point on it; anything below the steps is beaten
on both axes at once. The sliders never move this chart.
Quality per task
Judge reasoning — the audit trail
Every quality score links back to a stored judgment with reasoning, and
every judgment links back to a raw run (request, response, latency, pricing
snapshot, timestamp) preserved in that batch's archive.