Leaderboard and a signed savings receipt.
Teams want to know if their agent setup is actually efficient, and they want proof they can hand to someone else - not a vibe.
A receipt, not a vibe
Bench compares your measured savings against your own history and issues a signed, offline-verifiable receipt for the result - the same signature model used across every paid Surfil output.
Proof, not promises
Bench produces one metric type on the shared spine: signed benchmark receipts. Every paid output is signed (Ed25519) and verifiable offline with no account. Zero-trace: your source never leaves the device.
✓ VALID (offline · epoch 7)
Bench, as you'd actually see it
How Bench does it
Where Bench earns its place
Model choice per repo
Find which model is cheapest-per-fix on your codebase, not in a vendor benchmark.
Proof for the budget owner
Hand finance a signed savings receipt instead of a screenshot they can't check.
Tracking drift
See whether your setup is getting more or less efficient over time, measured.
Questions developers ask first
Is the leaderboard against other companies?
No - it runs against your own history only. We never fabricate peer comparisons.
What exactly is signed?
The measured facts of the run - token counts, dates, deltas - signed with Ed25519.
Can someone else verify my receipt?
Yes. surfil verify works offline for anyone holding the receipt - that's the point.
One spine - products compound
Add Bench to your agents.
Core installs with Starter; Bench plugs into the same interception point - no second layer, no new setup.