2026-07-02 · Surfil blog
How we benchmark honestly
Our methodology, false-positive/negative rates, and dates.
Most “X% better” claims in this space are untestable. Here's how Bench works instead.
We never publish a fixed injection-detection rate or a universal savings percentage. Every benchmark runs on your own repository, is dated, and produces a signed receipt you can re-verify. We report fixes-per-task, cost-per-fix, and a quality score - and we show the confidence interval.
False positives and false negatives are published alongside every run. If a model regresses next week, your receipt from this week still verifies - that's the point of signing the measured fact rather than a marketing number.