calibrate a confidence threshold from one inference pass

Threshold calibration ran the model over every tile once per threshold — three
GPU passes to compare 0.50, 0.25 and 0.15 on a hundred-tile raster. The answer
is already in a single run at the lowest value: detections above a higher cut
are a subset of it, and duplicate suppression walks candidates in descending
confidence, so a lower-confidence box can never displace a higher-confidence
one. The kept set above any cut is identical whichever threshold the run used,
which is what makes one pass sufficient rather than merely cheaper.

QA now takes calibration_thresholds and reads each operating point off the same
precision/recall walk it already performs, marking the F1-optimal cut. The lab
runs inference once and fills its table from the sweep.

The contract test asserted the per-threshold loop by name, pinning the waste it
was meant to describe. It now states what calibration owes an operator: a row
per requested threshold, from one run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Jens
2026-08-22 19:37:19 +02:00
co-authored by Claude Opus 5
parent 8a26007281
commit 2cd2c49389
11 changed files with 303 additions and 52 deletions
+12
View File
@@ -1913,6 +1913,18 @@ transaction timestamp, so ordering by `created_at` left the assignment — and
therefore the score and the false-positive evidence shown to a reviewer —
undefined between identical runs.
`calibration_thresholds` asks for named confidence cuts alongside the run's own
operating point. They are read off the same matching pass, so a sweep costs no
extra inference at all. Each row reports the tally, precision, recall and F1 at
that cut, and `best_f1_in_sweep` marks the F1-optimal one.
This replaces re-running the model once per threshold. Detections above a
higher cut are a subset of a lower-cut run, and duplicate suppression walks
candidates in descending confidence, so a lower-confidence box can never
displace a higher-confidence one: the kept set above any cut is identical
whichever threshold the run itself used. Three thresholds therefore cost one
GPU pass rather than three, and produce the same numbers.
The response also returns `precision_recall_curve`: precision, recall and F1 at
every confidence value present in the run, plus `average_precision`, `best_f1`
and `best_f1_threshold`. A single F1 describes one operating point and cannot