calibrate a confidence threshold from one inference pass
Threshold calibration ran the model over every tile once per threshold — three GPU passes to compare 0.50, 0.25 and 0.15 on a hundred-tile raster. The answer is already in a single run at the lowest value: detections above a higher cut are a subset of it, and duplicate suppression walks candidates in descending confidence, so a lower-confidence box can never displace a higher-confidence one. The kept set above any cut is identical whichever threshold the run used, which is what makes one pass sufficient rather than merely cheaper. QA now takes calibration_thresholds and reads each operating point off the same precision/recall walk it already performs, marking the F1-optimal cut. The lab runs inference once and fills its table from the sweep. The contract test asserted the per-threshold loop by name, pinning the waste it was meant to describe. It now states what calibration owes an operator: a row per requested threshold, from one run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1913,6 +1913,18 @@ transaction timestamp, so ordering by `created_at` left the assignment — and
|
||||
therefore the score and the false-positive evidence shown to a reviewer —
|
||||
undefined between identical runs.
|
||||
|
||||
`calibration_thresholds` asks for named confidence cuts alongside the run's own
|
||||
operating point. They are read off the same matching pass, so a sweep costs no
|
||||
extra inference at all. Each row reports the tally, precision, recall and F1 at
|
||||
that cut, and `best_f1_in_sweep` marks the F1-optimal one.
|
||||
|
||||
This replaces re-running the model once per threshold. Detections above a
|
||||
higher cut are a subset of a lower-cut run, and duplicate suppression walks
|
||||
candidates in descending confidence, so a lower-confidence box can never
|
||||
displace a higher-confidence one: the kept set above any cut is identical
|
||||
whichever threshold the run itself used. Three thresholds therefore cost one
|
||||
GPU pass rather than three, and produce the same numbers.
|
||||
|
||||
The response also returns `precision_recall_curve`: precision, recall and F1 at
|
||||
every confidence value present in the run, plus `average_precision`, `best_f1`
|
||||
and `best_f1_threshold`. A single F1 describes one operating point and cannot
|
||||
|
||||
Reference in New Issue
Block a user