rank model variants on average precision, and say when they are not comparable

The workbench ranks model variants by a stored F1, each measured at that
variant's own confidence threshold. A conservatively calibrated detector then
looks worse than a liberal one without detecting anything differently: the
number says as much about the threshold as about the model. POST
/detection/runs/compare ranks on average precision instead, which describes the
whole ranking a model produced, and keeps each run's own-threshold F1 visible
next to it so the difference between the two readings is auditable.

Comparability comes before the ranking. Runs over different source rasters,
scored against different references, without a proven inference footprint, or
covering a different evaluated population are not alternatives to one another,
and no metric makes them so. The report names which of those applies and still
returns the numbers — they are simply not a ranking.

Each run is scored through the same QA path the workbench uses, so a comparison
and the persisted quality checks cannot drift apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Jens
2026-08-22 19:47:14 +02:00
co-authored by Claude Opus 5
parent 2cd2c49389
commit f6eced1b94
6 changed files with 497 additions and 0 deletions
+4
View File
@@ -56,6 +56,8 @@ from .detection import (
DetectionRunListResponse,
DetectionRunRead,
DetectionRunRequest,
DetectionComparisonRequest,
DetectionComparisonResponse,
DetectionRunResponse,
ModelAssetListResponse,
ModelAssetRead,
@@ -238,6 +240,8 @@ __all__ = [
"DetectionRunListResponse",
"DetectionRunRead",
"DetectionRunRequest",
"DetectionComparisonRequest",
"DetectionComparisonResponse",
"DetectionRunResponse",
"ModelAssetListResponse",
"ModelAssetRead",