6.0 KiB
AI-assisted building checkpoint review — 2026-08-09
Outcome
GeoIntel still runs the proven active building checkpoint. A new 12-epoch GPU
fine-tune was rejected because it reduced mAP and introduced three detections
on pure-background validation tiles. A wider checkpoint comparison identified
the older reviewedexp6 checkpoint as the strongest non-protected validation
challenger, but it was not promoted.
The production runtime was verified healthy after the review with:
- image
geointel-all-in-one:0209167cfd37-wip25e7de62cde9-ai; - active weights SHA-256
a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1; - NVIDIA RTX 4080 SUPER inference; and
- the existing Kempen validation-scope enforcement.
Metric review
The checkpoint matrix used only the 36-image non-protected validation split:
18 positive images and 18 declared pure-background images. The active model
measured mAP50 0.345426 and mAP50-95 0.140719. The reviewedexp6
challenger measured 0.368177 and 0.155438 respectively.
At confidence 0.25 and match IoU 0.25, the active model measured F1
0.633058; the challenger measured 0.642599. Both produced zero detections
on the 18 pure-background validation tiles at this threshold. At confidence
0.15, however, the challenger still reproduces the two background detections
that blocked its July promotion.
Visual review
The error sheets were inspected at original resolution. Green denotes a matched prediction, red a false positive, magenta a false negative and yellow the reference footprint. The challenger reduces missed-building pressure, particularly in Westerlo, but dense Turnhout tiles still show extensive low-IoU disagreement.
A major part of that disagreement is not safely resolved by additional epochs: the aerial image shows a roof while GRB GBG represents the building footprint at ground level. Perspective, roof overhang and acquisition-date differences can therefore shift visible roofs relative to the authoritative footprint. These cases remain review candidates; they were not relabelled automatically.
This is explicitly an AI-assisted inspection, not a signed human review. No human-review gate or label-acceptance record was fabricated.
Fail-closed production decision
The existing seven-zone API benchmark was started with the challenger at
confidence 0.25, but the current provenance gate stopped it before model
loading. The historical checkpoint predates the required neighbouring
.geointel-model.json runtime manifest and its governed database snapshot.
That control was not bypassed and no production inference from the unbound
checkpoint was accepted.
The challenger may only be reconsidered after its original training evidence is migrated without invented lineage, followed by the same seven positive zones and three pure-background controls. Until then, the active checkpoint is the only safe production choice.
Machine-readable hashes, metrics, paths and the exact decision are in
artifacts/evidence/accuracy/model-training/20260809-v68-checkpoint-and-threshold-review.json.
Provenance migration follow-up
The original July training directory was audited after the threshold review.
It retains exact weights, Ultralytics arguments, result curves, dataset YAML,
dataset summary and a structural quality audit. This recovers useful facts,
including seed 0, deterministic mode, the base-model hash and all dataset
counts.
It does not retain the complete evidence required to construct a current production sidecar truthfully. In particular, the exact training commit, container digest, dependency/runtime receipt, immutable corpus and label release manifests, independent split audit and accepted human-review ledger are absent. The current review validator reproduces four failures:
accepted_human_review_evidence_missing;review_audit_manifest_not_immutable;review_audit_spatial_leakage_not_ok; andreview_complete_not_true.
No UUID, upstream checksum, historical runtime or reviewer decision was
invented. Consequently no runtime sidecar or database source snapshot was
created, and the challenger remains unavailable to production inference. The
full machine-readable audit is retained in
artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-provenance-migration-audit.json.
Completed AI-assisted AOI ledger
At the user's explicit request, the existing 64-tile contact sheet was reviewed again at original resolution and recorded as AI-assisted evidence. The sheet represents 28 AOIs and renders 21,611 labels. Twenty-six represented training AOIs were classified as eligible for experimental training; Grobbendonk and Vosselaar remain non-protected validation only. Lommel-heide and Postel-bos were visually retained as pure-empty examples, while six sparse contexts remain explicitly distinct from pure background.
The ledger identifies openai-codex as an AI assistant and sets human=false.
It explicitly cannot satisfy the human-review release gate. The production
training wrapper reproduced that boundary and no bypass was added. The exact
ledger is retained in
artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-ai-assisted-review-ledger.json.
The renderer was subsequently paginated and the review expanded from the 64-tile sample to all 252 retained tiles. Four immutable pages now account for all 79,192 labels with zero missing images, missing label files, invalid rows or low-variance tiles. No exact or IoU>=0.90 duplicate box pair was found. A separate containment probe identified 43 potentially nested pairs across 36 tiles (0.0543% relative to rendered labels); these remain human-adjudication candidates and were neither rewritten nor automatically excluded.
The 36 flagged tiles were then rendered on a separate high-resolution sheet. AI-assisted inspection found no systematic duplicate-label pattern: the relationships predominantly represent adjacent or complex building components in dense GRB contexts. All tiles remain available for experimental analysis, while the exact 43 pairs stay visible for human release adjudication.