Files
geointel/docs/reviews/2026-08-09-ai-assisted-checkpoint-review.md
T
Jens 2de438b9cc
GeoIntel release gates / Compile, test, contracts and builds (push) Canceled after 0s
GeoIntel release gates / Python and npm vulnerability policy (push) Canceled after 0s
GeoIntel release gates / GIS image, SBOM and container scan (push) Canceled after 0s
derive cleaner min-4px YOLO corpus
2026-08-09 19:51:17 +02:00

190 lines
10 KiB
Markdown

# AI-assisted building checkpoint review — 2026-08-09
## Outcome
GeoIntel still runs the proven active building checkpoint. A new 12-epoch GPU
fine-tune was rejected because it reduced mAP and introduced three detections
on pure-background validation tiles. A wider checkpoint comparison identified
the older `reviewedexp6` checkpoint as the strongest non-protected validation
challenger, but it was not promoted.
The production runtime was verified healthy after the review with:
- image `geointel-all-in-one:0209167cfd37-wip25e7de62cde9-ai`;
- active weights SHA-256
`a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1`;
- NVIDIA RTX 4080 SUPER inference; and
- the existing Kempen validation-scope enforcement.
## Metric review
The checkpoint matrix used only the 36-image non-protected validation split:
18 positive images and 18 declared pure-background images. The active model
measured mAP50 `0.345426` and mAP50-95 `0.140719`. The `reviewedexp6`
challenger measured `0.368177` and `0.155438` respectively.
A later exact cross-tile reconstruction established that this 512 px corpus
uses stride 256. Across the full corpus, 74,113 interior label rows represent
34,082 unique reconstructed objects; 40,031 rows are overlap repetitions and
one object can occur four times. No reconstructed object crosses the
train/validation boundary, but the tile rows are not statistically independent.
The matrix therefore remains useful only for relative non-protected candidate
ranking. Its mAP values are not an independent object-level accuracy estimate
or a basis for narrow confidence claims.
At confidence `0.25` and match IoU `0.25`, the active model measured F1
`0.633058`; the challenger measured `0.642599`. Both produced zero detections
on the 18 pure-background validation tiles at this threshold. At confidence
`0.15`, however, the challenger still reproduces the two background detections
that blocked its July promotion.
## Visual review
The error sheets were inspected at original resolution. Green denotes a
matched prediction, red a false positive, magenta a false negative and yellow
the reference footprint. The challenger reduces missed-building pressure,
particularly in Westerlo, but dense Turnhout tiles still show extensive
low-IoU disagreement.
A major part of that disagreement is not safely resolved by additional epochs:
the aerial image shows a roof while GRB GBG represents the building footprint
at ground level. Perspective, roof overhang and acquisition-date differences
can therefore shift visible roofs relative to the authoritative footprint.
These cases remain review candidates; they were not relabelled automatically.
This is explicitly an AI-assisted inspection, not a signed human review. No
human-review gate or label-acceptance record was fabricated.
## Fail-closed production decision
The existing seven-zone API benchmark was started with the challenger at
confidence `0.25`, but the current provenance gate stopped it before model
loading. The historical checkpoint predates the required neighbouring
`.geointel-model.json` runtime manifest and its governed database snapshot.
That control was not bypassed and no production inference from the unbound
checkpoint was accepted.
The challenger may only be reconsidered after its original training evidence
is migrated without invented lineage, followed by the same seven positive
zones and three pure-background controls. Until then, the active checkpoint is
the only safe production choice.
Machine-readable hashes, metrics, paths and the exact decision are in
`artifacts/evidence/accuracy/model-training/20260809-v68-checkpoint-and-threshold-review.json`.
## Provenance migration follow-up
The original July training directory was audited after the threshold review.
It retains exact weights, Ultralytics arguments, result curves, dataset YAML,
dataset summary and a structural quality audit. This recovers useful facts,
including seed `0`, deterministic mode, the base-model hash and all dataset
counts.
It does not retain the complete evidence required to construct a current
production sidecar truthfully. In particular, the exact training commit,
container digest, dependency/runtime receipt, immutable corpus and label
release manifests, independent split audit and accepted human-review ledger
are absent. The current review validator reproduces four failures:
- `accepted_human_review_evidence_missing`;
- `review_audit_manifest_not_immutable`;
- `review_audit_spatial_leakage_not_ok`; and
- `review_complete_not_true`.
No UUID, upstream checksum, historical runtime or reviewer decision was
invented. Consequently no runtime sidecar or database source snapshot was
created, and the challenger remains unavailable to production inference. The
full machine-readable audit is retained in
`artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-provenance-migration-audit.json`.
## Completed AI-assisted AOI ledger
At the user's explicit request, the existing 64-tile contact sheet was
reviewed again at original resolution and recorded as AI-assisted evidence.
The sheet represents 28 AOIs and renders 21,611 labels. Twenty-six represented
training AOIs were classified as eligible for experimental training;
Grobbendonk and Vosselaar remain non-protected validation only. Lommel-heide
and Postel-bos were visually retained as pure-empty examples, while six sparse
contexts remain explicitly distinct from pure background.
The ledger identifies `openai-codex` as an AI assistant and sets `human=false`.
It explicitly cannot satisfy the human-review release gate. The production
training wrapper reproduced that boundary and no bypass was added. The exact
ledger is retained in
`artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-ai-assisted-review-ledger.json`.
The renderer was subsequently paginated and the review expanded from the
64-tile sample to all 252 retained tiles. Four immutable pages now account for
all 79,192 labels with zero missing images, missing label files, invalid rows
or low-variance tiles. No exact or IoU>=0.90 duplicate box pair was found. A
separate containment probe identified 43 potentially nested pairs across 36
tiles (0.0543% relative to rendered labels); these remain human-adjudication
candidates and were neither rewritten nor automatically excluded.
The 36 flagged tiles were then rendered on a separate high-resolution sheet.
AI-assisted inspection found no systematic duplicate-label pattern: the
relationships predominantly represent adjacent or complex building components
in dense GRB contexts. All tiles remain available for experimental analysis,
while the exact 43 pairs stay visible for human release adjudication.
A second focused render now binds the relationship manifest back to the exact
YOLO label-row indices. It highlights all 86 implicated rows: cyan for exact
duplicates, red for near duplicates and magenta for possible nesting. The sheet
contains only magenta relationship marks, confirming visually and
machine-readably that there are no exact or high-IoU duplicates in this set.
The nesting cases are geographically dispersed and commonly combine a larger
GRB object envelope with a smaller component. They are not deterministic
rewrite candidates. Excluding the 36 complete tiles would also discard 17,167
unflagged labels, so no automatic tile exclusion or label mutation was made.
The checksum-bound highlighted sheet and summary are recorded in the AI review
ledger.
## Row-level geometry and overlap follow-up
The complete 79,192-label corpus was additionally audited at label-row level.
The immutable manifest identifies 6,657 unique rows with at least one geometric
training-risk signal: 1,812 rows have a width or height below 4 pixels, 26 have
aspect ratio at least 8 and 5,079 touch a tile edge. The categories were
rendered separately. All 23 extreme-aspect tiles and all 198 small-dimension
tiles were inspected, plus a 64-tile edge sample.
The extreme-aspect group predominantly shows plausible elongated sheds and
building components. Sixteen of its 26 rows are also below 4 pixels, allowing
the resolution floor to address most ambiguous extremes without deleting valid
long structures. The small-dimension sheets contain many visually marginal
miniature targets. The next immutable experimental corpus should therefore use
the exporter's normal minimum dimension of at least 4 pixels instead of this
legacy corpus's 3-pixel override. This finding does not justify mutating the
frozen corpus or retroactively changing its checkpoint.
The edge sample shows expected clipped buildings under the 0.35 minimum-visible
policy. Raising that value may reduce partial-target pressure, but it must be a
separately versioned ablation because removing all 5,079 rows without checking
the original visible fraction would be unsound.
## Derived min-4px experimental corpus
The row-level finding was converted into a new immutable derived corpus rather
than changing the historical dataset. The derivation is checksum-bound to the
79,192-label source summary and removes only labels whose smallest 512-tile
dimension is below 4 pixels. It retains all 252 tiles and 77,380 labels; no
positive tile became empty, and no source file was overwritten.
Fresh audits report zero sub-4-pixel labels, zero invalid or missing label
files, zero exact/high-IoU duplicates and zero reconstructed cross-split
object groups. Extreme-aspect labels fall from 26 to 10, possible nesting from
43 to 41 and tile-edge rows from 5,079 to 4,840. The complete four-page visual
render was inspected again. Pure-background tiles remain visibly empty, sparse
contexts remain distinct and the ten remaining elongated labels predominantly
match plausible long structures.
The generic quality auditor still reports `needs_attention` because normalized
median box area `0.000641` is below its conservative `0.001` warning threshold.
That corresponds to median dimensions near 12 by 13 pixels at source tile
resolution and is not a failed minimum-dimension gate. The warning remains
visible; it was not suppressed or relabelled as success.
This min-4px version is the preferred experimental successor to the legacy
min-3px corpus. It explicitly remains ineligible for governed training until
real human review and all release-contract evidence exist.