Files
geointel/docs/reviews/2026-08-09-ai-assisted-checkpoint-review.md
T
Jens b068a5e065
GeoIntel release gates / Compile, test, contracts and builds (push) Canceled after 0s
GeoIntel release gates / Python and npm vulnerability policy (push) Canceled after 0s
GeoIntel release gates / GIS image, SBOM and container scan (push) Canceled after 0s
evaluate models on fresh regional calibration AOIs
2026-08-09 22:36:27 +02:00

17 KiB

AI-assisted building checkpoint review — 2026-08-09

Outcome

GeoIntel still runs the proven active building checkpoint. A new 12-epoch GPU fine-tune was rejected because it reduced mAP and introduced three detections on pure-background validation tiles. A wider checkpoint comparison identified the older reviewedexp6 checkpoint as the strongest non-protected validation challenger, but it was not promoted.

The production runtime was verified healthy after the review with:

  • image geointel-all-in-one:0209167cfd37-wip25e7de62cde9-ai;
  • active weights SHA-256 a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1;
  • NVIDIA RTX 4080 SUPER inference; and
  • the existing Kempen validation-scope enforcement.

Metric review

The checkpoint matrix used only the 36-image non-protected validation split: 18 positive images and 18 declared pure-background images. The active model measured mAP50 0.345426 and mAP50-95 0.140719. The reviewedexp6 challenger measured 0.368177 and 0.155438 respectively.

A later exact cross-tile reconstruction established that this 512 px corpus uses stride 256. Across the full corpus, 74,113 interior label rows represent 34,082 unique reconstructed objects; 40,031 rows are overlap repetitions and one object can occur four times. No reconstructed object crosses the train/validation boundary, but the tile rows are not statistically independent. The matrix therefore remains useful only for relative non-protected candidate ranking. Its mAP values are not an independent object-level accuracy estimate or a basis for narrow confidence claims.

At confidence 0.25 and match IoU 0.25, the active model measured F1 0.633058; the challenger measured 0.642599. Both produced zero detections on the 18 pure-background validation tiles at this threshold. At confidence 0.15, however, the challenger still reproduces the two background detections that blocked its July promotion.

Visual review

The error sheets were inspected at original resolution. Green denotes a matched prediction, red a false positive, magenta a false negative and yellow the reference footprint. The challenger reduces missed-building pressure, particularly in Westerlo, but dense Turnhout tiles still show extensive low-IoU disagreement.

A major part of that disagreement is not safely resolved by additional epochs: the aerial image shows a roof while GRB GBG represents the building footprint at ground level. Perspective, roof overhang and acquisition-date differences can therefore shift visible roofs relative to the authoritative footprint. These cases remain review candidates; they were not relabelled automatically.

This is explicitly an AI-assisted inspection, not a signed human review. No human-review gate or label-acceptance record was fabricated.

Fail-closed production decision

The existing seven-zone API benchmark was started with the challenger at confidence 0.25, but the current provenance gate stopped it before model loading. The historical checkpoint predates the required neighbouring .geointel-model.json runtime manifest and its governed database snapshot. That control was not bypassed and no production inference from the unbound checkpoint was accepted.

The challenger may only be reconsidered after its original training evidence is migrated without invented lineage, followed by the same seven positive zones and three pure-background controls. Until then, the active checkpoint is the only safe production choice.

Machine-readable hashes, metrics, paths and the exact decision are in artifacts/evidence/accuracy/model-training/20260809-v68-checkpoint-and-threshold-review.json.

Provenance migration follow-up

The original July training directory was audited after the threshold review. It retains exact weights, Ultralytics arguments, result curves, dataset YAML, dataset summary and a structural quality audit. This recovers useful facts, including seed 0, deterministic mode, the base-model hash and all dataset counts.

It does not retain the complete evidence required to construct a current production sidecar truthfully. In particular, the exact training commit, container digest, dependency/runtime receipt, immutable corpus and label release manifests, independent split audit and accepted human-review ledger are absent. The current review validator reproduces four failures:

  • accepted_human_review_evidence_missing;
  • review_audit_manifest_not_immutable;
  • review_audit_spatial_leakage_not_ok; and
  • review_complete_not_true.

No UUID, upstream checksum, historical runtime or reviewer decision was invented. Consequently no runtime sidecar or database source snapshot was created, and the challenger remains unavailable to production inference. The full machine-readable audit is retained in artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-provenance-migration-audit.json.

Completed AI-assisted AOI ledger

At the user's explicit request, the existing 64-tile contact sheet was reviewed again at original resolution and recorded as AI-assisted evidence. The sheet represents 28 AOIs and renders 21,611 labels. Twenty-six represented training AOIs were classified as eligible for experimental training; Grobbendonk and Vosselaar remain non-protected validation only. Lommel-heide and Postel-bos were visually retained as pure-empty examples, while six sparse contexts remain explicitly distinct from pure background.

The ledger identifies openai-codex as an AI assistant and sets human=false. It explicitly cannot satisfy the human-review release gate. The production training wrapper reproduced that boundary and no bypass was added. The exact ledger is retained in artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-ai-assisted-review-ledger.json.

The renderer was subsequently paginated and the review expanded from the 64-tile sample to all 252 retained tiles. Four immutable pages now account for all 79,192 labels with zero missing images, missing label files, invalid rows or low-variance tiles. No exact or IoU>=0.90 duplicate box pair was found. A separate containment probe identified 43 potentially nested pairs across 36 tiles (0.0543% relative to rendered labels); these remain human-adjudication candidates and were neither rewritten nor automatically excluded.

The 36 flagged tiles were then rendered on a separate high-resolution sheet. AI-assisted inspection found no systematic duplicate-label pattern: the relationships predominantly represent adjacent or complex building components in dense GRB contexts. All tiles remain available for experimental analysis, while the exact 43 pairs stay visible for human release adjudication.

A second focused render now binds the relationship manifest back to the exact YOLO label-row indices. It highlights all 86 implicated rows: cyan for exact duplicates, red for near duplicates and magenta for possible nesting. The sheet contains only magenta relationship marks, confirming visually and machine-readably that there are no exact or high-IoU duplicates in this set. The nesting cases are geographically dispersed and commonly combine a larger GRB object envelope with a smaller component. They are not deterministic rewrite candidates. Excluding the 36 complete tiles would also discard 17,167 unflagged labels, so no automatic tile exclusion or label mutation was made. The checksum-bound highlighted sheet and summary are recorded in the AI review ledger.

Row-level geometry and overlap follow-up

The complete 79,192-label corpus was additionally audited at label-row level. The immutable manifest identifies 6,657 unique rows with at least one geometric training-risk signal: 1,812 rows have a width or height below 4 pixels, 26 have aspect ratio at least 8 and 5,079 touch a tile edge. The categories were rendered separately. All 23 extreme-aspect tiles and all 198 small-dimension tiles were inspected, plus a 64-tile edge sample.

The extreme-aspect group predominantly shows plausible elongated sheds and building components. Sixteen of its 26 rows are also below 4 pixels, allowing the resolution floor to address most ambiguous extremes without deleting valid long structures. The small-dimension sheets contain many visually marginal miniature targets. The next immutable experimental corpus should therefore use the exporter's normal minimum dimension of at least 4 pixels instead of this legacy corpus's 3-pixel override. This finding does not justify mutating the frozen corpus or retroactively changing its checkpoint.

The edge sample shows expected clipped buildings under the 0.35 minimum-visible policy. Raising that value may reduce partial-target pressure, but it must be a separately versioned ablation because removing all 5,079 rows without checking the original visible fraction would be unsound.

Derived min-4px experimental corpus

The row-level finding was converted into a new immutable derived corpus rather than changing the historical dataset. The derivation is checksum-bound to the 79,192-label source summary and removes only labels whose smallest 512-tile dimension is below 4 pixels. It retains all 252 tiles and 77,380 labels; no positive tile became empty, and no source file was overwritten.

Fresh audits report zero sub-4-pixel labels, zero invalid or missing label files, zero exact/high-IoU duplicates and zero reconstructed cross-split object groups. Extreme-aspect labels fall from 26 to 10, possible nesting from 43 to 41 and tile-edge rows from 5,079 to 4,840. The complete four-page visual render was inspected again. Pure-background tiles remain visibly empty, sparse contexts remain distinct and the ten remaining elongated labels predominantly match plausible long structures.

The generic quality auditor still reports needs_attention because normalized median box area 0.000641 is below its conservative 0.001 warning threshold. That corresponds to median dimensions near 12 by 13 pixels at source tile resolution and is not a failed minimum-dimension gate. The warning remains visible; it was not suppressed or relabelled as success.

This min-4px version is the preferred experimental successor to the legacy min-3px corpus. It explicitly remains ineligible for governed training until real human review and all release-contract evidence exist.

Non-overlapping checkpoint re-evaluation

The 36-tile checkpoint matrix was traced to a 512 px validation corpus with stride 256. It contained nine views each of Turnhout, Westerlo, Postel-bos and Arendonk-heide. Visual inspection then exposed that all Arendonk-heide views are blank/no-data imagery, not meaningful pure-background observations. The historical claim of 18 background images is therefore corrected: nine were blank no-data and nine represented Postel under overlap.

A training-disabled validation view now covers each 1024 px AOI with four non-overlapping 512 px tiles. Four blank Arendonk representatives are excluded with explicit reason codes. The resulting set contains eight positive tiles from Turnhout/Westerlo, four real Postel background tiles and 2,551 labels. Exact reconstruction finds 2,384 unique interior objects, 167 edge labels and zero repeated interior objects. The view has empty train directories, a checksum-bound NO_TRAINING.json, and the training wrapper rejects that marker.

On the Tower RTX 4080 SUPER, the active checkpoint measures precision 0.542632, recall 0.435317, mAP50 0.342034 and mAP50-95 0.141316. The reviewedexp6 challenger measures 0.581715, 0.462300, 0.368364 and 0.155318 respectively. At background confidence 0.15 the active checkpoint has zero Postel detections and the challenger has one. At 0.25 both have zero. Thus the refined evidence confirms the challenger's relative metric advantage and its threshold sensitivity, but still does not authorize promotion.

Non-overlap is not overstated as statistical independence: the positive tiles remain adjacent and come from only two AOIs, edge objects can remain split, and there is only one real background AOI. The production model remains unchanged.

Training/evaluation membership correction

The evaluation AOIs were subsequently checked against the exact tile summaries of both compared checkpoints. turnhout and westerlo do not occur in either train split. postel_bos, however, occurs in both: the active corpus contains 20 training AOIs and the challenger corpus 26, with Postel included in each. The Postel image is real rather than blank/no-data, but it is training-seen.

Consequently the positive Turnhout/Westerlo metrics remain a non-protected, adjacent-AOI candidate ranking; the Postel detection counts are only training-seen sanity/regression observations. They are not independent pure-background validation and cannot support release, threshold or generalisation claims. The earlier wording about a "real background AOI" must be read with this correction.

The checkpoint evaluator now accepts exact training summaries and, in governed mode, fails before PyTorch import, model loading or GPU inference when any evaluation AOI overlaps any supplied train split. The reproduced gate blocked on postel_bos for both checkpoints. Its immutable machine-readable record is artifacts/evidence/accuracy/model-training/20260809-v70-evaluation-independence-gate.json.

Full model-lineage correction

The preceding correction still considered only the final fine-tune corpus of each checkpoint. Exact retained Ultralytics arguments establish a longer ancestry: the active checkpoint was initialized from geointel-building-yolov8s-aoi1024expandedminpx4vis035e50, which was initialized from the generic yolov8s.pt; the challenger was then initialized from the active checkpoint. The copied model assets and retained best.pt files match byte-for-byte at each building-model stage.

The ancestral expanded corpus exposes all three evaluation AOIs: postel_bos as train, and turnhout plus westerlo as validation. Therefore none of the v69 AOIs is independent of the complete model family. Turnhout/Westerlo can still be used as familiar regression diagnostics, but their metrics are not a fresh candidate-ranking result and must not support accuracy, uncertainty, generalisation or release claims.

The gate now requires every ancestral corpus summary and checks every recorded split, including validation and calibration. A reproduced Tower run blocked on all three AOIs before PyTorch import, model loading or GPU inference. Evidence: artifacts/evidence/accuracy/model-training/20260809-v71-full-lineage-independence-gate.json. The byte-matching parent/output chain is retained separately in artifacts/evidence/accuracy/model-training/20260809-v71-model-lineage-receipt.json. Fresh geographically separated AOIs with complete spatial and lineage checks are required for the next meaningful evaluation.

Fresh V72 calibration portfolio

Six AOIs were pre-registered before inference and acquired again from the governed regional orthophoto and building adapters: two each in Vlaanderen, Wallonië and Brussel. All six are absent from every split of the retained model lineage. A raster-bound spatial audit resolved all 31 ancestral AOIs and found the nearest new AOI at 4,869.63 metres, above the frozen 2 km floor.

The initial generic edge-cover tiling would have counted 54 tiles because its final-row/final-column coverage nearly duplicated the preceding 512 px cells. That view was rejected. The corrected evaluation-only exporter retains exactly 24 full grid cells, excludes 30 overlapping edge-cover views with reason codes, and writes no training tiles. Both the corpus and tile view contain bound NO_TRAINING.json markers.

All 24 tiles and 348 retained labels were visually inspected. Havelange is a positive village context despite its pre-registered field-cal-bg name. Assenede and Laeken contain sparse official building labels and are difficult contexts, not pure-empty backgrounds. The evaluator now refuses to accept a pure-background prefix when any corresponding label file is non-empty.

On the RTX 4080 SUPER the active model scores mAP50 0.056779 and mAP50-95 0.021750; the challenger scores 0.044705 and 0.016004. At confidence 0.25 and match IoU 0.25 their F1 scores are 0.231527 and 0.235885. These low values expose substantial domain-generalisation failure. The active model stays in production because the challenger has weaker mAP, only a marginal F1 gain, missing historical release provenance and no valid promotion bundle.

The active model's aggregate calibration F1 peaks at threshold 0.30 (0.250356), but its Flemish subgroup becomes worse than at lower thresholds. No threshold was changed because a global average may not mask that subgroup regression. PICC and UrbIS remain diagnostic-only while their regional semantic harmonisation contracts are pending; no national release claim is made.

Complete evidence is retained under artifacts/evidence/accuracy/model-training/20260809-v72-fresh-calibration/ and summarized by 20260809-v72-fresh-calibration-summary.json.