304 lines
17 KiB
Markdown
304 lines
17 KiB
Markdown
# AI-assisted building checkpoint review — 2026-08-09
|
|
|
|
## Outcome
|
|
|
|
GeoIntel still runs the proven active building checkpoint. A new 12-epoch GPU
|
|
fine-tune was rejected because it reduced mAP and introduced three detections
|
|
on pure-background validation tiles. A wider checkpoint comparison identified
|
|
the older `reviewedexp6` checkpoint as the strongest non-protected validation
|
|
challenger, but it was not promoted.
|
|
|
|
The production runtime was verified healthy after the review with:
|
|
|
|
- image `geointel-all-in-one:0209167cfd37-wip25e7de62cde9-ai`;
|
|
- active weights SHA-256
|
|
`a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1`;
|
|
- NVIDIA RTX 4080 SUPER inference; and
|
|
- the existing Kempen validation-scope enforcement.
|
|
|
|
## Metric review
|
|
|
|
The checkpoint matrix used only the 36-image non-protected validation split:
|
|
18 positive images and 18 declared pure-background images. The active model
|
|
measured mAP50 `0.345426` and mAP50-95 `0.140719`. The `reviewedexp6`
|
|
challenger measured `0.368177` and `0.155438` respectively.
|
|
|
|
A later exact cross-tile reconstruction established that this 512 px corpus
|
|
uses stride 256. Across the full corpus, 74,113 interior label rows represent
|
|
34,082 unique reconstructed objects; 40,031 rows are overlap repetitions and
|
|
one object can occur four times. No reconstructed object crosses the
|
|
train/validation boundary, but the tile rows are not statistically independent.
|
|
The matrix therefore remains useful only for relative non-protected candidate
|
|
ranking. Its mAP values are not an independent object-level accuracy estimate
|
|
or a basis for narrow confidence claims.
|
|
|
|
At confidence `0.25` and match IoU `0.25`, the active model measured F1
|
|
`0.633058`; the challenger measured `0.642599`. Both produced zero detections
|
|
on the 18 pure-background validation tiles at this threshold. At confidence
|
|
`0.15`, however, the challenger still reproduces the two background detections
|
|
that blocked its July promotion.
|
|
|
|
## Visual review
|
|
|
|
The error sheets were inspected at original resolution. Green denotes a
|
|
matched prediction, red a false positive, magenta a false negative and yellow
|
|
the reference footprint. The challenger reduces missed-building pressure,
|
|
particularly in Westerlo, but dense Turnhout tiles still show extensive
|
|
low-IoU disagreement.
|
|
|
|
A major part of that disagreement is not safely resolved by additional epochs:
|
|
the aerial image shows a roof while GRB GBG represents the building footprint
|
|
at ground level. Perspective, roof overhang and acquisition-date differences
|
|
can therefore shift visible roofs relative to the authoritative footprint.
|
|
These cases remain review candidates; they were not relabelled automatically.
|
|
|
|
This is explicitly an AI-assisted inspection, not a signed human review. No
|
|
human-review gate or label-acceptance record was fabricated.
|
|
|
|
## Fail-closed production decision
|
|
|
|
The existing seven-zone API benchmark was started with the challenger at
|
|
confidence `0.25`, but the current provenance gate stopped it before model
|
|
loading. The historical checkpoint predates the required neighbouring
|
|
`.geointel-model.json` runtime manifest and its governed database snapshot.
|
|
That control was not bypassed and no production inference from the unbound
|
|
checkpoint was accepted.
|
|
|
|
The challenger may only be reconsidered after its original training evidence
|
|
is migrated without invented lineage, followed by the same seven positive
|
|
zones and three pure-background controls. Until then, the active checkpoint is
|
|
the only safe production choice.
|
|
|
|
Machine-readable hashes, metrics, paths and the exact decision are in
|
|
`artifacts/evidence/accuracy/model-training/20260809-v68-checkpoint-and-threshold-review.json`.
|
|
|
|
## Provenance migration follow-up
|
|
|
|
The original July training directory was audited after the threshold review.
|
|
It retains exact weights, Ultralytics arguments, result curves, dataset YAML,
|
|
dataset summary and a structural quality audit. This recovers useful facts,
|
|
including seed `0`, deterministic mode, the base-model hash and all dataset
|
|
counts.
|
|
|
|
It does not retain the complete evidence required to construct a current
|
|
production sidecar truthfully. In particular, the exact training commit,
|
|
container digest, dependency/runtime receipt, immutable corpus and label
|
|
release manifests, independent split audit and accepted human-review ledger
|
|
are absent. The current review validator reproduces four failures:
|
|
|
|
- `accepted_human_review_evidence_missing`;
|
|
- `review_audit_manifest_not_immutable`;
|
|
- `review_audit_spatial_leakage_not_ok`; and
|
|
- `review_complete_not_true`.
|
|
|
|
No UUID, upstream checksum, historical runtime or reviewer decision was
|
|
invented. Consequently no runtime sidecar or database source snapshot was
|
|
created, and the challenger remains unavailable to production inference. The
|
|
full machine-readable audit is retained in
|
|
`artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-provenance-migration-audit.json`.
|
|
|
|
## Completed AI-assisted AOI ledger
|
|
|
|
At the user's explicit request, the existing 64-tile contact sheet was
|
|
reviewed again at original resolution and recorded as AI-assisted evidence.
|
|
The sheet represents 28 AOIs and renders 21,611 labels. Twenty-six represented
|
|
training AOIs were classified as eligible for experimental training;
|
|
Grobbendonk and Vosselaar remain non-protected validation only. Lommel-heide
|
|
and Postel-bos were visually retained as pure-empty examples, while six sparse
|
|
contexts remain explicitly distinct from pure background.
|
|
|
|
The ledger identifies `openai-codex` as an AI assistant and sets `human=false`.
|
|
It explicitly cannot satisfy the human-review release gate. The production
|
|
training wrapper reproduced that boundary and no bypass was added. The exact
|
|
ledger is retained in
|
|
`artifacts/evidence/accuracy/model-training/20260809-reviewedexp6-ai-assisted-review-ledger.json`.
|
|
|
|
The renderer was subsequently paginated and the review expanded from the
|
|
64-tile sample to all 252 retained tiles. Four immutable pages now account for
|
|
all 79,192 labels with zero missing images, missing label files, invalid rows
|
|
or low-variance tiles. No exact or IoU>=0.90 duplicate box pair was found. A
|
|
separate containment probe identified 43 potentially nested pairs across 36
|
|
tiles (0.0543% relative to rendered labels); these remain human-adjudication
|
|
candidates and were neither rewritten nor automatically excluded.
|
|
|
|
The 36 flagged tiles were then rendered on a separate high-resolution sheet.
|
|
AI-assisted inspection found no systematic duplicate-label pattern: the
|
|
relationships predominantly represent adjacent or complex building components
|
|
in dense GRB contexts. All tiles remain available for experimental analysis,
|
|
while the exact 43 pairs stay visible for human release adjudication.
|
|
|
|
A second focused render now binds the relationship manifest back to the exact
|
|
YOLO label-row indices. It highlights all 86 implicated rows: cyan for exact
|
|
duplicates, red for near duplicates and magenta for possible nesting. The sheet
|
|
contains only magenta relationship marks, confirming visually and
|
|
machine-readably that there are no exact or high-IoU duplicates in this set.
|
|
The nesting cases are geographically dispersed and commonly combine a larger
|
|
GRB object envelope with a smaller component. They are not deterministic
|
|
rewrite candidates. Excluding the 36 complete tiles would also discard 17,167
|
|
unflagged labels, so no automatic tile exclusion or label mutation was made.
|
|
The checksum-bound highlighted sheet and summary are recorded in the AI review
|
|
ledger.
|
|
|
|
## Row-level geometry and overlap follow-up
|
|
|
|
The complete 79,192-label corpus was additionally audited at label-row level.
|
|
The immutable manifest identifies 6,657 unique rows with at least one geometric
|
|
training-risk signal: 1,812 rows have a width or height below 4 pixels, 26 have
|
|
aspect ratio at least 8 and 5,079 touch a tile edge. The categories were
|
|
rendered separately. All 23 extreme-aspect tiles and all 198 small-dimension
|
|
tiles were inspected, plus a 64-tile edge sample.
|
|
|
|
The extreme-aspect group predominantly shows plausible elongated sheds and
|
|
building components. Sixteen of its 26 rows are also below 4 pixels, allowing
|
|
the resolution floor to address most ambiguous extremes without deleting valid
|
|
long structures. The small-dimension sheets contain many visually marginal
|
|
miniature targets. The next immutable experimental corpus should therefore use
|
|
the exporter's normal minimum dimension of at least 4 pixels instead of this
|
|
legacy corpus's 3-pixel override. This finding does not justify mutating the
|
|
frozen corpus or retroactively changing its checkpoint.
|
|
|
|
The edge sample shows expected clipped buildings under the 0.35 minimum-visible
|
|
policy. Raising that value may reduce partial-target pressure, but it must be a
|
|
separately versioned ablation because removing all 5,079 rows without checking
|
|
the original visible fraction would be unsound.
|
|
|
|
## Derived min-4px experimental corpus
|
|
|
|
The row-level finding was converted into a new immutable derived corpus rather
|
|
than changing the historical dataset. The derivation is checksum-bound to the
|
|
79,192-label source summary and removes only labels whose smallest 512-tile
|
|
dimension is below 4 pixels. It retains all 252 tiles and 77,380 labels; no
|
|
positive tile became empty, and no source file was overwritten.
|
|
|
|
Fresh audits report zero sub-4-pixel labels, zero invalid or missing label
|
|
files, zero exact/high-IoU duplicates and zero reconstructed cross-split
|
|
object groups. Extreme-aspect labels fall from 26 to 10, possible nesting from
|
|
43 to 41 and tile-edge rows from 5,079 to 4,840. The complete four-page visual
|
|
render was inspected again. Pure-background tiles remain visibly empty, sparse
|
|
contexts remain distinct and the ten remaining elongated labels predominantly
|
|
match plausible long structures.
|
|
|
|
The generic quality auditor still reports `needs_attention` because normalized
|
|
median box area `0.000641` is below its conservative `0.001` warning threshold.
|
|
That corresponds to median dimensions near 12 by 13 pixels at source tile
|
|
resolution and is not a failed minimum-dimension gate. The warning remains
|
|
visible; it was not suppressed or relabelled as success.
|
|
|
|
This min-4px version is the preferred experimental successor to the legacy
|
|
min-3px corpus. It explicitly remains ineligible for governed training until
|
|
real human review and all release-contract evidence exist.
|
|
|
|
## Non-overlapping checkpoint re-evaluation
|
|
|
|
The 36-tile checkpoint matrix was traced to a 512 px validation corpus with
|
|
stride 256. It contained nine views each of Turnhout, Westerlo, Postel-bos and
|
|
Arendonk-heide. Visual inspection then exposed that all Arendonk-heide views
|
|
are blank/no-data imagery, not meaningful pure-background observations. The
|
|
historical claim of 18 background images is therefore corrected: nine were
|
|
blank no-data and nine represented Postel under overlap.
|
|
|
|
A training-disabled validation view now covers each 1024 px AOI with four
|
|
non-overlapping 512 px tiles. Four blank Arendonk representatives are excluded
|
|
with explicit reason codes. The resulting set contains eight positive tiles
|
|
from Turnhout/Westerlo, four real Postel background tiles and 2,551 labels.
|
|
Exact reconstruction finds 2,384 unique interior objects, 167 edge labels and
|
|
zero repeated interior objects. The view has empty train directories, a
|
|
checksum-bound `NO_TRAINING.json`, and the training wrapper rejects that marker.
|
|
|
|
On the Tower RTX 4080 SUPER, the active checkpoint measures precision
|
|
`0.542632`, recall `0.435317`, mAP50 `0.342034` and mAP50-95 `0.141316`.
|
|
The reviewedexp6 challenger measures `0.581715`, `0.462300`, `0.368364` and
|
|
`0.155318` respectively. At background confidence 0.15 the active checkpoint
|
|
has zero Postel detections and the challenger has one. At 0.25 both have zero.
|
|
Thus the refined evidence confirms the challenger's relative metric advantage
|
|
and its threshold sensitivity, but still does not authorize promotion.
|
|
|
|
Non-overlap is not overstated as statistical independence: the positive tiles
|
|
remain adjacent and come from only two AOIs, edge objects can remain split, and
|
|
there is only one real background AOI. The production model remains unchanged.
|
|
|
|
## Training/evaluation membership correction
|
|
|
|
The evaluation AOIs were subsequently checked against the exact tile summaries
|
|
of both compared checkpoints. `turnhout` and `westerlo` do not occur in either
|
|
train split. `postel_bos`, however, occurs in both: the active corpus contains
|
|
20 training AOIs and the challenger corpus 26, with Postel included in each.
|
|
The Postel image is real rather than blank/no-data, but it is training-seen.
|
|
|
|
Consequently the positive Turnhout/Westerlo metrics remain a non-protected,
|
|
adjacent-AOI candidate ranking; the Postel detection counts are only
|
|
training-seen sanity/regression observations. They are not independent
|
|
pure-background validation and cannot support release, threshold or
|
|
generalisation claims. The earlier wording about a "real background AOI" must
|
|
be read with this correction.
|
|
|
|
The checkpoint evaluator now accepts exact training summaries and, in governed
|
|
mode, fails before PyTorch import, model loading or GPU inference when any
|
|
evaluation AOI overlaps any supplied train split. The reproduced gate blocked
|
|
on `postel_bos` for both checkpoints. Its immutable machine-readable record is
|
|
`artifacts/evidence/accuracy/model-training/20260809-v70-evaluation-independence-gate.json`.
|
|
|
|
## Full model-lineage correction
|
|
|
|
The preceding correction still considered only the final fine-tune corpus of
|
|
each checkpoint. Exact retained Ultralytics arguments establish a longer
|
|
ancestry: the active checkpoint was initialized from
|
|
`geointel-building-yolov8s-aoi1024expandedminpx4vis035e50`, which was initialized
|
|
from the generic `yolov8s.pt`; the challenger was then initialized from the
|
|
active checkpoint. The copied model assets and retained `best.pt` files match
|
|
byte-for-byte at each building-model stage.
|
|
|
|
The ancestral expanded corpus exposes all three evaluation AOIs: `postel_bos`
|
|
as train, and `turnhout` plus `westerlo` as validation. Therefore none of the
|
|
v69 AOIs is independent of the complete model family. Turnhout/Westerlo can
|
|
still be used as familiar regression diagnostics, but their metrics are not a
|
|
fresh candidate-ranking result and must not support accuracy, uncertainty,
|
|
generalisation or release claims.
|
|
|
|
The gate now requires every ancestral corpus summary and checks every recorded
|
|
split, including validation and calibration. A reproduced Tower run blocked on
|
|
all three AOIs before PyTorch import, model loading or GPU inference. Evidence:
|
|
`artifacts/evidence/accuracy/model-training/20260809-v71-full-lineage-independence-gate.json`.
|
|
The byte-matching parent/output chain is retained separately in
|
|
`artifacts/evidence/accuracy/model-training/20260809-v71-model-lineage-receipt.json`.
|
|
Fresh geographically separated AOIs with complete spatial and lineage checks
|
|
are required for the next meaningful evaluation.
|
|
|
|
## Fresh V72 calibration portfolio
|
|
|
|
Six AOIs were pre-registered before inference and acquired again from the
|
|
governed regional orthophoto and building adapters: two each in Vlaanderen,
|
|
Wallonië and Brussel. All six are absent from every split of the retained model
|
|
lineage. A raster-bound spatial audit resolved all 31 ancestral AOIs and found
|
|
the nearest new AOI at 4,869.63 metres, above the frozen 2 km floor.
|
|
|
|
The initial generic edge-cover tiling would have counted 54 tiles because its
|
|
final-row/final-column coverage nearly duplicated the preceding 512 px cells.
|
|
That view was rejected. The corrected evaluation-only exporter retains exactly
|
|
24 full grid cells, excludes 30 overlapping edge-cover views with reason codes,
|
|
and writes no training tiles. Both the corpus and tile view contain bound
|
|
`NO_TRAINING.json` markers.
|
|
|
|
All 24 tiles and 348 retained labels were visually inspected. Havelange is a
|
|
positive village context despite its pre-registered `field-cal-bg` name.
|
|
Assenede and Laeken contain sparse official building labels and are difficult
|
|
contexts, not pure-empty backgrounds. The evaluator now refuses to accept a
|
|
pure-background prefix when any corresponding label file is non-empty.
|
|
|
|
On the RTX 4080 SUPER the active model scores mAP50 `0.056779` and mAP50-95
|
|
`0.021750`; the challenger scores `0.044705` and `0.016004`. At confidence 0.25
|
|
and match IoU 0.25 their F1 scores are `0.231527` and `0.235885`. These low
|
|
values expose substantial domain-generalisation failure. The active model stays
|
|
in production because the challenger has weaker mAP, only a marginal F1 gain,
|
|
missing historical release provenance and no valid promotion bundle.
|
|
|
|
The active model's aggregate calibration F1 peaks at threshold 0.30
|
|
(`0.250356`), but its Flemish subgroup becomes worse than at lower thresholds.
|
|
No threshold was changed because a global average may not mask that subgroup
|
|
regression. PICC and UrbIS remain diagnostic-only while their regional semantic
|
|
harmonisation contracts are pending; no national release claim is made.
|
|
|
|
Complete evidence is retained under
|
|
`artifacts/evidence/accuracy/model-training/20260809-v72-fresh-calibration/`
|
|
and summarized by `20260809-v72-fresh-calibration-summary.json`.
|