Files
geointel/docs/BELGIUM_BUILDING_TRAINING_LOOP.md
T
Jens faeb58ef6d
GeoIntel release gates / Compile, test, contracts and builds (push) Successful in 1m49s
GeoIntel release gates / Python and npm vulnerability policy (push) Successful in 21s
GeoIntel release gates / Production AI image, SBOM and container scan (push) Successful in 5m39s
GeoIntel release gates / Deploy exact gated revision to Unraid (push) Failing after 58m43s
Initial public release
2026-08-31 21:56:53 +02:00

11 KiB

Belgian building detector: closed training loop

Meaning of complete

100% trained means that every frozen release gate below passes. It does not mean a fabricated 100% precision, recall or mAP score. A model that memorises a small test set is not complete.

The loop is:

  1. provision new, spatially independent AOIs from governed official services;
  2. freeze imagery, labels, metadata and checksums into a new corpus version;
  3. reject invalid, duplicate and sub-resolution labels and run spatial-leakage checks;
  4. train only on the train split with CUDA on the Tower NVIDIA GPU;
  5. use validation for early stopping and calibration only for threshold choice;
  6. evaluate the fixed threshold once on regional test and background-test data;
  7. attribute false positives and false negatives to a region, AOI and context;
  8. add new training-only examples for the observed failure modes and repeat;
  9. stop only when every objective gate passes; request human review afterward.

Protected calibration, test and background-test AOIs never become training data. A new iteration adds independent training AOIs instead.

Frozen release gates

Area Gate
Runtime CUDA required; NVIDIA device visible; no CPU fallback
Corpus Immutable manifest and artifacts with SHA-256 evidence
Geographic composition Each land region has at least 15 train, 2 val, 3 calibration, 3 test and 2 background-test AOIs
Contexts Dense urban, suburban, rural, industrial and difficult negative contexts represented
Leakage No intersecting AOIs across protected split roles
Label integrity No malformed rows; sub-resolution labels explicitly rejected
Temporal truth Unknown per-pixel dates remain unknown; acquisition dates may not masquerade as observation dates
Threshold selection Calibration set only; maximise the worst regional F1 before aggregate F1
Test aggregate F1 at least 0.55 at the frozen footprint/detection match IoU 0.25 contract
Dense-tile capacity Retain up to 1000 detections per tile; the library default of 300 is below observed Belgian urban label density
Regional test Every region: F1 at least 0.45, precision at least 0.50 and recall at least 0.40
Pure background Zero detections on every pure-empty tile at the selected threshold
Production Exact candidate checksum and fail-closed promotion report required
Final review Human accepts every queued AOI contact sheet after all automated gates pass

These are minimum release gates, not performance targets. Raising a confidence threshold until detections disappear cannot pass because regional recall is a simultaneous gate.

Current gap inventory

The active v31 rotated-holdout corpus contains 142 independent AOIs and 26,041 accepted source building labels. Its automated corpus audit has no failures: spatial leakage is zero, temporal identity is explicit and 272 sub-resolution plus 393 post-imagery labels are rejected rather than silently trained. The tile-quality audit records 51,244 visible training instances, zero invalid or missing label files, zero low-variance positive tiles and 214 negative tiles. It covers coastal, ribbon-development, farmland, park, forest, industrial, rail, quarry and additional urban contexts. Its remaining known gaps are:

  • dated 2025 imagery is used for Flanders and Brussels and the dated 2024 SPW campaign for Wallonia. Exact flight days remain a later metadata refinement, but all corpus relations are now measured periods rather than download dates;
  • GRB/PICC/UrbIS describe ground footprints, whereas visible roofs can remain displaced. The existing detector QA contract therefore uses IoU 0.25; the threshold is frozen and cannot be relaxed per candidate;
  • the active v38 loop still has to prove the frozen regional calibration gates; iteration 2 improved aggregate calibration F1 to 0.573 but remained below the Flemish F1/precision/recall gates and the Walloon precision gate;
  • Oostende coastal fabric is the dominant observed Flemish calibration failure. Independent train-only coastal positives and port/dunes negatives are therefore selected by the failure-driven sampler without opening or copying protected calibration data;
  • sparse hard contexts pass the empty-image test more easily than dense urban recall, so both gates must remain independent;
  • building boxes are a valid first detector contract, but footprint-perfect geometry ultimately requires a separately validated segmentation model.

The active production model remains unchanged while any gate fails.

The v31 expansion added eighteen independent train-only AOIs after the v30 rotated-holdout calibration isolated the remaining domain gaps. Nine Flemish AOIs cover industrial roofs, ribbon development and coastal urban fabric. Nine Walloon AOIs add dense urban, rural-town, regional-architecture and difficult industrial/rail/quarry negatives. Existing validation, calibration, test and background-test AOIs remain frozen. The provisioner can load a short-lived operator session from a mode-0600 token file, allowing governed acquisition on an authenticated runtime without exposing credentials in process arguments.

Reproducible evidence

  • corpus assembler: scripts/assemble_belgium_building_corpus.py;
  • corpus auditor: scripts/audit_belgium_building_corpus.py;
  • tile exporter/auditor/contact sheets: the operator_yolo scripts;
  • per-AOI evaluator: scripts/evaluate_belgium_building_candidate.py;
  • calibration-only selection and release gates: scripts/assess_belgium_building_training_iteration.py.
  • checkpointed CUDA orchestration: scripts/run_belgium_building_training_loop.py.

Every failed assessment returns continue_training_loop. Only a report with training_complete may proceed to final human review and guarded activation. Optimizer, initial learning rate, image size and geometric augmentation are explicit loop inputs. This permits a conservative aerial-imagery finetune (for example AdamW with mosaic disabled) without changing calibration, test or release gates.

After a failed assessment, scripts/build_failure_driven_yolo_sampling.py creates a checksummed, train-only sampling manifest. Positive tiles from regions that fail F1 or recall are repeated, while true negative train tiles are repeated when a regional precision gate or the pure-background gate fails. Calibration, test, background-test and validation AOIs are excluded by their frozen corpus split; the generated evidence records that no protected sample entered training. Per-AOI calibration evidence also identifies failed region/context pairs. Train-only AOIs with the same governed context receive a stronger repeat factor than the remaining failed region, so correction rounds target distinct failure modes without copying a protected AOI into training. The sampling evidence records both context sets and repeat factors. When no matching train context exists, regional sampling remains active and the missing context becomes a concrete input for the next immutable corpus expansion. Failure weighting may not let one region exceed 65% of the sampled entries. The deterministic cap removes only repeated entries and retains every unique train tile at least once; manifests record pre-cap counts, final counts and the number of dropped repeats. This keeps a weak region prominent without turning the national detector into a single-region expert. Precision correction expands failed semantic contexts into related negative families: coastal urban failures target port/dunes negatives, industrial failures target industrial/rail/port negatives, and ribbon/rural/regional architecture failures target their governed farmland, forest or quarry counterparts. These diagnostic negative repeats are ordered ahead of generic repeats so the regional cap cannot discard them first. The checkpointed orchestrator invokes this builder after every rejected iteration, stores its checksum in training-loop-state.json, and uses the resulting dataset YAML for the next checkpoint. A restart resumes both the candidate weights and that exact failure-driven training input. An already completed out-of-band checkpoint enters the same contract with --evaluate-initial-model: the first iteration skips fitting, copies and hashes the checkpoint, and begins at calibration. A rejection then follows the identical failure-driven CUDA path and cannot open protected test evidence early. Tower's v37 supervisor binds the completed results.png artifact to a versioned JSON argv list. The handoff starts the orchestrator detached exactly once; shell strings are not accepted. Subsequent iterations retain the frozen 180-degree aerial rotation, vertical/horizontal flip, scale and translation parameters rather than silently reverting to generic augmentation defaults. Before fitting a state-pending iteration, the orchestrator checks its canonical run directory for weights/last.pt. If present, it resumes that exact CUDA checkpoint instead of restarting from the prior candidate. A host-side parent supervisor monitors the exact loop-process marker and persisted loop state; on container recreation it restores the current orchestrator script and launches the versioned JSON argv command. Invalid state and bounded launch exhaustion fail closed. The orchestrator refuses to start unless every automated frozen-dataset gate passes and the corpus contains zero blank/low-variance positive tiles. The separate train tile-quality audit is a required checksummed loop input; missing invalid-label, missing-label-file or low-variance evidence fails closed rather than being interpreted as zero. The audit status may remain needs_human_review while training and objective evaluation continue: final human sign-off is deliberately the last gate and can never be interpreted as model promotion approval in advance. For dated imagery, GRB BEGINDATUM and PICC DATE_CREAT are compared with the end of the imagery period. A feature created afterward is retained in the audit but excluded from training as created_after_imagery_period. UrbIS does not expose an equivalent feature creation field in this acquisition contract, so its remaining temporal relation stays an explicit sample-level limitation.

An opt-in visible-roof experiment can dissolve source footprints that truly touch or overlap; separated footprints are never bridged. The audit retains every contributing native feature identifier and reports both source-feature and visible-instance counts. This mode is not the default: the Belgium v8 experiment showed that unconditional touching-footprint dissolve can merge whole urban blocks and therefore must pass the same independent gates before it can replace native instances.

The compact-roof variant therefore merges a connected group only when it has at most 12 source footprints and fills at least 55% of its axis-aligned envelope. Larger or irregular connected groups retain their native instances and are marked native_instance_complex_touch_group. These fixed criteria prevent administrative row-house chains from becoming one ambiguous detector box while keeping the experiment deterministic and auditable.