Files
geointel/docs/PYTORCH_TRAINING_ROADMAP_BELGIUM.md
T
Jens 7a83df1e7f
GeoIntel release gates / Compile, test, contracts and builds (push) Canceled after 0s
GeoIntel release gates / Python and npm vulnerability policy (push) Canceled after 0s
GeoIntel release gates / GIS image, SBOM and container scan (push) Canceled after 0s
Audit Belgian corpus splits and record CUDA evidence
2026-07-26 22:29:02 +02:00

319 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GeoIntel PyTorch training roadmap — Belgium
Status: active programme
Owner: AI pipeline, with GIS and QA gates
Runtime: Tower NVIDIA GeForce RTX 4080 SUPER, PyTorch CUDA
Companion contract: `docs/PYTORCH_MODEL_PROGRAM.md`
## 1. Programme outcome
GeoIntel has a small portfolio of task-specific image models, each constrained
to proven classes, input imagery, geography and acquisition periods. A model is
called **fully trained** only when its dataset, split isolation, label review,
CUDA training, independent evaluation, calibration, deployment and monitoring
gates all pass. Completing an epoch count is not completion.
The programme does not train neural replacements for official GIS sources.
Road length, water area, land use, terrain, flood depth, bathymetry and
source-to-source change remain deterministic analyses.
## 2. Target portfolio
| Priority | Model contract | Output | Initial geographic claim | Current state |
|---|---|---|---|---|
| P0 | `building-detect-be` | building bounding boxes | Belgium land, per validated region | active model is building-only and Mol/Kempen-scoped |
| P1 | `building-segment-be` | building footprint masks | no claim until mask QA passes | `not_configured` |
| P2 | `solar-panel-detect-be` | photovoltaic installation boxes/polygons | no claim until reviewed labels exist | unvalidated asset; `not_configured` |
Additional classes require a written use case, a distinct label contract and
the same intake gates. They are not appended opportunistically to an existing
model.
## 3. Non-model capabilities
The following stay outside PyTorch unless a later RFC demonstrates a material
benefit over the authoritative source:
- roads, parcels, waterways and building footprints from GRB/PICC/UrbIS;
- land-cover and land-use classes from governed thematic rasters;
- elevation, slope, flood, bathymetry and raster indices;
- vector/raster temporal comparison and QA/QC metrics;
- the grounded assistant, which remains an Ollama language-model integration.
## 4. Definition of fully trained
A model/version is fully trained for one declared scope only when:
1. its immutable dataset manifest identifies every image, label source,
observation date, CRS, licence and checksum;
2. train, validation, calibration and test AOIs are spatially disjoint, including
a buffer larger than the maximum tile context;
3. automated label checks pass and sampled labels have completed human review;
4. at least one reproducible CUDA training run and one reproducibility rerun
produce checksum-bound evidence;
5. evaluation covers Flanders, Wallonia and Brussels separately, plus urban,
suburban, rural, industrial and pure-background contexts where applicable;
6. it beats the incumbent on the frozen portfolio without a material regional
or background regression;
7. confidence threshold, tiling and duplicate suppression are calibrated on a
calibration split, never on the final test split;
8. the model card, limitation text, rollback asset and production preflight are
complete;
9. a post-deployment shadow run confirms the offline results before scope is
widened.
## 5. Dataset topology
### Geographic strata
- Flanders: dense urban, ribbon development, villages, farms, industry and
forest/heath negatives.
- Wallonia: Liège/Charleroi urban fabric, villages, dispersed rural buildings,
industrial valleys, quarries and forest negatives.
- Brussels: dense blocks, detached residential, industrial/logistics sites,
parks and rail contexts.
- Boundary controls: language boundary and cross-provider seams.
- Hold-outs: complete municipalities or non-overlapping large AOIs never used
for training or hyperparameter selection.
### Source bindings
| Region | Imagery | Building reference | Required evidence |
|---|---|---|---|
| Flanders | governed Digitaal Vlaanderen orthophoto | GRB GBG | product edition and acquisition mismatch report |
| Wallonia | governed SPW orthophoto | SPW PICC buildings | WMS/catalog identity, CC BY attribution, campaign and PICC snapshot |
| Brussels | governed Paradigm/UrbIS orthophoto | UrbIS buildings | WMS layer identity, licence, edition and vector snapshot |
### Split contract
- `train`: model fitting only;
- `val`: epoch/model selection only;
- `calibration`: threshold, tile size, overlap and NMS tuning only;
- `test`: opened once per candidate family for the promotion decision;
- `background-test`: pure-empty and difficult contextual negatives;
- no source image, overlapping tile, building instance or buffered AOI may cross
split boundaries.
## 6. Work packages
### WP0 — Governance and frozen contracts
- [x] distinguish PyTorch tasks from deterministic GIS functions;
- [x] require CUDA in production training and record runtime evidence;
- [x] parameterize single-class export by class/source/layer;
- [ ] assign stable model IDs, semantic versions and model-card schema;
- [ ] replace weak historical promotion defaults with task-specific approved
gates while preserving old reports as historical evidence.
Exit: schema tests pass and no runtime can advertise an unproven class/scope.
### WP1 — National imagery acquisition
- [x] verify official SPW and Brussels imagery services and usage conditions;
- [x] implement bounded SPW orthophoto acquisition with the existing AOI worker;
- [x] implement bounded UrbIS orthophoto acquisition with the existing AOI worker;
- [x] expose provider edition, resolution, time, CRS, attribution and checksum;
- [x] reuse retry/cache/size limits and verify both live WMS providers;
- [ ] retain explicit `partial`/`not_configured` status when a service is down.
Exit: one checksum-bound RGB GeoTIFF per target stratum can be acquired without
manual URL editing and without crossing provider coverage zones.
### WP2 — Reference-label normalization
- [x] normalize GRB, PICC and UrbIS building semantics to canonical `building`;
- [x] retain original feature IDs and source classes;
- [x] quantify invalid geometry repairs, duplicates and temporal mismatches;
- [x] define exclusion rules for ruins, canopies, underground structures and
objects below the resolvable pixel size;
- [x] generate overlay/contact-sheet review packs.
Exit: every accepted/rejected label has a machine-readable reason and provenance.
### WP3 — Frozen Belgium building corpus
- [x] select and persist the initial geographic strata and split assignments;
- [x] acquire the candidate imagery and references through governed services;
- [x] export tiles at candidate resolutions with deterministic negative sampling;
- [x] run checksum, class, label-size and visual-variance audits;
- [ ] complete representative human label review;
- [ ] freeze dataset version `building-be-v1` and make it immutable.
Minimum composition is set by coverage, not one global tile count: every region
and context stratum must contain positive instances and appropriate negatives.
Any sparse stratum blocks the national claim.
Exit: corpus audit passes with zero leakage and no unreviewed critical findings.
### WP4 — Building detector training
- [ ] define a bounded experiment matrix: architecture size, initialization,
image size and augmentation policy;
- [ ] train all planned candidates on `cuda:0` with deterministic seeds where
supported, early stopping and immutable summaries;
- [ ] rerun the best configuration from a clean run directory;
- [ ] retain training curves, environment, dataset/model hashes and peak VRAM;
- [ ] do not overwrite the active production asset.
Initial matrix:
| Family | Initialization | Image sizes | Purpose |
|---|---|---|---|
| YOLO small | local generic `yolov8s.pt` | 640, 960 | reproducible baseline |
| Active-model fine-tune | current promoted building model | 640 | controlled national-domain extension |
| Larger challenger | locally approved medium model, if VRAM preflight passes | 640 | accuracy/latency trade-off |
Exit: all candidates have valid artifacts; failed/OOM runs remain recorded and
cannot be selected silently.
### WP5 — Independent evaluation and calibration
- [ ] run fixed candidate inference on frozen validation/calibration/test sets;
- [ ] calculate precision, recall, F1, AP50/AP50-95 and spatial QA at multiple
object sizes;
- [ ] report each region/context independently and macro-average them;
- [ ] audit false positives and false negatives with map/contact-sheet evidence;
- [ ] tune confidence/tile/overlap only on calibration data;
- [ ] compare against the active model with paired AOI-level results.
Provisional promotion gates (to be frozen before the test set is opened):
- zero detections in approved pure-empty background AOIs;
- no region below its approved minimum F1;
- macro precision and recall both meet the approved floor;
- candidate improves the primary macro metric or offers a documented operational
advantage without statistically/materially worse regional performance;
- inference fits RTX 4080 SUPER VRAM and bounded runtime limits;
- no unresolved high-severity visual-review findings.
Numeric floors are frozen from reviewed baseline distributions before final
testing; they are not relaxed after seeing candidate test results.
### WP6 — Promotion and deployment
- [ ] generate model card, promotion report and operator decision;
- [ ] copy, never overwrite, the immutable promoted `.pt` asset;
- [ ] configure exact class list and validated geographic scopes;
- [ ] rebuild the immutable image/runtime configuration;
- [ ] verify Docker GPU exposure, CUDA preflight and bounded live inference;
- [ ] run shadow comparison before widening the public scope;
- [ ] preserve one-command rollback to the previous asset/configuration.
Exit: API/UI show the exact supported class, regions and limitations derived
from evidence, not from file presence.
### WP7 — Building segmentation
- [ ] decide whether official footprint polygons and imagery alignment support a
trustworthy mask corpus;
- [ ] implement polygon-to-mask export with edge/occlusion rules;
- [ ] train a dedicated segmentation model; do not reuse detection boxes as masks;
- [ ] evaluate mask IoU, boundary quality, area bias and topology;
- [ ] keep `not_configured` unless segmentation-specific gates pass.
Exit: georeferenced persisted masks/polygons pass segmentation QA and export
contracts on independent regional samples.
### WP8 — Solar-panel intake
- [ ] write the user-facing analytical use case and canonical class ontology;
- [ ] secure authoritative or reviewed annotations with explicit licensing;
- [ ] create roof, industrial, ground-mounted and hard-negative strata;
- [ ] benchmark the existing local external model as a challenger only;
- [ ] train a Belgium-specific detector/segmenter if the data supports it;
- [ ] apply the same WP3WP6 promotion path.
Exit: solar capability remains `not_configured` until its own evidence passes;
building evidence cannot promote it.
### WP9 — Monitoring and controlled learning loop
- [ ] capture inference/QA drift by region, imagery edition and object size;
- [ ] maintain operator review queues without feeding decisions directly into
training labels;
- [ ] require reviewed label releases for every retraining cycle;
- [ ] retrain on a schedule only when drift/data evidence justifies it;
- [ ] retain previous datasets, models and promotion reports for rollback/audit.
Exit: no automatic self-training or silent production promotion exists.
## 7. Verification matrix
Each model release must pass:
- unit tests for transforms, labels, splits, hashes and class contracts;
- integration tests for persisted Dataset/Job/AnalysisRun/Detection or
Segmentation records;
- geospatial tests for CRS, tile transforms, clipping and duplicate suppression;
- GPU tests for Docker DeviceRequest, CUDA availability and fail-closed behavior;
- offline evaluation on frozen positive/background portfolios;
- live bounded smoke on Tower;
- frontend loading, progress, empty, error, limitation and export states;
- upgrade, restart, rollback and corrupted/missing-model failure tests.
## 8. GPU execution policy
- production training device: `cuda:0`;
- required flag: `TRAIN_REQUIRE_CUDA=true`;
- inference remains available only within the active validated scope while a
challenger trains;
- schedule long training outside operational validation windows;
- preflight free VRAM and choose batch size conservatively;
- never run two unbounded training jobs concurrently;
- persist partial-run evidence on interruption and restart as a new run unless
checkpoint resume is explicitly checksum-validated.
## 9. Artefact layout
```text
/app/storage/operator-data/model-program/<task>/<dataset-version>/
source-manifest.json
split-manifest.json
dataset.yaml
dataset-audit.json
visual-review/
/app/storage/training/<task>/<run-id>/
args.yaml
results.csv
weights/best.pt
training_summary.json
/app/storage/operator-data/model-review/<task>/<candidate-version>/
positive-portfolio/
background-portfolio/
calibration/
promotion-report.json
model-card.json
```
Large imagery, labels and weights remain outside Git. Code, schemas, fixture
contracts and documentation remain versioned in the repository.
## 10. Delivery sequence
1. WP0 contract freeze.
2. WP1 regional imagery providers.
3. WP2 source normalization.
4. WP3 frozen Belgium building corpus.
5. WP4WP5 training and independent evaluation.
6. WP6 building promotion and controlled scope widening.
7. WP7 segmentation only after the detector/data foundation is stable.
8. WP8 solar only after its label intake is approved.
9. WP9 continuously after the first national model release.
WP1WP3 are the current critical path. More epochs on the existing Kempen-only
corpus cannot satisfy the Belgian claim and are therefore not the next action.
## 11. Programme completion
The programme is complete when every target portfolio row is either:
- production-configured with immutable, independently verified evidence for its
declared classes and geography; or
- explicitly `not_configured` with the missing data/evidence named.
“Everything works” never means one universal model. It means every GeoIntel
capability uses the correct deterministic source or a task-specific, validated
model, and the platform communicates that distinction accurately.