8.7 KiB
GeoIntel Kempen — Storage Architecture v1.0
GeoIntel stores metadata in PostgreSQL/PostGIS and binary/geospatial files on filesystem storage or object storage.
Principles
- Database stores metadata, relationships and vector geometries.
- Filesystem/object storage stores original rasters, derived rasters, tiles, masks, reports and model artifacts.
- Every stored file must have a dataset/export/model record in the database.
- Never store large raster binary data directly in regular application tables in V1.
Root storage layout
storage/
uploads/
{project_id}/
rasters/
vectors/
lidar/
rasters/
derived/
{project_id}/{dataset_id}/
tiles/
{project_id}/{dataset_id}/{tile_set_id}/
masks/
{project_id}/{analysis_run_id}/
previews/
{project_id}/{dataset_id}/
exports/
{project_id}/
geojson/
csv/
reports/
coco/
yolo/
models/
detection/
segmentation/
training-runs/
cache/
grb/
osm/
sentinel/
dhmv/
Upload policy
When a file is uploaded:
- Save original file unchanged.
- Compute checksum.
- Extract metadata.
- Create dataset record.
- Create preview if applicable.
Required file metadata:
path
original_filename
mime_type
size_bytes
checksum_sha256
created_at
storage_backend
Derived data policy
Derived files must record:
- source dataset id(s)
- analysis run id or processing job id
- processing parameters
- software component version
- created_at
Segmentation mask artifacts
Sprint 9 stores segmentation masks as filesystem artifacts and segmentation polygons as authoritative PostGIS records.
Default mask path convention:
storage/masks/{project_id}/{analysis_run_id}/tile_{tile_index}/mask_{segmentation_id}.png
Optional run manifest convention:
storage/masks/{project_id}/{analysis_run_id}/manifest.json
Mask files are provenance/debug artifacts. QA, map display and GeoJSON output must use persisted segmentations.geometry rather than mask files.
Cleanup policy
Do not delete originals automatically. Derived outputs may be cleaned through explicit cache management.
Official temporal source artifacts
Waterinfo raw station layers, timeseries responses and checksum manifests live
under storage/operator-data/waterinfo/<scope>/. These are immutable source
evidence; queryable annual Point snapshots are normal Dataset/vector_feature
records. Bounded orthophotos are normal raster Dataset files. Their WMS URL,
product/layer, request/spatial hash, temporal validity and limitations are held
in source/provenance metadata. Browser PNG rendering is derived on request and
does not replace the stored GeoTIFF.
DHMV II DTM/DSM outputs are also normal raster Dataset files. The provider WCS
returns multipart coverage data; GeoIntel retains response and extracted
coverage SHA256 values in provenance, then stores one normalized, compressed,
Area-clipped GeoTIFF with its ordinary Dataset/DatasetVersion checksum. Source
metadata records the 1 m native product, 5 m default analysis grid, EPSG:31370,
-9999 nodata, TAW and acquisition period 2013-2015. Colour-relief PNGs are
derived browser views and are never authoritative. No raster binary is stored
in PostgreSQL and no DHMV file is treated as water depth or volume.
VMM flood-hazard scenario outputs follow the same raster Dataset policy. The
bounded acquisition service stores one normalized, compressed, Area-clipped
GeoTIFF per scenario/request identity. provision_regional_flood_hazards.py
does not create a parallel storage layout: every municipality/scenario result
is an ordinary Dataset and DatasetVersion with WCS request hashes, tile request
URLs, response/coverage/normalized checksums, EPSG:31370 bounds, scenario
metadata and explicit unsupported-volume flags. Repeat runs reuse matching
ready Datasets through the acquisition service cache.
BWK/Natura 2000 evidence lives under
storage/operator-evidence/bwk-natura2000-2025/mol/. The raw/ directory
contains immutable WFS pages; the adjacent manifest records their URLs,
checksums, boundary checksum, page/feature counts, clipping diagnostics and
official class totals. The final GeoJSON is the retained source artifact while
datasets and vector_features remain the canonical queryable PostGIS state.
Definitive agricultural-use parcel evidence lives under:
storage/operator-evidence/agricultural-use-parcels/{scope}/{year}/
agpa_{year}_*_public.zip
agricultural_use_parcels_{year}_{scope}.geojson
agricultural_use_parcels_{year}_crop_codes.json
agricultural_use_parcels_{year}_{scope}.manifest.json
The official ZIP is immutable source evidence and is never removed by normal cache cleanup. The extracted GeoPackage is temporary to avoid retaining a second full source copy. The normalized GeoJSON is the canonical upload artifact; Dataset and vector_feature rows remain the queryable PostGIS state. The manifest binds source, crop-code list and upload artifact checksums. A checksum conflict with an existing annual Dataset fails closed.
Buildings and Addresses Register snapshot evidence lives under:
storage/operator-evidence/buildings-addresses-register/{area-key}/{observed-date}/
raw/
buildings_page_*.json
building_units_page_*.json
addresses_page_*.json
buildings_addresses_register.geojson
buildings_addresses_register.manifest.json
Raw pages contain the unmodified official response and are retained only as
checksummed operator evidence. They can contain address labels and must never
be served as a map/API artifact. The normalized GeoJSON deliberately contains
one polygon per register building with lifecycle state, aggregate relation
counts and classified GRB reconciliation only. It enters PostGIS exclusively
through DatasetService and ordinary vector_features; the operator never
writes database rows directly. The manifest binds all raw pages, the normalized
artifact, exact Area boundary and every GRB partition used for reconciliation.
The Area key is deterministic (mol, geel, and so on). Reuse also requires
the same Area label, observation date and boundary checksum, preventing one
municipality's evidence from being imported under another Area.
Offline demo export artifacts can be inspected and cleaned with:
python scripts/cleanup_demo_artifacts.py
python scripts/cleanup_demo_artifacts.py --keep-latest 10 --export-type project_report_html
python scripts/cleanup_demo_artifacts.py --keep-latest 10 --max-delete 100 --apply
docker compose exec -T backend python scripts/cleanup_demo_artifacts.py
The script is dry-run by default, targets only the explicit
GeoIntel Demo - Building QA project unless an exact --project-name is
provided, keeps the newest export artifacts per matching project and refuses to
delete files outside STORAGE_ROOT. It cleans exports records/files only; it
does not remove original uploads, vector features, QA/QC rows, projects, areas,
tiles, rasters or masks. --max-delete defaults to 25 and blocks oversized
apply runs until the operator increases the cap after reviewing dry-run output.
Repeat --export-type to restrict cleanup to selected artifact kinds.
Live runtime validation for this maintenance path is available as a dry-run smoke:
bash scripts/verify_demo_cleanup_dry_run.sh
CLEANUP_MODE=container CLEANUP_CONTAINER=geointel bash scripts/verify_demo_cleanup_dry_run.sh
The smoke never passes --apply. It fails if the cleanup summary is not a
dry-run, if any export/file deletion is reported, or if the dry-run candidate
fields are missing.
Model storage
Model artifacts live under:
storage/models/
Database model registry records:
model_id
name
task_type
framework
path
classes_json
version
created_at
metrics_json
Exports
Every export is reproducible and linked to project/analysis.
Export record fields:
id
project_id
analysis_run_id
export_type
path
format
created_at
parameters_json
Local development default
Use local filesystem paths. Keep MinIO/object storage as future extension.
Regional BWK/Natura 2000 evidence
The regional state-2025 operator stores immutable evidence under:
storage/operator-evidence/bwk-natura2000-2025/regional/
kempen-transport-region/partitions/{nis_code}/
raw/*.json.gz
bwk_natura2000_2025.geojson
manifest.json
snapshots/kempen-transport-region/
bwk_natura2000_2025.geojson
manifest.json
Partition reuse requires the same source version, municipality identity, boundary checksum, normalized output checksum and all raw evidence checksums. Snapshot reuse additionally binds the ordered set of 28 partition output checksums. Only the assembled snapshot enters DatasetService/PostGIS; raw WFS responses and partition files remain operator evidence on persistent storage.