290 lines
10 KiB
Markdown
290 lines
10 KiB
Markdown
# Database Implementation Plan
|
|
|
|
Database: PostgreSQL + PostGIS.
|
|
|
|
## Rules
|
|
|
|
- Store geometries in PostGIS with explicit SRID.
|
|
- Preserve original CRS metadata even when normalized geometry is stored as EPSG:4326 or a local projected CRS.
|
|
- Prefer UUID primary keys.
|
|
- Store large raster/mask/model files in filesystem or object storage; store metadata and paths in PostgreSQL.
|
|
- Keep analysis outputs reproducible by storing parameters JSON.
|
|
|
|
## Required extensions
|
|
|
|
```sql
|
|
CREATE EXTENSION IF NOT EXISTS postgis;
|
|
CREATE EXTENSION IF NOT EXISTS postgis_topology;
|
|
CREATE EXTENSION IF NOT EXISTS "uuid-ossp";
|
|
```
|
|
|
|
## Core tables
|
|
|
|
### projects
|
|
|
|
- `id uuid primary key`
|
|
- `name text not null`
|
|
- `description text`
|
|
- `region text default 'Kempen'`
|
|
- `status text default 'active'`
|
|
- `created_at timestamptz`
|
|
- `updated_at timestamptz`
|
|
|
|
### areas
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `name text not null`
|
|
- `geometry geometry(MultiPolygon, 4326) not null`
|
|
- `original_crs text`
|
|
- `area_m2 double precision`
|
|
- `bbox geometry(Polygon, 4326)`
|
|
- `created_at timestamptz`
|
|
|
|
Spatial index required on `geometry`.
|
|
|
|
### datasets
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `area_id uuid nullable references areas(id)`
|
|
- `name text not null`
|
|
- `dataset_type text not null`
|
|
- `source text not null`
|
|
- `storage_path text`
|
|
- `derived_from_dataset_id uuid nullable references datasets(id)`
|
|
- `crs text`
|
|
- `bounds_json jsonb`
|
|
- `resolution_json jsonb`
|
|
- `bands_json jsonb`
|
|
- `metadata_json jsonb`
|
|
- `status text default 'created'`
|
|
- `created_at timestamptz`
|
|
|
|
### vector_features
|
|
|
|
Used for imported vector datasets and derived vector outputs when feature-level storage is needed. Original files remain source artifacts; this table is the queryable PostGIS state for vector features.
|
|
|
|
- `id uuid primary key`
|
|
- `dataset_id uuid references datasets(id) on delete cascade`
|
|
- `feature_class text`
|
|
- `source_feature_id text`
|
|
- `properties_json jsonb`
|
|
- `geometry geometry(Geometry, 4326) not null`
|
|
- `created_at timestamptz`
|
|
|
|
Required indexes:
|
|
|
|
- `dataset_id`
|
|
- GiST index on `geometry`
|
|
|
|
### analysis_runs
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `area_id uuid references areas(id)`
|
|
- `dataset_id uuid nullable references datasets(id)`
|
|
- `job_id uuid nullable references jobs(id)`
|
|
- `analysis_type text not null`
|
|
- `status text not null`
|
|
- `model_name text nullable`
|
|
- `model_version text nullable`
|
|
- `parameters_json jsonb not null`
|
|
- `result_json jsonb nullable`
|
|
- `created_at timestamptz`
|
|
- `started_at timestamptz`
|
|
- `finished_at timestamptz`
|
|
- `error_message text`
|
|
|
|
Analysis runs are domain lifecycle records. Jobs track execution state; analysis runs track reproducibility, model metadata, parameters and result summaries.
|
|
|
|
### detections
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `dataset_id uuid nullable references datasets(id)`
|
|
- `analysis_run_id uuid nullable references analysis_runs(id)`
|
|
- `job_id uuid nullable references jobs(id)`
|
|
- `model_name text not null`
|
|
- `model_version text nullable`
|
|
- `class_name text not null`
|
|
- `confidence double precision not null`
|
|
- `geometry geometry(Geometry, 4326)`
|
|
- `bbox_json jsonb`
|
|
- `source_tile_path text nullable`
|
|
- `properties_json jsonb`
|
|
- `created_at timestamptz`
|
|
|
|
Required indexes:
|
|
|
|
- `project_id`
|
|
- `dataset_id`
|
|
- `analysis_run_id`
|
|
- `class_name`
|
|
- GiST index on `geometry`
|
|
|
|
Sprint 8 persists detections as first-class PostGIS records. Detections are never stored only in `jobs.result_json`.
|
|
|
|
### segmentations
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `dataset_id uuid nullable references datasets(id)`
|
|
- `job_id uuid nullable references jobs(id)`
|
|
- `analysis_run_id uuid nullable references analysis_runs(id)`
|
|
- `model_name text not null`
|
|
- `model_version text nullable`
|
|
- `class_name text not null`
|
|
- `confidence double precision nullable`
|
|
- `geometry geometry(MultiPolygon, 4326) not null`
|
|
- `bbox_json jsonb`
|
|
- `area_m2 double precision`
|
|
- `mask_path text`
|
|
- `source_tile_path text`
|
|
- `tile_index integer`
|
|
- `properties_json jsonb`
|
|
- `provenance_json jsonb`
|
|
- `created_at timestamptz`
|
|
|
|
Required indexes:
|
|
|
|
- `project_id`
|
|
- `dataset_id`
|
|
- `analysis_run_id`
|
|
- `job_id`
|
|
- `class_name`
|
|
- GiST index on `geometry`
|
|
|
|
Sprint 9 persists segmentation outputs as first-class PostGIS records. Mask paths are artifact/provenance references only; map display, QA and GeoJSON output use `segmentations.geometry`.
|
|
|
|
### metrics
|
|
|
|
- `id uuid primary key`
|
|
- `quality_check_id uuid nullable references quality_checks(id)`
|
|
- `analysis_run_id uuid nullable references analysis_runs(id)`
|
|
- `metric_key text not null`
|
|
- `metric_value double precision`
|
|
- `metric_unit text`
|
|
- `label text`
|
|
- `metadata_json jsonb`
|
|
- `created_at timestamptz`
|
|
|
|
Metrics may belong to a quality check, an analysis run, or both. Sprint 7A persists QA/QC metrics through `quality_check_id`.
|
|
|
|
### quality_checks
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `job_id uuid nullable references jobs(id)`
|
|
- `analysis_run_id uuid nullable references analysis_runs(id)`
|
|
- `candidate_dataset_id uuid nullable references datasets(id)`
|
|
- `reference_dataset_id uuid references datasets(id)`
|
|
- `check_type text not null`
|
|
- `status text not null`
|
|
- `score double precision`
|
|
- `parameters_json jsonb`
|
|
- `findings_json jsonb`
|
|
- `created_at timestamptz`
|
|
- `completed_at timestamptz nullable`
|
|
|
|
Quality checks are domain records. Jobs track execution state; quality checks track the persisted QA/QC result; metrics track individual measurements.
|
|
|
|
Detection QA coverage and box-to-footprint diagnostics require no schema
|
|
change. The canonical metric rows remain precision, recall, F1, mean IoU and
|
|
false-positive/negative counts. Tile coverage population counts and the
|
|
diagnostic reference-envelope comparison are persisted in the existing
|
|
`quality_checks.findings_json`; `parameters_json.coverage_policy` records the
|
|
evaluation policy used for reproducibility.
|
|
|
|
### detection_reviews
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id) on delete cascade`
|
|
- `quality_check_id uuid references quality_checks(id) on delete cascade`
|
|
- `analysis_run_id uuid nullable references analysis_runs(id) on delete set null`
|
|
- `evidence_role text not null` (`false_positive` or `false_negative`)
|
|
- `evidence_feature_id text not null`
|
|
- `detection_id uuid nullable references detections(id) on delete set null`
|
|
- `reference_feature_id uuid nullable references vector_features(id) on delete set null`
|
|
- `decision text not null default 'unreviewed'`
|
|
- `notes text nullable`
|
|
- `reviewed_by text not null default 'operator'`
|
|
- `created_at timestamptz`
|
|
- `updated_at timestamptz`
|
|
|
|
The unique key is `(quality_check_id, evidence_role, evidence_feature_id)`.
|
|
Indexes cover project, quality check, analysis run and decision. Reviews
|
|
classify persisted QA evidence only; they do not replace or modify Detection,
|
|
VectorFeature, QualityCheck or Metric records. Unreviewed, reference-gap,
|
|
imagery-uncertain and QA-alignment cases are not training labels.
|
|
|
|
### exports
|
|
|
|
- `id uuid primary key`
|
|
- `project_id uuid references projects(id)`
|
|
- `analysis_run_id uuid nullable references analysis_runs(id)`
|
|
- `export_type text not null`
|
|
- `storage_path text not null`
|
|
- `metadata_json jsonb`
|
|
- `created_at timestamptz`
|
|
|
|
## Migration strategy
|
|
|
|
- Use Alembic.
|
|
- First migration creates extensions and core tables.
|
|
- Second migration adds spatial indexes.
|
|
- Seed script may create a sample project and sample area only if explicitly run.
|
|
|
|
## Sprint 7B provider-to-dataset mapping
|
|
|
|
Provider integration is a contract layer only in Sprint 7B. Providers do not write directly to `vector_features`; future provider output must flow through `DatasetService` and `VectorFeatureService` so dataset provenance, storage metadata and feature persistence remain consistent.
|
|
|
|
- `grb`: maps to `dataset_role='reference'`, `source_name='grb'`.
|
|
- `osm`: maps to `dataset_role='source'` by default, or `dataset_role='reference'` only when explicitly requested; `source_name='osm'`.
|
|
- `manual`: maps to `dataset_role='reference'`, `source_name='manual'`.
|
|
- `fixture`: maps to `dataset_role='reference'`, `source_name='fixture'`.
|
|
|
|
GRB and OSM live imports are intentionally `not_configured` in Sprint 7B. Manual and fixture reference datasets use existing upload and fixture flows.
|
|
|
|
Large explicit operator imports may invoke DatasetService directly inside the
|
|
backend container when a single multipart upload would require loading the
|
|
entire regional artifact into memory. The regional GRB building operator still
|
|
creates a normal Dataset and DatasetVersion and delegates every queryable row
|
|
to VectorFeatureService. It copies the immutable combined artifact in a
|
|
stream, loads one retained municipality partition at a time, flushes bounded
|
|
feature batches and commits only when the indexed count matches the manifest.
|
|
It does not expose a direct SQL/provider write path and does not alter the
|
|
public provider endpoint's `not_configured` status.
|
|
|
|
## Geometry normalization
|
|
|
|
- User-drawn polygons arrive as EPSG:4326.
|
|
- Uploaded vector data may arrive in another CRS; preserve original CRS and reproject to EPSG:4326 for storage.
|
|
- Area calculations should use a projected CRS suitable for Belgium, preferably EPSG:31370 or another documented Belgian projection.
|
|
|
|
## Out of scope for V1
|
|
|
|
- Raster-in-database storage.
|
|
- Multi-tenant row-level security.
|
|
- User accounts.
|
|
- Full model registry tables.
|
|
|
|
## Temporal dataset foundation
|
|
|
|
Historical observations remain normal `datasets` and `vector_features`; there
|
|
is no parallel temporal feature store. Snapshots are grouped by
|
|
`datasets.temporal_series_key` and carry `observed_at`, `valid_from`,
|
|
`valid_to`, `temporal_granularity` and `source_version`. Observation time is
|
|
kept separate from ingestion time (`imported_at`).
|
|
|
|
`dataset_versions` records immutable storage provenance for every upload and
|
|
derived output: dataset-local `version`, storage path, source version,
|
|
observation/validity dates, checksum and source/provenance JSON. The
|
|
`(dataset_id, version)` pair is unique. Temporal series lookup is indexed by
|
|
`(project_id, temporal_series_key, observed_at)` and source-feature lookup by
|
|
`(dataset_id, source_feature_id)`.
|
|
|
|
Time-series comparison is read-only and aggregates persisted geometry inside a
|
|
requested bbox. Object-level added/removed/modified evidence is only valid for
|
|
sources that explicitly declare stable source feature identifiers.
|