Files
geointel/artifacts/evidence/accuracy/P1/commands-and-results.md
T

8.2 KiB

Phase 1 command and result ledger

Captured: 2026-08-01
Audited repository baseline: 0c019bb22f816db1e4b7a68379bcad08924d9a21
Branch: codex/geointel-accuracy-program

This ledger records the commands and claim boundaries behind the retained evidence. JSON, JUnit, SQL and full text logs in this directory are the source of truth when this summary and a raw artifact differ.

Repository and static inventory

python scripts/run_accuracy_phase1_baseline.py --output-dir artifacts/evidence/accuracy/P1

Result: completed. The baseline found 2,495 tracked files, including a tracked geointel/ mirror with 1,153 files; 68 paired files differ from root. No model checkpoint is tracked in the local checkout. The mirror is excluded by the root .dockerignore; the residual risk is source/import/maintenance ambiguity, not official all-in-one build-context inclusion.

Evidence:

  • phase1-baseline-summary.json
  • repository-inventory.json
  • local-artifact-inventory.json
  • static-risk-signals.json

Deterministic contract reproductions

python scripts/reproduce_accuracy_phase1_findings.py

Result: 7 of 7 expected violations reproduced. They cover cross-theme coverage contamination, metres-as-degrees buffering, Lambert-as-4326 persistence, caller-spoofed source authority, mutable-name model scope, mutable-name legal scope and silently ignored Area PATCH geometry.

Evidence: forensic-reproductions.json.

Tower runtime, database and storage references

The collector was streamed into the running geointel container and executed with read-only SQL and a 30-second statement timeout:

docker exec -i geointel env PYTHONPATH=/app/backend python -

Result:

  • Python 3.11.2, PyTorch 2.11.0+cu128, CUDA 12.8, Ultralytics 8.4.99;
  • NVIDIA GeForce RTX 4080 SUPER visible on cuda:0;
  • Alembic head 202607260001;
  • 5,816 direct database storage references checked, 0 missing;
  • four Geel detections contain Lambert-domain coordinates while the geometry column reports SRID 4326;
  • no database or storage row was modified.

A broader recursive audit_data_operations.py storage scan was stopped by its 244-second execution timeout and produced no retained result. It is not counted as a pass. The bounded, direct-reference collector above completed and is the only storage-completeness claim made here.

Evidence:

  • tower-runtime-database-snapshot.json
  • tower-runtime-database-snapshot-detailed.json

Real GPU inference smoke

The read-only collector used the production YoloDetectionAdapter, the active model and one existing Geel tile:

python - --model-path /app/models/geointel-building-yolov8s-smallbld-minpx3-img640-ft30.pt \
  --tile-path /app/storage/tiles/cb80638d-dbef-48ac-b19c-cec7c3efc96e/ae0ff76d-70c0-404f-b777-54d14517179a/b191c968-7d56-4e0d-afbb-8b5baaa62470/tile_0000.tif \
  --manifest-path /app/storage/tiles/cb80638d-dbef-48ac-b19c-cec7c3efc96e/ae0ff76d-70c0-404f-b777-54d14517179a/b191c968-7d56-4e0d-afbb-8b5baaa62470/manifest.json \
  --confidence 0.5 --image-size 640 --max-detections 1000 --device cuda:0 --seed 20260801

Result: passed; 17 raw building detections; 0.883694486 s synchronized inference; model SHA-256 a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1.

Claim boundary: this proves one runtime execution. It does not prove that any box is correct, that confidence is calibrated, or that the model generalizes.

Evidence: tower-gpu-inference-smoke.json.

ML/data-lineage inventory

The bounded collector read V56/V58/V62/V66 plus model, checkpoint, report and manifest metadata from the mounted Tower volume:

docker exec -i geointel python - --app-root /app

Result:

  • 26 model assets, 229 training checkpoints, 424 training JSON reports and 36 operator manifests inventoried;
  • V56 has 180 AOIs and 0 completed human reviews;
  • three pure-empty background-test AOIs: Flanders 2, Wallonia 1, Brussels 0;
  • zero exact cross-split raster-hash duplicates;
  • bounded 64-bit dHash screen: 180 rasters, minimum Hamming distance 17, zero cross-split pairs at or below 4;
  • minimum cross-split AOI bbox distance 95.720336 m and 24 pairs below 2 km;
  • V58/V62 evidence is calibration-only; at threshold 0.15 aggregate F1 is 0.512905 while Flanders recall and F1 are 0;
  • no protected-test, background-test promotion or national-release evidence.

Claim boundary: exact/dHash and bbox-distance screens do not establish municipality, flight-strip, instance, semantic or imagery-edition independence.

Evidence: tower-ml-data-lineage-snapshot.json and tower-key-artifact-hashes.json (V62 best/last checkpoints and V56 tile audit).

Backend tests

Full root-context suite

$env:PYTHONPATH='C:\Projects\geointel\backend;C:\Projects\geointel'
python -m pytest backend/tests -q -p no:cacheprovider \
  -W error::DeprecationWarning \
  --junitxml=artifacts/evidence/accuracy/P1/backend-full-suite.junit.xml

Result: exit 1; 1,180 passed and 17 failed in 70.63 seconds. The 17 failures are stale source/contract assertions, including expectations that conflict with the new explicit user-selected model flow. They remain release-blocking until replaced with reviewed behavior tests.

Evidence:

  • backend-full-suite.txt
  • backend-full-suite.junit.xml

Actual backend CI working directory

cd backend
python -m pytest -W error::DeprecationWarning

Result: exit 2 during collection; 1,194 items collected plus one import error:

ModuleNotFoundError: No module named 'scripts.render_operator_polygon_label_qa'

Evidence: backend-ci-entrypoint.txt.

Phase 1 tooling

python -m pytest tests/test_accuracy_phase1_baseline.py -q -p no:cacheprovider

Result: 4 passed.

Evidence: phase1-tooling-tests.txt.

Lint

python -m ruff check scripts/collect_accuracy_phase1_inference_smoke.py \
  scripts/collect_accuracy_phase1_ml_lineage.py \
  scripts/collect_accuracy_phase1_runtime.py \
  scripts/reproduce_accuracy_phase1_findings.py \
  scripts/run_accuracy_phase1_baseline.py tests/test_accuracy_phase1_baseline.py

Result: all new Phase 1 code passed.

python -m ruff check backend scripts tests

Result: exit 1; 112 findings: E402 13, E701 2, E702 69, F401 23, F403 1, F811 2 and F841 2.

Evidence:

  • repository-ruff-baseline.txt
  • repository-ruff-baseline.json

Frontend

cd frontend
npm run test:unit
npm run typecheck
npm run build

Results:

  • Vitest: 16 files, 51 tests passed;
  • TypeScript typecheck: passed;
  • production build: passed, 1,896 modules transformed.

The attempted generic npm test -- --run failed because no test script exists; the configured command is test:unit. The test strategy also calls for lint, but npm run lint fails with Missing script: "lint".

Evidence:

  • frontend-vitest.txt (failed non-existent generic command, retained)
  • frontend-vitest-unit.txt (correct configured command, passed)
  • frontend-typecheck.txt
  • frontend-build.txt
  • frontend-lint.txt

API, migrations and golden QA

python scripts/audit_api_contracts.py

Result: 147 implemented routes match documentation; 10 explicitly tracked non-envelope endpoints.

Evidence: openapi-contract-audit.txt.

cd backend
python -m alembic heads
python -m alembic upgrade head --sql

Result: one head, 202607260001; complete offline upgrade rendered 496 SQL/log lines. This is not a local live-PostGIS migration test.

Evidence:

  • alembic-heads.txt
  • alembic-offline-upgrade.sql
python scripts/run_golden_qa_benchmark.py --json

Result: two runs have semantically identical metric results but different file hashes because run identities use UUID4.

Evidence:

  • golden-qa-run-1.json
  • golden-qa-run-2.json
  • golden-qa-reproducibility.json

Explicitly not executed as a success gate

  • no protected test or background-test portfolio was opened;
  • no training, calibration fitting, checkpoint promotion or active-model change;
  • no label was accepted on behalf of a human reviewer;
  • no production database migration or data repair;
  • no dataset, checkpoint, cache, output or user-owned untracked file deleted, rewritten or moved;
  • no segmentation, SAM or solar asset was inferred to be production-ready from file presence.