8.2 KiB
Phase 1 command and result ledger
Captured: 2026-08-01
Audited repository baseline: 0c019bb22f816db1e4b7a68379bcad08924d9a21
Branch: codex/geointel-accuracy-program
This ledger records the commands and claim boundaries behind the retained evidence. JSON, JUnit, SQL and full text logs in this directory are the source of truth when this summary and a raw artifact differ.
Repository and static inventory
python scripts/run_accuracy_phase1_baseline.py --output-dir artifacts/evidence/accuracy/P1
Result: completed. The baseline found 2,495 tracked files, including a tracked
geointel/ mirror with 1,153 files; 68 paired files differ from root. No model
checkpoint is tracked in the local checkout. The mirror is excluded by the root
.dockerignore; the residual risk is source/import/maintenance ambiguity, not
official all-in-one build-context inclusion.
Evidence:
phase1-baseline-summary.jsonrepository-inventory.jsonlocal-artifact-inventory.jsonstatic-risk-signals.json
Deterministic contract reproductions
python scripts/reproduce_accuracy_phase1_findings.py
Result: 7 of 7 expected violations reproduced. They cover cross-theme coverage contamination, metres-as-degrees buffering, Lambert-as-4326 persistence, caller-spoofed source authority, mutable-name model scope, mutable-name legal scope and silently ignored Area PATCH geometry.
Evidence: forensic-reproductions.json.
Tower runtime, database and storage references
The collector was streamed into the running geointel container and executed
with read-only SQL and a 30-second statement timeout:
docker exec -i geointel env PYTHONPATH=/app/backend python -
Result:
- Python 3.11.2, PyTorch 2.11.0+cu128, CUDA 12.8, Ultralytics 8.4.99;
- NVIDIA GeForce RTX 4080 SUPER visible on
cuda:0; - Alembic head
202607260001; - 5,816 direct database storage references checked, 0 missing;
- four Geel detections contain Lambert-domain coordinates while the geometry column reports SRID 4326;
- no database or storage row was modified.
A broader recursive audit_data_operations.py storage scan was stopped by its
244-second execution timeout and produced no retained result. It is not counted
as a pass. The bounded, direct-reference collector above completed and is the
only storage-completeness claim made here.
Evidence:
tower-runtime-database-snapshot.jsontower-runtime-database-snapshot-detailed.json
Real GPU inference smoke
The read-only collector used the production YoloDetectionAdapter, the active
model and one existing Geel tile:
python - --model-path /app/models/geointel-building-yolov8s-smallbld-minpx3-img640-ft30.pt \
--tile-path /app/storage/tiles/cb80638d-dbef-48ac-b19c-cec7c3efc96e/ae0ff76d-70c0-404f-b777-54d14517179a/b191c968-7d56-4e0d-afbb-8b5baaa62470/tile_0000.tif \
--manifest-path /app/storage/tiles/cb80638d-dbef-48ac-b19c-cec7c3efc96e/ae0ff76d-70c0-404f-b777-54d14517179a/b191c968-7d56-4e0d-afbb-8b5baaa62470/manifest.json \
--confidence 0.5 --image-size 640 --max-detections 1000 --device cuda:0 --seed 20260801
Result: passed; 17 raw building detections; 0.883694486 s synchronized
inference; model SHA-256
a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1.
Claim boundary: this proves one runtime execution. It does not prove that any box is correct, that confidence is calibrated, or that the model generalizes.
Evidence: tower-gpu-inference-smoke.json.
ML/data-lineage inventory
The bounded collector read V56/V58/V62/V66 plus model, checkpoint, report and manifest metadata from the mounted Tower volume:
docker exec -i geointel python - --app-root /app
Result:
- 26 model assets, 229 training checkpoints, 424 training JSON reports and 36 operator manifests inventoried;
- V56 has 180 AOIs and 0 completed human reviews;
- three pure-empty background-test AOIs: Flanders 2, Wallonia 1, Brussels 0;
- zero exact cross-split raster-hash duplicates;
- bounded 64-bit dHash screen: 180 rasters, minimum Hamming distance 17, zero cross-split pairs at or below 4;
- minimum cross-split AOI bbox distance 95.720336 m and 24 pairs below 2 km;
- V58/V62 evidence is calibration-only; at threshold 0.15 aggregate F1 is 0.512905 while Flanders recall and F1 are 0;
- no protected-test, background-test promotion or national-release evidence.
Claim boundary: exact/dHash and bbox-distance screens do not establish municipality, flight-strip, instance, semantic or imagery-edition independence.
Evidence: tower-ml-data-lineage-snapshot.json and
tower-key-artifact-hashes.json (V62 best/last checkpoints and V56 tile audit).
Backend tests
Full root-context suite
$env:PYTHONPATH='C:\Projects\geointel\backend;C:\Projects\geointel'
python -m pytest backend/tests -q -p no:cacheprovider \
-W error::DeprecationWarning \
--junitxml=artifacts/evidence/accuracy/P1/backend-full-suite.junit.xml
Result: exit 1; 1,180 passed and 17 failed in 70.63 seconds. The 17 failures are stale source/contract assertions, including expectations that conflict with the new explicit user-selected model flow. They remain release-blocking until replaced with reviewed behavior tests.
Evidence:
backend-full-suite.txtbackend-full-suite.junit.xml
Actual backend CI working directory
cd backend
python -m pytest -W error::DeprecationWarning
Result: exit 2 during collection; 1,194 items collected plus one import error:
ModuleNotFoundError: No module named 'scripts.render_operator_polygon_label_qa'
Evidence: backend-ci-entrypoint.txt.
Phase 1 tooling
python -m pytest tests/test_accuracy_phase1_baseline.py -q -p no:cacheprovider
Result: 4 passed.
Evidence: phase1-tooling-tests.txt.
Lint
python -m ruff check scripts/collect_accuracy_phase1_inference_smoke.py \
scripts/collect_accuracy_phase1_ml_lineage.py \
scripts/collect_accuracy_phase1_runtime.py \
scripts/reproduce_accuracy_phase1_findings.py \
scripts/run_accuracy_phase1_baseline.py tests/test_accuracy_phase1_baseline.py
Result: all new Phase 1 code passed.
python -m ruff check backend scripts tests
Result: exit 1; 112 findings: E402 13, E701 2, E702 69, F401 23, F403 1, F811 2 and F841 2.
Evidence:
repository-ruff-baseline.txtrepository-ruff-baseline.json
Frontend
cd frontend
npm run test:unit
npm run typecheck
npm run build
Results:
- Vitest: 16 files, 51 tests passed;
- TypeScript typecheck: passed;
- production build: passed, 1,896 modules transformed.
The attempted generic npm test -- --run failed because no test script
exists; the configured command is test:unit. The test strategy also calls
for lint, but npm run lint fails with Missing script: "lint".
Evidence:
frontend-vitest.txt(failed non-existent generic command, retained)frontend-vitest-unit.txt(correct configured command, passed)frontend-typecheck.txtfrontend-build.txtfrontend-lint.txt
API, migrations and golden QA
python scripts/audit_api_contracts.py
Result: 147 implemented routes match documentation; 10 explicitly tracked non-envelope endpoints.
Evidence: openapi-contract-audit.txt.
cd backend
python -m alembic heads
python -m alembic upgrade head --sql
Result: one head, 202607260001; complete offline upgrade rendered 496 SQL/log
lines. This is not a local live-PostGIS migration test.
Evidence:
alembic-heads.txtalembic-offline-upgrade.sql
python scripts/run_golden_qa_benchmark.py --json
Result: two runs have semantically identical metric results but different file hashes because run identities use UUID4.
Evidence:
golden-qa-run-1.jsongolden-qa-run-2.jsongolden-qa-reproducibility.json
Explicitly not executed as a success gate
- no protected test or background-test portfolio was opened;
- no training, calibration fitting, checkpoint promotion or active-model change;
- no label was accepted on behalf of a human reviewer;
- no production database migration or data repair;
- no dataset, checkpoint, cache, output or user-owned untracked file deleted, rewritten or moved;
- no segmentation, SAM or solar asset was inferred to be production-ready from file presence.