audit: establish accuracy phase 1 baseline
This commit is contained in:
@@ -0,0 +1,266 @@
|
||||
# Phase 1 command and result ledger
|
||||
|
||||
Captured: 2026-08-01
|
||||
Audited repository baseline: `0c019bb22f816db1e4b7a68379bcad08924d9a21`
|
||||
Branch: `codex/geointel-accuracy-program`
|
||||
|
||||
This ledger records the commands and claim boundaries behind the retained
|
||||
evidence. JSON, JUnit, SQL and full text logs in this directory are the source
|
||||
of truth when this summary and a raw artifact differ.
|
||||
|
||||
## Repository and static inventory
|
||||
|
||||
```powershell
|
||||
python scripts/run_accuracy_phase1_baseline.py --output-dir artifacts/evidence/accuracy/P1
|
||||
```
|
||||
|
||||
Result: completed. The baseline found 2,495 tracked files, including a tracked
|
||||
`geointel/` mirror with 1,153 files; 68 paired files differ from root. No model
|
||||
checkpoint is tracked in the local checkout. The mirror is excluded by the root
|
||||
`.dockerignore`; the residual risk is source/import/maintenance ambiguity, not
|
||||
official all-in-one build-context inclusion.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `phase1-baseline-summary.json`
|
||||
- `repository-inventory.json`
|
||||
- `local-artifact-inventory.json`
|
||||
- `static-risk-signals.json`
|
||||
|
||||
## Deterministic contract reproductions
|
||||
|
||||
```powershell
|
||||
python scripts/reproduce_accuracy_phase1_findings.py
|
||||
```
|
||||
|
||||
Result: 7 of 7 expected violations reproduced. They cover cross-theme coverage
|
||||
contamination, metres-as-degrees buffering, Lambert-as-4326 persistence,
|
||||
caller-spoofed source authority, mutable-name model scope, mutable-name legal
|
||||
scope and silently ignored Area PATCH geometry.
|
||||
|
||||
Evidence: `forensic-reproductions.json`.
|
||||
|
||||
## Tower runtime, database and storage references
|
||||
|
||||
The collector was streamed into the running `geointel` container and executed
|
||||
with read-only SQL and a 30-second statement timeout:
|
||||
|
||||
```text
|
||||
docker exec -i geointel env PYTHONPATH=/app/backend python -
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- Python 3.11.2, PyTorch 2.11.0+cu128, CUDA 12.8, Ultralytics 8.4.99;
|
||||
- NVIDIA GeForce RTX 4080 SUPER visible on `cuda:0`;
|
||||
- Alembic head `202607260001`;
|
||||
- 5,816 direct database storage references checked, 0 missing;
|
||||
- four Geel detections contain Lambert-domain coordinates while the geometry
|
||||
column reports SRID 4326;
|
||||
- no database or storage row was modified.
|
||||
|
||||
A broader recursive `audit_data_operations.py` storage scan was stopped by its
|
||||
244-second execution timeout and produced no retained result. It is not counted
|
||||
as a pass. The bounded, direct-reference collector above completed and is the
|
||||
only storage-completeness claim made here.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tower-runtime-database-snapshot.json`
|
||||
- `tower-runtime-database-snapshot-detailed.json`
|
||||
|
||||
## Real GPU inference smoke
|
||||
|
||||
The read-only collector used the production `YoloDetectionAdapter`, the active
|
||||
model and one existing Geel tile:
|
||||
|
||||
```text
|
||||
python - --model-path /app/models/geointel-building-yolov8s-smallbld-minpx3-img640-ft30.pt \
|
||||
--tile-path /app/storage/tiles/cb80638d-dbef-48ac-b19c-cec7c3efc96e/ae0ff76d-70c0-404f-b777-54d14517179a/b191c968-7d56-4e0d-afbb-8b5baaa62470/tile_0000.tif \
|
||||
--manifest-path /app/storage/tiles/cb80638d-dbef-48ac-b19c-cec7c3efc96e/ae0ff76d-70c0-404f-b777-54d14517179a/b191c968-7d56-4e0d-afbb-8b5baaa62470/manifest.json \
|
||||
--confidence 0.5 --image-size 640 --max-detections 1000 --device cuda:0 --seed 20260801
|
||||
```
|
||||
|
||||
Result: passed; 17 raw building detections; 0.883694486 s synchronized
|
||||
inference; model SHA-256
|
||||
`a9088b8491dfae36694b53e9e9406cb4e3511d334a5712fa34f75078a47759c1`.
|
||||
|
||||
Claim boundary: this proves one runtime execution. It does not prove that any
|
||||
box is correct, that confidence is calibrated, or that the model generalizes.
|
||||
|
||||
Evidence: `tower-gpu-inference-smoke.json`.
|
||||
|
||||
## ML/data-lineage inventory
|
||||
|
||||
The bounded collector read V56/V58/V62/V66 plus model, checkpoint, report and
|
||||
manifest metadata from the mounted Tower volume:
|
||||
|
||||
```text
|
||||
docker exec -i geointel python - --app-root /app
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- 26 model assets, 229 training checkpoints, 424 training JSON reports and
|
||||
36 operator manifests inventoried;
|
||||
- V56 has 180 AOIs and 0 completed human reviews;
|
||||
- three pure-empty background-test AOIs: Flanders 2, Wallonia 1, Brussels 0;
|
||||
- zero exact cross-split raster-hash duplicates;
|
||||
- bounded 64-bit dHash screen: 180 rasters, minimum Hamming distance 17,
|
||||
zero cross-split pairs at or below 4;
|
||||
- minimum cross-split AOI bbox distance 95.720336 m and 24 pairs below 2 km;
|
||||
- V58/V62 evidence is calibration-only; at threshold 0.15 aggregate F1 is
|
||||
0.512905 while Flanders recall and F1 are 0;
|
||||
- no protected-test, background-test promotion or national-release evidence.
|
||||
|
||||
Claim boundary: exact/dHash and bbox-distance screens do not establish
|
||||
municipality, flight-strip, instance, semantic or imagery-edition independence.
|
||||
|
||||
Evidence: `tower-ml-data-lineage-snapshot.json` and
|
||||
`tower-key-artifact-hashes.json` (V62 best/last checkpoints and V56 tile audit).
|
||||
|
||||
## Backend tests
|
||||
|
||||
### Full root-context suite
|
||||
|
||||
```powershell
|
||||
$env:PYTHONPATH='C:\Projects\geointel\backend;C:\Projects\geointel'
|
||||
python -m pytest backend/tests -q -p no:cacheprovider \
|
||||
-W error::DeprecationWarning \
|
||||
--junitxml=artifacts/evidence/accuracy/P1/backend-full-suite.junit.xml
|
||||
```
|
||||
|
||||
Result: exit 1; 1,180 passed and 17 failed in 70.63 seconds. The 17 failures are
|
||||
stale source/contract assertions, including expectations that conflict with the
|
||||
new explicit user-selected model flow. They remain release-blocking until
|
||||
replaced with reviewed behavior tests.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `backend-full-suite.txt`
|
||||
- `backend-full-suite.junit.xml`
|
||||
|
||||
### Actual backend CI working directory
|
||||
|
||||
```powershell
|
||||
cd backend
|
||||
python -m pytest -W error::DeprecationWarning
|
||||
```
|
||||
|
||||
Result: exit 2 during collection; 1,194 items collected plus one import error:
|
||||
|
||||
```text
|
||||
ModuleNotFoundError: No module named 'scripts.render_operator_polygon_label_qa'
|
||||
```
|
||||
|
||||
Evidence: `backend-ci-entrypoint.txt`.
|
||||
|
||||
### Phase 1 tooling
|
||||
|
||||
```powershell
|
||||
python -m pytest tests/test_accuracy_phase1_baseline.py -q -p no:cacheprovider
|
||||
```
|
||||
|
||||
Result: 4 passed.
|
||||
|
||||
Evidence: `phase1-tooling-tests.txt`.
|
||||
|
||||
## Lint
|
||||
|
||||
```powershell
|
||||
python -m ruff check scripts/collect_accuracy_phase1_inference_smoke.py \
|
||||
scripts/collect_accuracy_phase1_ml_lineage.py \
|
||||
scripts/collect_accuracy_phase1_runtime.py \
|
||||
scripts/reproduce_accuracy_phase1_findings.py \
|
||||
scripts/run_accuracy_phase1_baseline.py tests/test_accuracy_phase1_baseline.py
|
||||
```
|
||||
|
||||
Result: all new Phase 1 code passed.
|
||||
|
||||
```powershell
|
||||
python -m ruff check backend scripts tests
|
||||
```
|
||||
|
||||
Result: exit 1; 112 findings: E402 13, E701 2, E702 69, F401 23, F403 1,
|
||||
F811 2 and F841 2.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `repository-ruff-baseline.txt`
|
||||
- `repository-ruff-baseline.json`
|
||||
|
||||
## Frontend
|
||||
|
||||
```powershell
|
||||
cd frontend
|
||||
npm run test:unit
|
||||
npm run typecheck
|
||||
npm run build
|
||||
```
|
||||
|
||||
Results:
|
||||
|
||||
- Vitest: 16 files, 51 tests passed;
|
||||
- TypeScript typecheck: passed;
|
||||
- production build: passed, 1,896 modules transformed.
|
||||
|
||||
The attempted generic `npm test -- --run` failed because no `test` script
|
||||
exists; the configured command is `test:unit`. The test strategy also calls
|
||||
for lint, but `npm run lint` fails with `Missing script: "lint"`.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `frontend-vitest.txt` (failed non-existent generic command, retained)
|
||||
- `frontend-vitest-unit.txt` (correct configured command, passed)
|
||||
- `frontend-typecheck.txt`
|
||||
- `frontend-build.txt`
|
||||
- `frontend-lint.txt`
|
||||
|
||||
## API, migrations and golden QA
|
||||
|
||||
```powershell
|
||||
python scripts/audit_api_contracts.py
|
||||
```
|
||||
|
||||
Result: 147 implemented routes match documentation; 10 explicitly tracked
|
||||
non-envelope endpoints.
|
||||
|
||||
Evidence: `openapi-contract-audit.txt`.
|
||||
|
||||
```powershell
|
||||
cd backend
|
||||
python -m alembic heads
|
||||
python -m alembic upgrade head --sql
|
||||
```
|
||||
|
||||
Result: one head, `202607260001`; complete offline upgrade rendered 496 SQL/log
|
||||
lines. This is not a local live-PostGIS migration test.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `alembic-heads.txt`
|
||||
- `alembic-offline-upgrade.sql`
|
||||
|
||||
```powershell
|
||||
python scripts/run_golden_qa_benchmark.py --json
|
||||
```
|
||||
|
||||
Result: two runs have semantically identical metric results but different file
|
||||
hashes because run identities use UUID4.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `golden-qa-run-1.json`
|
||||
- `golden-qa-run-2.json`
|
||||
- `golden-qa-reproducibility.json`
|
||||
|
||||
## Explicitly not executed as a success gate
|
||||
|
||||
- no protected test or background-test portfolio was opened;
|
||||
- no training, calibration fitting, checkpoint promotion or active-model change;
|
||||
- no label was accepted on behalf of a human reviewer;
|
||||
- no production database migration or data repair;
|
||||
- no dataset, checkpoint, cache, output or user-owned untracked file deleted,
|
||||
rewritten or moved;
|
||||
- no segmentation, SAM or solar asset was inferred to be production-ready from
|
||||
file presence.
|
||||
Reference in New Issue
Block a user