Document remote-sensing YOLO candidate benchmark
GeoIntel CI / docs-smoke (push) Has been cancelled
GeoIntel CI / contract-smoke (push) Has been cancelled

This commit is contained in:
Codex
2026-07-07 23:13:36 +02:00
parent 9bd6752128
commit 4d2ba1ea5c
4 changed files with 75 additions and 0 deletions
+12
View File
@@ -7,6 +7,18 @@
# Changelog
## Sprint 134 External remote-sensing YOLO candidate benchmark (2026-07-07)
- Evaluated the Hugging Face `agademer/yolo-remote-sensing-photovoltaic` YOLOv8l detection checkpoint as an explicit operator-provided runtime model asset.
- Downloaded `yolo-remote-sensing-photovoltaic-v8l-solar-farms-and-cities-v20260331-detect-1000_epochs.pt` to the Tower runtime as `/app/models/yolo-remote-sensing-photovoltaic-v8l-detect-1000.pt`; the model file is not committed to Git.
- The model catalog exposes it as `yolo-remote-sensing-photovoltaic-v8l-detect-1000-pt` with SHA256 `242ff4ab889569278f0eb9fcd22eb2c4bf2a52e48d05d89cc7cfa7941165d203`.
- Live YOLO preflight loaded the model successfully with `status=ready`, `model_load_ok=true`, `manifest_valid=true`, `tile_paths_exist=true`, `will_download_models=false` and `will_run_inference=false`.
- Live 45-run dense QA matrix compared the external YOLOv8l candidate with `geointel-building-yolov8n-expanded160e50-pt` and `geointel-building-yolov8n-hardneg160r8e40-pt` across Geel, Mol, Turnhout, Retie and Kasterlee-bos.
- Result: the external candidate was very conservative and missed most dense GRB buildings. It scored F1 `0.0` on Geel, `0.010582010582010581` on Mol, `0.019070321811680575` on Turnhout and `0.0` on Retie, while expanded160e50 remained the dense-AOI winner.
- Live 27-run background matrix showed the external candidate was cleaner on Kasterlee-bos than local YOLOv8n candidates, with 1/2/5 detections at thresholds `0.25`/`0.15`/`0.05`, but it leaked 0/1/3 detections on Postel-bos and was therefore not uniformly cleaner than hardneg160r8e40.
- Decision: keep the model as runtime evidence only. It should not become the V1 default because recall is too low for operational extraction. The next pass should train a higher-capacity local model, starting from a stronger base and using the existing dense plus hard-negative benchmark gates.
- No API contract change, provider fetching, fake detections, model auto-provisioning, repository-stored weights or app-side model training behavior was introduced.
## Sprint 133 Hard-negative-balanced YOLO candidate (2026-07-07)
- Added `--background-negative-repeat` / `OPERATOR_YOLO_BACKGROUND_NEGATIVE_REPEAT` support to `scripts/export_operator_yolo_tile_dataset.py` so train-split background-candidate negative tiles can be repeated deterministically without duplicating validation tiles.