Conversation
Drop the two full-val (5000-image) reference runs from this first batch. They are valid, but a full-val record for the same model/hardware/runtime/precision makes the site prefer that record for the model's leaderboard row, which would move yolov9t and rfdetr-s onto the Spark full-val accuracy. Keep the canonical mini500 runs that the protocol asks for latency. The full-val references are kept locally and available on request: yolov9t 38.09, rfdetr-s 52.93 (they reproduce the project's Port Fidelity values within 0.1). Rebuilt generated/verified-results.v1.json.
|
Review notes on this first batch, written before you dig in. Everything below comes from reading the branch, re-running the repo validators locally and cross-checking the existing corpus. I have not re-measured on the Spark. What matches the protocol
Two things I deliberately held back, and would like your call on
Hardware metadata
CI Both workflows on this branch sit at Not touching further models until you confirm the format and the hardware-metadata question. |
--limit. Harness defaults kept: batch 1, conf 0.001, IoU 0.6, max_det 300, native input sizes 640 / 512.c25f6dffb521ea60bc0f63ae3dffb168a7edc466, the tagged release commit) and clean harnessf55a528cdd714259405f08b22c44d32be4fc7ab7; the PR proposes adding that release commit to the support matrix.Unknownandgpu_memory_gb: 0.0for unified memory. These fields are unchanged and need maintainer review.validate_submission.pyandbuild_verified_results.pypassed. Exactly two new records; no existing record was dropped.