Canonical quality
Project mAP50 remains primary, with precision, recall, and per-class AP50 as supporting evidence.
Measured execution profile
Record Params(M), peak VRAM, headroom, throughput, physical-batch feasibility, and stability where measured.
Object-level pair audit
Record both correct, A only, B only, both fail, and differences in class and localization errors.
| Diagnostic | Use | Boundary |
|---|---|---|
| Parameter Efficiency | mAP50 / Params(M) | Diagnostic only; does not replace project mAP50. |
| VRAM Efficiency | mAP50 / peak training VRAM (GB) | Use only where peak VRAM was actually measured. |
| VRAM headroom | Measured remaining capacity | More informative than parameter count alone for expansion feasibility. |
| Throughput | Images/s or equivalent project record | Report alongside batch and precision context. |
| Stability | Qualification/training behavior | Keep separate; do not collapse into a synthetic score. |
Co-DINO
Remains in a different performance tier from all accepted challengers so far.
DEIM
Converged very early; project-best 0.563005, substantially below Co-DINO.
Stable-DINO R50
Stable matching did not improve the matched corrected DINO R50 baseline.
Native vs project metrics
Native metrics are diagnostic only. Project checkpoint selection is controlled by marine_competition.metrics.full_metrics.
A candidate remains interesting when it is not clearly dominated across both performance and resource cost. Use a multi-axis Pareto frontier—not raw mAP alone and not parameter efficiency alone.