FUTURE · ALL CANONICAL MODELS · CROSS-MODEL EVIDENCE GATE

Full Shared Audit — All Accepted Canonical Models

Audit every accepted canonical model under the frozen project evaluator after the planned baseline queue completes—not only the current top three.

ALL ACCEPTED CANONICAL MODELSPROJECT MAP50 AUTHORITYNO MASTER SCORE
project mAP50precision / recallper-class AP50tiny / small / medium / largecrowding / densitylong-tail / rare classesbackground false positivesmissesclass substitutionslocalizationduplicate predictionsparameter countpeak training / inference VRAMphysical-batch feasibility / headroomthroughput / training stabilitypairwise both-correct / A-only / B-only / both-fail
ACCURACY

Canonical quality

Project mAP50 remains primary, with precision, recall, and per-class AP50 as supporting evidence.

RESOURCES

Measured execution profile

Record Params(M), peak VRAM, headroom, throughput, physical-batch feasibility, and stability where measured.

COMPLEMENTARITY

Object-level pair audit

Record both correct, A only, B only, both fail, and differences in class and localization errors.

Resource-efficiency diagnostics
DiagnosticUseBoundary
Parameter EfficiencymAP50 / Params(M)Diagnostic only; does not replace project mAP50.
VRAM EfficiencymAP50 / peak training VRAM (GB)Use only where peak VRAM was actually measured.
VRAM headroomMeasured remaining capacityMore informative than parameter count alone for expansion feasibility.
ThroughputImages/s or equivalent project recordReport alongside batch and precision context.
StabilityQualification/training behaviorKeep separate; do not collapse into a synthetic score.
KEY FINDING

Co-DINO

Remains in a different performance tier from all accepted challengers so far.

KEY FINDING

DEIM

Converged very early; project-best 0.563005, substantially below Co-DINO.

KEY FINDING

Stable-DINO R50

Stable matching did not improve the matched corrected DINO R50 baseline.

METRIC POLICY

Native vs project metrics

Native metrics are diagnostic only. Project checkpoint selection is controlled by marine_competition.metrics.full_metrics.

A candidate remains interesting when it is not clearly dominated across both performance and resource cost. Use a multi-axis Pareto frontier—not raw mAP alone and not parameter efficiency alone.