Post-baseline decision structure
High quality, weak-slice strength, Pareto efficiency, untapped convergence, or architecture-specific opportunity.
Strong individual quality plus empirically complementary errors.
Good baseline plus measured headroom, expertizable structure, and stable training behavior.
Evidence-qualified optimization
Expected ≈4–6 initial models. Depth may differ by model; a smaller subset can receive deeper optimization after early targeted experiments.
Quality + complementarity
Expected ≈2–4 models. Large models are allowed; predictions can be generated independently without simultaneous GPU residency.
Resource-aware substrate
Expected ≈1–2 models. The largest or highest-mAP detector is not automatically the best expertization target.
Tiny-object bottleneck
Test resolution, native multiscale, crop/tiling, or scale-aware augmentation.
Long-tail bottleneck
Test repeat-factor or class-aware sampling, class-aware augmentation, or loss weighting.
Crowding bottleneck
Test query count, resolution, matching settings, or crop strategy.
Localization bottleneck
Test regression-specific knobs, resolution, or geometry augmentation.