Retuning a motorcycle-helmet compliance pipeline for a gate camera where the road — and the riders on it — are often forty pixels tall.
Config-only, in-place tuning of the existing detect → track → associate → count loop — no architecture rewrite, no retraining. Each step is an .env knob, reversible on its own.
A rider or helmet only ever matters here as an attribute of a motorcycle — never tracked as its own object. Splitting detection into motorcycle-first, rider-second, helmet-third means a frame with no motorcycle at all skips two entire inference calls, not just their results.
On a fixed gate camera, half the view is often sky, building, or the wrong side of the road. DETECTION_ZONE crops inference to one half before the model ever sees the rest, cutting cost roughly in half and removing a source of false positives outright — drawn on-screen as a red boundary line.
A motorcycle on screen is, definitionally, being ridden. If the person detector missed the rider in every frame of a track's life, the count previously read 0 riders, 0 helmets, 0 violations — silently correct-looking, and silently wrong. Flooring the rider count to at least one fixes it.
Step 01's gating, drawn as the decision it actually is — the branch that determines whether two more model calls happen at all.
| Setting | V5 | V6 |
|---|---|---|
| Model | yolov8n.pt | yolov8s.pt |
| Inference size | 640px | 960px |
| Confidence / IoU | 0.35 / 0.50 | 0.25 / 0.45 |
| Helmet confidence | 0.35 | 0.20 |
| ByteTrack high / low / new | 0.50 / 0.10 / 0.60 | 0.40 / 0.05 / 0.50 |
| Min frames to count | 5 | 3 |
| Rider/helmet overlap floor | 0.25 (hardcoded) | 0.15 |
| Min detection box area | — none — | 9px² |
| Contrast preprocessing | — none — | CLAHE |
| Rider/helmet tracked by ID | yes, with motorcycle | no — motorcycle only |
| Detection region | full frame | full / half-left / half-right |
| Min rider assumption | — none — | 1 per motorcycle |
| Dev-machine throughput | n/a | 4.8 → 12.1 FPS (cpu → mps) |