
Computer vision · live system
Greenhouse person detection and monitoring
live, my detector since 21 Sep 2026
Person detection and site control for my family's greenhouse, on a 4G line behind CGNAT. I trained the detector it now uses.
94.1% vs 49.8% person recall at a 3.8% false-positive rate, my detector against stock YOLO11s
Details
Problem
The family needs an alarm when a person is on site, day or night, and everything has to go over the 4G line. At the false-positive rate production ran at (3.8%), stock YOLO11s found 49.8% of people.

What I did
- The system: six cameras (one PTZ), an ESP32 climate node, automated irrigation, and a web app the family uses on their phones for alerts, live view and an event calendar. Detection runs on a GPU server at home.
- Went through a week of live alarms one by one (57 false, 61 true). It showed that my offline test could not fail the model: no frame in its negative set scored above the alarm threshold (the highest was 0.49, the threshold 0.52). The same audit turned up three code bugs, among them a check that failed open when the HD frame grab timed out.
- My first fine-tuned detectors beat stock YOLO only on a leaky split. One I briefly deployed was rolled back when its whole gain turned out to come from the part of the test set closest to its training data.
- What worked: a D-FINE-X teacher labelled 125,900 frames. After dropping uncertain ones and mining static scene clutter as hard negatives, a 19.4M-parameter D-FINE-M student trained on 13,668 of them in 70 minutes on an RTX 4060 Ti.
- It went live on 21 September 2026, after a shadow run that raised zero extra alarms (740 live polls and 8,505 archive frames).
- An adapter left the production code untouched, separate score calibrations kept false alarms on the old scale, and it falls back to the old model after three crashes.
- Before touching the 4G line I measured what lower resolution costs: at the same false-positive rate, recall was 95% at 1080p and 94% at 352x288. So detection moved to the DVR sub-stream.
- Built on-demand live view (HEVC to H.264 HLS) and debugged it on a real phone.
- Wrote a GPU lease: while a batch job has the card, live inference moves to the CPU, and it moves back afterwards, even if the job crashes.

Result
- On 2,918 frames I labelled myself, from days that do not overlap training: 94.1% of people found against 49.8% for stock YOLO11s, at a 3.8% false-positive rate. AUC 0.981 against 0.863. At night, both at a 0.50 threshold, 93.5% against 38.1%.
- 4G traffic went from 4.2 to 1.45 GB a day, and detection delay from 11 to 5 seconds.
- Switching cameras in live view went from up to 180 s to 16 s.
Limits
- The main test table comes from July and August frames. The September check set has 249 frames and only 15 positives so far.
- On DVR alarm frames the student reaches 58% recall against the teacher's 75% (48 positives).
- A follow-up model gained 0.0104 AUC on the September set, but its AUC spread 0.0180 across three training seeds, so I withdrew the claim and it did not ship.
- The fallback has already fired once: after three crashes on 24 September it switched to the old model, and the new one was back the same evening. I have not looked into the cause yet.
Tools
- Python
- PyTorch
- D-FINE
- ONNX Runtime
- OpenCV
- FFmpeg
- HLS
- ESP32
- systemd




