Skip to content
Alperen Uğur Erden

Alperen Uğur Erden

I work on computer vision and machine learning, mostly object detection and model evaluation, with PyTorch and OpenCV.

Looking for ML / computer vision internship or part-time remote work

Physics undergrad at Bilkent, graduating 2028.

Turkey (UTC+3)alperenugur04@gmail.com

I build systems that have to run when nobody is watching. The main one watches my family's greenhouse over a 4G line, and the person detector I trained for it has been live since 21 September 2026.

Getting there was mostly catching my own bad measurements. My first detectors only looked better because validation leaked, and one I deployed was rolled back.

I also run thelecturenotes.com alone, a course site for Bilkent students. It rebuilds every day, and the daily build deploys only when its browser tests pass.

Case study: the greenhouse detector

What happened between my first fine-tuned detector and the one in production, in order. Most of the steps are about measurement.

  1. Julyfound

    A test that could not fail

    I went through a week of live alarms one by one (57 false, 61 true). No frame in my offline negative set scored above the alarm threshold (the highest was 0.49, the threshold 0.52), so that test could never fail. The same audit turned up three code bugs.

  2. Augustrolled back

    Better only on a leaky split

    My first fine-tuned detectors beat stock YOLO only on a leaky split. One I briefly deployed was rolled back when its whole gain turned out to come from the part of the test set closest to its training data.

  3. Septembertrained

    Teacher and student

    A D-FINE-X teacher labelled 125,900 frames. After dropping uncertain ones and mining static scene clutter as hard negatives, a 19.4M-parameter D-FINE-M student trained on 13,668 of them in 70 minutes on an RTX 4060 Ti.

  4. 21 Sepshipped

    Live, after a shadow run

    The shadow run raised zero extra alarms (740 live polls and 8,505 archive frames). An adapter left the production code untouched, separate score calibrations kept false alarms on the old scale, and it falls back to the old model after three crashes.

  5. 23 Sepwithdrawn

    A gain I took back

    A follow-up model gained 0.0104 AUC on the September set, but its AUC spread 0.0180 across three training seeds, so I withdrew the claim and it did not ship.

  6. 24 Sepfell back

    The fallback fired

    After three crashes it switched to the old model, and the new one was back the same evening. I have not looked into the cause yet.

Person recall: stock YOLO11s against my detector

Bar chart. On all 2,918 test frames, at a 3.8% false-positive rate, stock YOLO11s found 49.8% of people and my detector 94.1%. On night frames, both at a 0.50 threshold, stock YOLO11s found 38.1% and my detector 93.5%.All test frames, at a 3.8% false-positive rateStock YOLO11s49.8%My detector (D-FINE-M)94.1%Night frames, both at a 0.50 thresholdStock YOLO11s38.1%My detector (D-FINE-M)93.5%
Note 2,918 frames I labelled myself, from July and August days that do not overlap training. Source: Greenhouse

4G traffic per day

Bar chart. 4G traffic was 4.2 GB a day with frames from the main stream and 1.45 GB a day after detection moved to the DVR sub-stream.Before: frames from the main stream4.2 GBAfter: frames from the DVR sub-stream1.45 GB
Note Measured before the switch: at the same false-positive rate, recall was 95% at 1080p and 94% at 352x288. Detection delay went from 11 to 5 seconds. Source: Greenhouse

Other results

Each number comes with its caveat. The full write-ups are in the project entries below.

Player tracking: IDF1 on a 2-minute window

Bar chart of IDF1 on a 2-minute window with 18 identities: raw tracklets 0.515, online BotSort 0.574, my offline stitching 0.835. The score is optimistic.Raw tracklets0.515Online BotSort0.574My offline stitching0.835
Note Optimistic: the rejected identities, all in the crowded far-goal area, are left out of the ground truth. Over a full 32-minute match the same method fragments into 698 identities. Source: Football tracking

Radiomics, my analysis: AUC against a label-shuffle control

Bar chart of AUC under patient-grouped cross-validation. Label-shuffle control over 753 permutations 0.513, close to chance. Mean density alone 0.776. Texture random forest 0.985 plus or minus 0.018.Label-shuffle control (753 permutations)0.513Mean density alone0.776Texture random forest0.985 ± 0.018chance
Note Data and clinical question are my collaborator's. 180 lesions from 120 patients, single centre, p = 0.0013 against the shuffle control. Location and scan protocol alone reach AUC 0.96 where recorded, so it needs external validation. Source: Radiomics analysis

gemma3:12b on 80 yes/no constraint checks

Unit chart, one square per check. Asked as a yes/no judgement, gemma3:12b got 27 of 80 wrong. Asked to translate each constraint to Python for code to check, it got 0 of 80 wrong.Yes/no judgement27 of 80 wrongTranslated to Python, checked by code0 of 80 wrongwrongright
Note Answers are computed in code. On the Python route qwen2.5:14b missed 3. The translation idea is not new: it is in PAL and NL4Opt. Source: LLM arithmetic

Projects

My own projects, not client work. Most of them run on servers I set up and look after myself.

Two phone screens from the family web app, in Turkish: a 3D map of the site with numbered camera pins and live temperature and humidity, and a month calendar of detected visits.

Computer vision · live system

Greenhouse person detection and monitoring

live, my detector since 21 Sep 2026

Person detection and site control for my family's greenhouse, on a 4G line behind CGNAT. I trained the detector it now uses.

94.1% vs 49.8% person recall at a 3.8% false-positive rate, my detector against stock YOLO11s

Details

Problem

The family needs an alarm when a person is on site, day or night, and everything has to go over the 4G line. At the false-positive rate production ran at (3.8%), stock YOLO11s found 49.8% of people.

Diagram of the greenhouse system. At the site, six cameras (five on a DVR, one PTZ), an ESP32 climate node and a smart plug for the irrigation pump connect to an old Android phone as the edge node, behind a 4G router with no inbound ports. A home server with an RTX 4060 Ti pulls from it over Tailscale, with a cloud VM as fallback, and runs ingest, the D-FINE-M detector, a DVR alarm referee, on-demand live view and a state API. The family app reaches it through a Cloudflare Tunnel and Worker.
How the data moves. Moving the camera feed to the DVR sub-stream cut 4G traffic from 4.2 to 1.45 GB a day. Open full size

What I did

  • The system: six cameras (one PTZ), an ESP32 climate node, automated irrigation, and a web app the family uses on their phones for alerts, live view and an event calendar. Detection runs on a GPU server at home.
  • Went through a week of live alarms one by one (57 false, 61 true). It showed that my offline test could not fail the model: no frame in its negative set scored above the alarm threshold (the highest was 0.49, the threshold 0.52). The same audit turned up three code bugs, among them a check that failed open when the HD frame grab timed out.
  • My first fine-tuned detectors beat stock YOLO only on a leaky split. One I briefly deployed was rolled back when its whole gain turned out to come from the part of the test set closest to its training data.
  • What worked: a D-FINE-X teacher labelled 125,900 frames. After dropping uncertain ones and mining static scene clutter as hard negatives, a 19.4M-parameter D-FINE-M student trained on 13,668 of them in 70 minutes on an RTX 4060 Ti.
  • It went live on 21 September 2026, after a shadow run that raised zero extra alarms (740 live polls and 8,505 archive frames).
  • An adapter left the production code untouched, separate score calibrations kept false alarms on the old scale, and it falls back to the old model after three crashes.
  • Before touching the 4G line I measured what lower resolution costs: at the same false-positive rate, recall was 95% at 1080p and 94% at 352x288. So detection moved to the DVR sub-stream.
  • Built on-demand live view (HEVC to H.264 HLS) and debugged it on a real phone.
  • Wrote a GPU lease: while a batch job has the card, live inference moves to the CPU, and it moves back afterwards, even if the job crashes.
Four phone screens from the family web app, in Turkish: a 3D map of the site with numbered camera pins and live temperature and humidity, a month calendar of detected visits, the climate tab with a 24-hour temperature chart and a frost check for the night, and the irrigation calendar.
The family web app. Camera views are left out on purpose. Open full size

Result

  • On 2,918 frames I labelled myself, from days that do not overlap training: 94.1% of people found against 49.8% for stock YOLO11s, at a 3.8% false-positive rate. AUC 0.981 against 0.863. At night, both at a 0.50 threshold, 93.5% against 38.1%.
  • 4G traffic went from 4.2 to 1.45 GB a day, and detection delay from 11 to 5 seconds.
  • Switching cameras in live view went from up to 180 s to 16 s.

Limits

  • The main test table comes from July and August frames. The September check set has 249 frames and only 15 positives so far.
  • On DVR alarm frames the student reaches 58% recall against the teacher's 75% (48 positives).
  • A follow-up model gained 0.0104 AUC on the September set, but its AUC spread 0.0180 across three training seeds, so I withdrew the claim and it did not ship.
  • The fallback has already fired once: after three crashes on 24 September it switched to the old model, and the new one was back the same evening. I have not looked into the cause yet.

Tools

  • Python
  • PyTorch
  • D-FINE
  • ONNX Runtime
  • OpenCV
  • FFmpeg
  • HLS
  • ESP32
  • systemd
thelecturenotes.com on a laptop and a phone. The laptop shows a MATH 101 study-notes page with a plot of secant lines closing in on a tangent. The phone shows the schedule builder with a week of sections and zero clashes.

Web platform · live

thelecturenotes.com

live since May 2026, run alone

A course site for Bilkent students, with about 3,000 past exams and notes and a schedule builder.

1,286 Google clicks in the four weeks to 22 Sep 2026, up from 122, registration season included

site thelecturenotes.com

Details

Problem

The site shows sections and seat counts, and during registration they are only useful if they are current. I run it alone, so the checks run every day without me.

What I did

  • A page for every course, about 3,000 past exams and notes, and a schedule builder that flags section clashes and suggests electives that fit around required courses. Most of the material was imported in bulk from existing archives; new uploads go through a moderation queue.
  • An LLM classifier that reads file content to catch material filed under the wrong course, and a vision step that transcribes scanned pages.
  • A daily pipeline pulls the university's public course and section data (972 courses, 1,971 sections this term), rebuilds 3,548 static pages, runs a 39-suite browser and data test gate, and deploys only when it passes.
  • Moved the Django 5 + PostgreSQL API off Fly.io onto my own server behind a Cloudflare Tunnel. Static pages are on Cloudflare Pages.
  • Used Search Console to find the problems: the wrong landing page, truncated titles, and course pages that did not link to each other. The median course page went from 0 to 7 internal links, and dead internal links from 112 to 0.
thelecturenotes.com on a laptop and a phone. The laptop shows a MATH 101 study-notes page with a plot of secant lines closing in on a tangent. The phone shows the schedule builder with a week of sections and zero clashes.
Left: LLM-drafted study notes. Right: the schedule builder. Open full size

Result

  • Search Console, four weeks to 22 September 2026: 1,286 clicks and 36,584 impressions, against 122 and 2,653 the four weeks before. Registration season is part of that jump.
  • The test gate has already stopped a deploy that carried two real bugs.
  • Found and fixed a caching bug that had frozen seat and instructor data for about four weeks: on registration day the site showed sections 58.6% full overall when the real figure was 89%.
The schedule builder on a desktop screen: five courses laid out on a Monday to Friday grid, with totals for credits, ECTS and weekly hours above it, and zero clashes.
A five-course week in the schedule builder, with no clashes. Open full size

Limits

  • Study notes for 6 courses (75 sections) are LLM-drafted. I built the pipeline and the pages, not the maths in them, and the maths has not been independently checked.
  • The browser test gate runs in the daily pipeline. A deploy triggered by a push runs only the backend tests.

Tools

  • Python
  • Django
  • PostgreSQL
  • Docker
  • GitHub Actions
  • Cloudflare Pages
  • Cloudflare R2
Top-down map of a football pitch with each tracked player drawn as a red dot and the camera position marked in one corner.

Computer vision · research prototype

Player tracking for amateur football

paused since August 2026

Turns match video from fixed venue cameras into one track per player. Over a full 32-minute match it still breaks into 698 identities.

IDF1 0.835 on a 2-minute window, against 0.574 for BotSort

Details

Problem

The venue films matches with fixed fisheye cameras. The goal was one track per player, placed on the pitch.

What I did

  • Fisheye correction and pitch calibration from line detection, then player detection with RF-DETR.
  • Offline tracklet stitching with appearance (OSNet re-ID), speed and time-overlap constraints.
  • Built the identity ground truth by having LLM vision agents review tracklet image cards in several passes, then ran a round that tried to reject every merged identity. 7 of 16 were rejected.
Top-down map of a football pitch with each tracked player drawn as a red dot and the camera position marked in one corner.
Player positions after fisheye correction and the pitch homography, drawn on a top-down pitch. Open full size

Result

On a 2-minute window with 18 identities, offline stitching reaches IDF1 0.835, against 0.515 for raw tracklets and 0.574 for online BotSort on the same window.

Limits

  • The score is optimistic. The rejected identities, all in the crowded far-goal area, are left out of the ground truth, and its first grouping used the same appearance threshold as the stitcher.
  • Over a full 32-minute match the same method fragments into 698 identities, and the 16 largest cover only 40%.
  • Calibration on an unseen pitch still needs a manual step. The second camera is aligned (0.77 m) but only partly fused.

Tools

  • Python
  • PyTorch
  • OpenCV
  • RF-DETR
  • OSNet re-ID
  • homography
  • MOT metrics
Four robustness panels from the manuscript draft: a confound bound for each texture feature; AUC of a logistic model and a random forest before and after removing the recorded scan protocol, on the 110 lesions where it was recorded; AUC as a growing share of labels is randomly flipped; and a learning curve over the number of training lesions.

Medical imaging ML · my analysis

Analysis for a collaborator's radiomics study

analysis done, manuscript draft

The clinical question and the lesion data are my collaborator's. I built the classification and validation, caught a patient-level leak, and drafted the manuscript.

AUC 0.985 patient-grouped nested CV; mean density alone 0.776

Details

Problem

Sclerotic bone lesions on CT, in a breast cancer cohort, have to be sorted into metastases and bone islands. My collaborator's question was whether texture features add anything over mean density. The data: 180 lesions from 120 patients, single centre, contoured by two radiologists.

What I did

  • Built and compared the classification models on the radiomic features, under patient-grouped nested cross-validation.
  • Regrouped the folds by patient. The first version split per lesion, and 31 of 120 patients contribute more than one lesion, so the same patient could sit on both sides of a split.
  • Ran a label-shuffle control over 753 permutations, and checked how well recorded lesion location and scan protocol alone separate the two groups.
  • Drafted the manuscript.

Result

  • Random forest AUC 0.985 ± 0.018 with patient-grouped folds. Mean density alone reaches 0.776.
  • The per-lesion version had reported 0.989. The leak was small here, but only the grouped number is defensible.
  • The label-shuffle control averages 0.513 (p = 0.0013).

Limits

  • Single centre. Where it was recorded, lesion location and scan protocol alone separate the two groups at AUC 0.96, so that confound is bounded but not ruled out. The result needs external validation.
  • The mean-density baseline depends on the method: 0.776 under nested cross-validation, 0.886 in a bootstrap panel.
Four robustness panels from the manuscript draft: a confound bound for each texture feature; AUC of a logistic model and a random forest before and after removing the recorded scan protocol, on the 110 lesions where it was recorded; AUC as a growing share of labels is randomly flipped; and a learning curve over the number of training lesions.
Robustness checks from my analysis, in the manuscript draft. Removing the recorded protocol hurts the linear model far more than the random forest, which bounds the confound but does not rule it out. Open full size

Tools

  • Python
  • scikit-learn
  • PyRadiomics
  • nested cross-validation
  • permutation tests

LLM evaluation · research

Where local LLMs get arithmetic wrong

experiments ongoing, September 2026

Experiments on local models from 3B to 32B, on arithmetic tasks whose answers are computed in code.

27 of 80 → 0 of 80 gemma3:12b errors, yes/no judgement against translation to Python

Details

Problem

I measured how often local LLMs get it wrong when asked whether a numeric constraint holds, and whether asking a different way helps.

What I did

  • Ran local models from 3B to 32B (gemma3, qwen2.5, qwen3, served with Ollama).
  • Asked each constraint two ways: as a yes/no judgement, and as a translation into a Python expression that code then evaluates.
  • For word problems, had the model label what each number means and compiled the constraints in code, and compared that with asking it to write the constraints directly.
  • Sampled 400 points from a 7-axis space of 2,430 task configurations to map where each model fails.

Result

  • Yes/no: gemma3:12b got 27 of 80 wrong. Translated to Python: it got all 80 right, and qwen2.5:14b missed 3.
  • Word problems: labelling and compiling beat writing the constraints directly, Jaccard 0.44 to 0.63 over 30 paraphrased problems.
  • gemma3:12b and qwen2.5:14b fail in mostly the same places (r = 0.917 between their pass-rate maps). qwen2.5:3b fails somewhere else (r 0.55 to 0.65).

Limits

  • Scoring the same collected data with 7 evaluators gave 4 different winners, so I don't quote a sampling-method gain without naming the evaluator.
  • None of the methods is new: the translation idea is in PAL and NL4Opt, the boundary sampling in Bryan et al. (2005).

Tools

  • Python
  • Ollama
  • Hugging Face Transformers
  • CUDA
Home page of yurttanayriliyorum.com, in Turkish: a headline asking what happens to your things when you leave the dorm, a school search box, and a row of university logos.

Web app · live

yurttanayriliyorum.com

live, last code change August 2026

A move-out marketplace for students, set up for 50 Turkish campuses and built solo on Cloudflare Workers and D1.

97 routes one Cloudflare Worker, about 4,900 lines

site yurttanayriliyorum.com

Details

Problem

Students moving out of dorms need a way to pass things on to other students at the same school, with sign-in limited to students.

What I did

  • One Cloudflare Worker (about 4,900 lines, 97 routes) on D1, with a subdomain per school, set up for 50 Turkish campuses.
  • Student-email OTP sign-in, with codes stored only as SHA-256 hashes keyed with a server secret, and rate limits.
  • KVKK notices (Turkey's data-protection law).
Home page of yurttanayriliyorum.com, in Turkish: a headline asking what happens to your things when you leave the dorm, a school search box, and a row of university logos.
The home page. Sign-in is a one-time code sent to a student email address. Open full size

Limits

I don't publish usage numbers for it, because the ones I have are not verified.

Tools

  • JavaScript
  • Cloudflare Workers
  • D1
  • SQL

More projects

LLM agents · code on GitHub

llm-agent-runtime: autonomous agent daemon

paused since June 2026

An agent daemon that splits a goal into work packets and can change files on the machine; most of the code limits what it can do.

8 LLM backends from five providers; shell commands run in a bubblewrap sandbox

code llm-agent-runtime

Details

Problem

An agent that edits files and runs shell commands on a real machine needs limits that hold when the model gets things wrong.

What I did

  • A goal is split into work packets, and each packet is routed to a panel of eight LLM backends from five providers (five APIs and three CLI or subscription wrappers). The executor can actually change files on the machine.
  • Each packet gets its autonomy level when the goal is decomposed, before anything runs.
  • Shell commands run inside a bubblewrap sandbox with a writable-path allowlist, and the MCP client refuses write tools, such as sending email or creating files, unless the call is explicitly marked as a write.
  • A rule-based self-audit checks telemetry and repo state without calling an LLM.

Tools

  • Python
  • asyncio
  • SQLite
  • bubblewrap
  • multi-provider routing
  • sandboxed execution

LLM pipeline · code on GitHub

papers-net: multi-model research pipeline

stopped in July 2026

A pipeline that sends one research problem to several LLM providers, scores their answers and logs every run.

9,403 runs May to July 2026, 4,817 successful

code papers-net-core

Details

Problem

I wanted to compare LLM providers on the same research problems and keep only the answers that survived scoring.

What I did

  • Each problem goes to several providers. Their answers are scored and merged, and the surviving results are kept as material for papers.
  • Every run is logged in SQLite, with a dashboard for run history and provider comparison.

Result

9,403 runs between May and July 2026 across five provider families (OpenAI, Google, Anthropic, DeepSeek, OpenRouter), 4,817 of them successful.

Limits

4,586 runs failed, almost half.

Tools

  • Python
  • SQLite
  • FastAPI
  • multi-provider orchestration

Embedded Linux · code on GitHub

Android 10 on a 2012 phone

shelved after August 2026

LineageOS 17.1 (Android 10) ported to a 2012 Galaxy S3 mini with a Linux 3.4 kernel and 2013 vendor blobs.

about 124 s to a full boot, after tracking down 10 root causes

code old-kernel-modern-android

Details

Problem

Get Android 10 running on a phone whose kernel and vendor blobs predate it by years, then see whether the phone can run detection on-device.

What I did

  • Ported LineageOS 17.1 and tracked down 10 root causes on the way, among them a vendor kernel patch that randomised every mmap and fragmented the address space.
  • Measured on-device detection with YOLO-FastestV2 int8.

Result

  • A full boot in about 124 s.
  • Detection ran at 252 ms per frame.

Limits

The detector found 1 of 20 players in a football frame, so it is not usable for tracking.

Tools

  • AOSP build
  • init
  • HAL
  • binder
  • Linux kernel debugging

Other

Firmware-level diagnosis of a read-only NVMe SSD

not recovered

A 1 TB NVMe drive locked itself read-only. I traced the cause to its boot die, one of 16 NAND dies, which fails erases and pushes the firmware into permanent write protection. Not recovered: the fix needs vendor firmware I don't have.

  • Linux
  • NVMe
  • PCIe
  • vfio
  • reverse engineering

Skills

Computer vision and ML

  • Computer Vision
  • Object Detection
  • Deep Learning
  • Machine Learning
  • Model Evaluation
  • Image Processing
  • Data Annotation
  • Statistics

Used in: Greenhouse, Football tracking, Radiomics analysis, LLM arithmetic

ML libraries and tools

  • PyTorch
  • OpenCV
  • ONNX Runtime
  • scikit-learn
  • NumPy
  • Pandas
  • FFmpeg

Used in: Greenhouse, Football tracking, Radiomics analysis

LLMs

  • Large Language Models
  • AI Agents
  • Prompt Engineering
  • Ollama

Used in: LLM arithmetic, llm-agent-runtime, papers-net, thelecturenotes.com

Backend and infrastructure

  • Django
  • FastAPI
  • REST APIs
  • PostgreSQL
  • SQL
  • Docker
  • Linux
  • systemd
  • Git
  • GitHub Actions
  • Cloudflare Pages, Workers, D1, R2

Used in: thelecturenotes.com, yurttanayriliyorum.com, Greenhouse

Languages

  • Python
  • SQL
  • JavaScript
  • Bash
  • LaTeX

Education

Bilkent University

B.S. in Physics · 2022–2028 (expected)

Relevant coursework: Artificial Intelligence (CS461), Machine Learning (CS464), Probability and Statistics (MATH230), Programming in Python (CS115), Advanced Calculus for Physics (MATH242).

  • Bilkent Comprehensive Merit Scholarship (full)
  • TÜBİTAK Undergraduate Scholarship in Basic Sciences
  • TEV Outstanding Achievement Scholarship

Papers

  • Texture Radiomics Outperforms Mean CT Attenuation for Sclerotic Bone Lesions

    Manuscript draft. The clinical question and data are my collaborator's; I did the analysis and drafted the text.

  • Minimizers of the Seidel Quadratic Form: Exact Optimum Counts for n ≤ 14

    Sole author, preprint.

    doi 10.5281/zenodo.22018437