Back to the work

03Independent software · Desktop · Open source

LabelMate

A local AI companion for YOLO labeling: the model proposes boxes, the person keeps the decision.

Inference
Local · Ultralytics YOLO
Interface
Python · PySide6
Output
YOLO-format labels
Licence
MIT · open source

My contribution

Designed and built on my own. The tool grew out of a real labeling need during my workplace training at Boğaziçi Savunma; LabelMate is its generalised, public form. I built it to keep labeling local and offline—protecting the data—and to make the process faster.

03Recorded demo · the real applicationYOLOv8n · COCO
A real screen recording of the application: proposals arrive, the model takes a chest of drawers for an oven, and the user deletes the proposal and moves on. The images are from the COCO dataset and the model is a pretrained YOLOv8n.

01Problem

Drawing every box by hand.

An object detector learns only as well as its labels. Boxing every object in every image by hand is slow, repetitive work, and as the data grows, labeling becomes the heavy part.

LabelMate turns the order around: the model proposes first, and the person reviews and corrects. Images never leave the machine; inference runs locally with a YOLO model the user chooses.

01The same image · before and after24 / 128
The same image with five coloured proposal boxes and confidence scores around the cake, the cup and the table.
Cake and a coffee cup on a wooden table; the image is dimmed with “Running YOLO Assist…” in the middle.
Two frames from the real recording: while the model runs in the background, and once its five proposals arrive. The image is from the COCO dataset.

02Human and model

The model proposes; the person decides.

Every proposal is a starting point: accepted, corrected or deleted.

02Workflow01 / 04
The LabelMate window: three proposal boxes on a photo of cake; the image list on the left, the annotation panel on the right.

When an image opens, the model runs in the background; proposals arrive with a class and a confidence score (cake 0.48, dining table 0.25, fork 0.37).

The cake box selected, with handles on its corners and edges; the side panel shows class cake, confidence 0.482 and normalised centre and size values.

A selected box is moved and resized with eight handles; the side panel shows its class, confidence and normalised position.

A selected “oven 0.37” box around a chest of drawers; the side panel shows class oven and confidence 0.373.

The model takes a chest of drawers for an oven (oven 0.37). The proposal reaches a person before anything is saved.

The same image with the box deleted; the side panel shows 0 boxes.

The user deletes the proposal. With “Empty Label” on, moving to the next image saves this one with an empty label file that says “no object here”.

When an image opens, the model runs in the background; proposals arrive with a class and a confidence score (cake 0.48, dining table 0.25, fork 0.37).

The images are browsed from the keyboard, and an image with a saved label turns green in the list. The model can be wrong, so every proposal reaches a person before it is saved.

From the keyboard

Space
save, next
A / D
previous / next
Del
delete box
Ctrl+Z
undo
Ctrl+R
re-run assist

The recording uses a pretrained YOLOv8n on COCO images. No usage, time-saving or accuracy measurement was reported.

03Under the hood

The interface stays responsive; the model works in the background.

Every image takes the same path.

03One image’s pathPySide6 · QThread

When a model is chosen · once

  1. Model choiceModelDialog.pt / .onnx, or a preset YOLO11 / YOLOv8 download
  2. LoadingYOLOModel.loadUltralytics YOLO
  3. Warm-upWarmupWorkera first inference on a blank 640×640 image

For every image

  1. Image opens_open_imagelist and canvas update
  2. Saved label?read_labelif so it loads; the model does not run again
  3. InferenceInferenceWorkera separate thread (QThread)
  4. Proposalsresults_readya Qt signal; dropped if another image is open
  5. Person editsCanvasselect · move · resize · delete · draw
  6. Savewrite_labelon Space, or when the image changes
  7. labels/<name>.txtnormalised YOLO lines
Names from the repository’s code: app/main_window.py, app/workers.py, app/yolo_inference.py, app/label_io.py. The flow is explanatory; it contains no timing measurement.

04Canvas

A move on screen, a change in the data.

The canvas works in three coordinate planes: the screen, image pixels, and the normalised values in the file. A mouse move is first mapped to image pixels, the edit happens there, and the result is normalised to the 0–1 range.

Boxes are clamped to the image and their corners sorted, so a box drawn backwards stays valid. Zooming keeps the point under the cursor in place.

Before every edit, a full copy of the boxes goes onto the undo stack: up to 50 steps back and forward.

04Three coordinate planesExplanatory
Explanatory drawing; the transforms are from app/canvas.py and app/bbox.py. The example values are those the side panel shows for the selected cake box in the recording.

05Output

From editing to training data.

One text file per image, one box per line. The class number, centre and size are normalised to the image; the confidence score is not written to the file.

Class names come from classes.txt in the output folder, or from the model. Optionally, the images are copied into the output folder too.

05YOLO label fileYOLO · TXT

Output folder

  • output_folder/
  • labels/
  • image001.txt
  • image002.txt
  • images/optional
  • classes.txt

One line, one box

00.5123000.4382000.0341000.028900

class: 0 · centre x: 0.512300 · centre y: 0.438200 · width: 0.034100 · height: 0.028900

Folder structure and example line from the repository’s documentation; values lie in 0–1 and are written with six decimals.

06Outcome

From an internal need to an open tool.

  1. 01

    Local

    Images never leave the machine; inference runs on the user’s computer with the chosen YOLO model.

  2. 02

    Person in control

    The model only proposes; saving, correcting or deleting is the person’s call.

  3. 03

    Responsive

    Warm-up, inference and model downloads run in background threads.

  4. 04

    Open

    MIT licence; a Windows setup script that creates a virtual environment and a desktop shortcut.

Scope: optimised for single-class workflows, with multi-class through classes.txt. Image folders only; video is not supported. GPU use is left to Ultralytics. No usage, time-saving or platform-test measurement was reported.

Independent project · 2026 · MIT licence

Contact

Let’s buildthe next system together.

As an engineer focused on turning AI models into systems that work in the real world, I’m open to new opportunities and technical collaborations. If you work on computer vision, edge AI or applied machine learning, I’d be glad to connect.

Muhammed Ali Yıldırım

Applied AI / ML Engineering

Back to top