한국어

KNOWLEDGE · OBJECT DETECTION

How many looks does it take to find every object in a photo?

YOLOv1 divides a photo into a 7×7 grid and passes it through a neural network just once to predict where every object is and what it is, all at the same time.

Detectors before YOLO proposed about 2,000 candidate regions where an object might be and checked them one by one. YOLO replaced that whole process with a single prediction, which is where the name comes from: You Only Look Once. Work through four interactive scenes to see how that is possible.

  1. 01 GridWho finds the object
  2. 02 ArchitectureWhat does the computing
  3. 03 InferenceFrom 98 boxes to a few
  4. 04 TrainingWhat gets penalized

01 Grid

The cell where an object's center falls is responsible for it

The photo is divided into 7 cells across and 7 down, 49 cells in all. The one cell that contains an object's center is responsible for that object. Each cell outputs 2 boxes and 20 class probabilities: 30 numbers in total.

Drag the dog or the bicycle. With a keyboard, select the picture, move with the arrow keys (Shift moves one cell at a time), and press Space to switch objects.

Selected object

dog → row 5, column 3

x, y (within the cell)
0.10, 0.48
w, h (relative to the image)
0.36, 0.30
√w, √h (what is predicted)
0.60, 0.55

Which of this cell's 30 numbers get filled

1–5: box 16–10: box 211–30: 20 classes

Of the two boxes, the one that overlaps the ground truth more takes the object, and its confidence learns that overlap (IOU). The other box and the other 48 cells learn a confidence of 0.

Source: paper §2, Figure 2, §2.4

02 Architecture

How a 448 × 448 image becomes 7 × 7 × 30

24 convolutional layers extract features while shrinking the width and height to 1/64, and 2 fully connected layers produce the 7 × 7 × 30 numbers. The first 20 layers are pretrained on ImageNet image classification; the remaining 4 convolutional layers and the fully connected layers are added for detection.

Each block is the output of one layer. The square face is its width and height; its thickness is the number of channels (both on a log scale).

24 conv layers · 2 fully connected · 271,703,550 parameters · about 20.3 billion multiply-adds per image

Width and height shrink 64-fold, from 448 to 7, while the channels grow from 3 to 1024. Pick a layer, and lines show which window of its input passes through every channel to become one cell of its output.

Drag to rotate; zoom with the wheel or two fingers. Click a block to go to that layer. Keyboard: arrows rotate, +/− zoom, [ ] move between layers.

First 20 layers, pretrained 4 conv layers added for detection Fully connected Pooling, input and output

0 / 31 · Input

A 448 × 448 image with 3 RGB channels. The original photo is resized to this size.

Source: paper §2.1, §2.2, Figure 3 · Numbers counted on this architecture implemented in PyTorch

03 Inference

How 98 boxes shrink to a few

One pass through the network gives 2 boxes per cell, 98 in all. A box's score is class probability × confidence. Boxes with low scores are dropped, and among boxes that overlap on the same object, NMS (non-maximum suppression) keeps only the one with the highest score.

Raise the score threshold, then turn on NMS.

Network output98
Above threshold98
Final98

These boxes are real outputs from a YOLO implemented from scratch following the paper (a small version with 8 conv layers). It was trained on pictures of squares, circles, and triangles and run on scenes it had never seen.

Source: paper §2, Eq. 1, §2.3

04 Training

Which mistakes cost the most

Training minimizes a sum of squared errors over five things: box center, box size, the confidence of boxes with objects, the confidence of boxes without objects, and the class. Used as is, though, a plain sum of squares goes wrong in two places.

Size error: the same error hurts small boxes more

Δw is the difference between the predicted width and the ground-truth width w. It is not a difference of square roots: it is how far off w itself is, with the image width taken as 1. The same Δw makes the overlap (IOU) of a small box drop sharply but barely changes a large one. That is why YOLO predicts √w instead of w and computes the penalty on the difference of square roots.

Ground truth · width w Prediction · width w + Δw

Predicted width w + Δw 0.15Overlap IOU 0.667

Penalty = λcoord × (prediction − ground truth)², λcoord = 5

w as is0.0125

5 × (0.15 − 0.10)²

√w (YOLO)0.0253

5 × (√0.15 − √0.10)² = 5 × (0.387 − 0.316)²

Penalty with Δw fixed at 0.05, varying only the ground-truth width w

With the same Δw (0.05), the difference in √w is 0.071 for a small box (w = 0.1) and 0.028 for a large box (w = 0.8). The penalty is the square of that difference, so the small box's penalty is 6.7 times larger. Using w as is, the two penalties would be equal (0.0125).

Empty boxes: most of the 98 have no object

Two boxes in each of the 7 × 7 cells make 98 boxes. In the cell with an object's center, only one box takes that object; every other box learns a confidence of 0. Empty boxes far outnumber the others, so left alone they drown out the signal from the boxes responsible for objects. The paper multiplies the empty boxes' penalty by λnoobj = 0.5.

Bar height is the weight of the confidence penalty. Drag to rotate; click a cell to place or remove an object (the box you click takes it). Keyboard: arrows rotate.

Responsible box · weight 1 Empty box · weight λnoobj Cell with an object's center

Responsible boxes 2 · empty boxes 96

Click a cell to place an object, and this line shows what each of the cell's two boxes learns.

Weight of the confidence penalty

Each empty box weighs little (λnoobj = 0.50), but all 96 together carry 96.0% of the confidence-penalty weight (98.0% with λnoobj = 1). The paper shrinks the empty boxes' penalty with λnoobj = 0.5 and, at the same time, scales up the coordinate penalty with λcoord = 5.

Source: paper §2.2, Eq. 3

Limits

The price of looking once

MythLooking only once means looking carelessly.

ActuallyBecause it sees the whole image at once, it uses the surrounding context. In the paper, YOLO made less than half as many background errors (mistaking background for an object) as Fast R-CNN.

Source: paper §1, §2.4, §4.2

Questions

Frequently asked questions

What does YOLO stand for?

YOLO stands for You Only Look Once. It is an object detection method, which finds where the objects in an image are and what they are. Instead of checking candidate regions one by one, it passes the whole image through a neural network once and predicts every object at the same time.

What is the YOLOv1 architecture?

It has 24 convolutional layers followed by 2 fully connected layers. It takes a 448 × 448 image and outputs a 7 × 7 × 30 tensor. Counted on an implementation that follows the paper, it has about 272 million parameters, and about 76% of them are in the first fully connected layer.

Why is the YOLOv1 output 7 × 7 × 30?

The image is divided into a 7 × 7 grid, and each cell predicts 2 boxes and 20 class probabilities. Each box is 5 numbers (x, y, w, h, and confidence), so each cell has 2 × 5 + 20 = 30 numbers.

How many boxes does YOLOv1 predict per image, and how are they reduced?

It predicts 7 × 7 cells × 2 = 98 boxes. Each box is scored by class probability × confidence and low-scoring boxes are discarded; among boxes that overlap on the same object, non-maximum suppression (NMS) keeps only the one with the highest score.

Why does YOLO predict the square root of the width instead of the width?

Because the same size error makes the overlap (IOU) of a small box drop sharply but barely affects a large box. With the penalty computed on the difference of square roots, getting the width wrong by the same 0.05 costs a small box (w = 0.1) about 6.7 times as much as a large box (w = 0.8). The paper notes that this only partially addresses the problem.

What are λcoord and λnoobj in the YOLO loss?

They are weights in the loss function (Eq. 3). Box coordinate errors are multiplied by λcoord = 5 to increase them, and the confidence errors of boxes without objects are multiplied by λnoobj = 0.5 to reduce them. Most cells contain no object, so without these weights the signal pushing confidence toward 0 would overpower the signal from cells with objects and make training unstable.

How fast is YOLOv1?

According to the paper, it processes 45 images per second on a Titan X GPU, and the smaller Fast YOLO processes 155 images per second. That is fast enough to process streaming video in real time with less than 25 milliseconds of latency.

What are the limitations of YOLOv1?

Each cell predicts only 2 boxes and 1 class, so small objects that appear in groups, like a flock of birds, are easy to miss. It struggles with objects in aspect ratios or arrangements not seen in training, and its most common error is inaccurate localization.

Source: paper abstract, §1, §2, §2.2, §2.3, §2.4, Table 1 · The parameter count and the penalty ratio come from this page's implementation and calculation

Further reading

The layer structure, compute figures, and predicted boxes on this page come from a PyTorch implementation that follows the paper.