한국어

KNOWLEDGE · OBJECT DETECTION

How does YOLOv3 find large and small objects in one pass?

YOLOv3 predicts boxes on three grids, 13 × 13, 26 × 26 and 52 × 52, with three anchors per cell, so large and small objects are shared among them, and for each ground-truth object it teaches only the best-fitting of the nine anchors that it holds an object.

YOLOv2 divides a 416 × 416 image into a 13 × 13 grid and outputs 845 boxes from five anchor boxes per cell. One cell covers 32 × 32 pixels of the image, which is coarse for small objects, and on COCO its AP for small objects was 5.0. YOLOv3 added a number of small changes. This page lets you try two of them: predicting on grids of three sizes (multi-scale prediction) and what each box learns (bounding box prediction).

COCO results of YOLOv2 and YOLOv3 (paper Table 3, test-dev)

  1. Small objects APS01·025.0 → 18.3
  2. Medium objects APM22.4 → 35.4
  3. Large objects APL35.5 → 41.9
  4. AP50 (IOU 0.5)44.0 → 57.9
  5. AP (averaged over IOU 0.5–0.95)21.6 → 33.0

YOLOv2YOLOv3 608 × 608

Bars run from 0 to 60. The comparison mixes several changes, including the backbone (Darknet-53) and the input size; the paper has no table that adds them one at a time. The paper credits the gain on small objects to the new multi-scale predictions (§3).

  1. 01 Three gridsA grid for each object size
  2. 02 Feature pathsUnfold and concatenate
  3. 03 What a box learns1, ignore, 0

01 Multi-scale prediction

Three grids, three anchors per cell

YOLOv3 looks at a 416 × 416 image through three grids of different sizes: 13 × 13 (cells 32 pixels wide), 26 × 26 (16 pixels) and 52 × 52 (8 pixels). Every cell of every grid has three anchor boxes, so there are nine anchors in all. The large three are used by 13 × 13, the middle three by 26 × 26 and the small three by 52 × 52. For each object, the one anchor of the nine whose shape fits best takes the object, and the cell holding the object's center on that anchor's grid predicts the box.

Drag the bus, the dog and the bird. Grow or shrink the selected object with the size slider, and the anchor and grid that take it change. With a keyboard, focus the picture, move with the arrow keys (Shift for big steps), resize with + −, and switch objects with the space bar.

Boxes: 13 × 13 × 3 = 507 · 26 × 26 × 3 = 2,028 · 52 × 52 × 3 = 8,112 · total 10,647

Selected object

Bus → 13 × 13 grid · cell at row 7, column 9 · anchor 156 × 198

Overlap with each of the nine anchors (IOU with the centers aligned). The anchor with the largest one takes the object, and its grid predicts it.

There is no separate rule that picks the grid by size. Once the anchor with the most similar shape is chosen, the grid follows. That is why small objects are predicted on 52 × 52, whose cells are 8 pixels, and large objects on 13 × 13, whose cells are 32 pixels. YOLOv2 output all of its 845 boxes (13 × 13 × 5 anchors) on 13 × 13; YOLOv3 outputs 10,647 over three grids.

Source: paper §2.3 Predictions Across Scales · the nine anchors are the COCO values in §2.3 · three anchors per grid from the authors' Darknet config cfg/yolov3.cfg (mask 6,7,8 / 3,4,5 / 0,1,2) · choosing the anchor that takes an object among all nine follows the authors' code (yolo_layer.c)

02 Upsampling and concatenation

Deep features are unfolded onto the finer grids

Predicting on three grids needs 13 × 13, 26 × 26 and 52 × 52 feature maps. The 26 × 26 and 52 × 52 maps early in the network are fine but have passed through few layers; the 13 × 13 map at the end has passed through many layers but is coarse. YOLOv3 unfolds the late features by 2× at a time (upsampling) and concatenates them with the earlier maps. The paper says this gets more meaningful semantic information from the upsampled features and finer-grained information from the earlier map.

Stacking the 107 layers as a pyramid

YOLOv3 has 107 layers (numbered 0–106 in the authors' config file). The picture below draws every layer as one plate, stacks the plates by resolution and splits them into three columns, backbone, neck and head, the way object detectors are usually drawn. A plate's width is its resolution and its thickness its channels. The backbone (Darknet-53, layers 0–74) stacks upward as it halves the image and becomes a pyramid. The neck starts at the top and comes down one level at a time. At every level it takes the backbone map of the same height from the side (layers 74, 61, 36), and it unfolds the features of layers 79 and 91 with a 1 × 1 convolution and a 2× upsample, sends them a level down and concatenates them there (layers 86, 98). So the neck widens on its way down. The head predicts on the neck's three levels with a 3 × 3 and a 1 × 1 convolution. These three paths (bottom-up, top-down, lateral) are the shape the paper calls similar to feature pyramid networks (FPN).

Follow the step tabs. “Layer map” shows it from the front with the layer numbers; “Pyramid” shows it at an angle. Hover over or tap a plate to see what that layer is, and tap ①–⑤ to take the picture below to that step. Drag to rotate. Keyboard: arrow keys rotate. The widths are not to scale: a plate shrinks by the same amount each time the resolution halves. The thickness is proportional to the channels.

1 / 5

The whole of YOLOv3, every layer drawn as one plate and stacked by resolution. On the left, the backbone narrows like a pyramid from the image to 13 × 13; in the middle, the neck widens again from 13 × 13 to 52 × 52 on its way down, an upside-down pyramid; on the right, the head predicts at three heights. Turn it to “Layer map” to see it from the front with the layer numbers.

Backbone maps the neck takes: layer 74 (13 × 13 × 1024), layer 61 (26 × 26 × 512), layer 36 (52 × 52 × 256) · predictions of the head: 13 × 13, 26 × 26 and 52 × 52 × 255

  1. ① Where it comes fromThe 13 × 13 × 512 output of layer 79 (1 × 1 × 512), the last layer of the neck's 13 × 13 row. It comes before the head's two layers (80 and 81), so it is 2 layers before the output convolution that makes the prediction (layer 81). This is the paper's “feature map from 2 layers previous”.
  2. ② Reduce256 1 × 1 convolutions (layer 84) make it 13 × 13 × 256.
  3. ③ UnfoldA 2× upsample (layer 85). Every value is copied into a 2 × 2 block of four cells, giving 26 × 26 × 256.
  4. ④ Where it mergesIt is concatenated along the channels with layer 61 of Darknet-53 (the last residual block at 26 × 26, 26 × 26 × 512) in layer 86: 256 + 512 = 768 channels.
  5. ⑤ PredictFive convolutions in the neck (layers 87–91), then the head's 3 × 3 convolution (layer 92) and output convolution (layer 93), predict 26 × 26 × 255 (layer 94). Repeating the same design gives 52 × 52: layer 91 → 1 × 1 × 128 → 2× upsample → concatenate with layer 36 (52 × 52 × 256, 384 channels) → 52 × 52 × 255 (layer 106).

How the unfolding and concatenation work

Watch ②–⑤ on the path from 13 × 13 to 26 × 26. Upsampling copies each value into a 2 × 2 block of four cells as it is (nearest neighbor). It has no learned weights and creates no new values. The fine 26 × 26 detail comes from the earlier map it is concatenated with.

Follow along with the step tabs. The part of the map above where the current step happens lights up. Tap a cell to follow where it goes. Drag to rotate. Keyboard: arrow keys rotate, Shift + arrow keys move the picked cell.

1 / 6

① The 13 × 13 × 512 features from layer 79, the last layer of the neck's 13 × 13 row, 2 layers before the head's output convolution (layer 81). Each cell covers 32 × 32 pixels of the image.

Picked cell: row 7, column 7 of 13 × 13 · 512 values

Number of values: 13 × 13 × 512 = 86,528

YOLOv2's passthrough went the other way. It folded the 26 × 26 map into channels 2 × 2 at a time and attached it to 13 × 13 (fine features to the coarse grid). YOLOv3 unfolds the deep features of the coarse grid to 26 × 26 and 52 × 52 and predicts on those grids too. It also differs from FPN: FPN adds the two maps; YOLOv3 concatenates them.

The colors are a concept picture for the explanation. They assume the deep features hold where the dog is at 13 × 13 resolution and the earlier map holds its outline at 26 × 26 resolution. Real feature values are not shapes a person can read like this.

Source: paper §2.3 · layer order and shapes from the authors' Darknet config cfg/yolov3.cfg (route −4, upsample 2, route −1, 61 / route −1, 36) · upsampling from the authors' code (upsample_layer.c, nearest neighbor) · the names backbone, neck and head come from the YOLOv4 paper §2.1 and §3.4 (the YOLOv3 paper does not use them), and the boundary between neck and head follows MMDetection's YOLOv3 implementation (YOLOV3Neck, YOLOV3Head)

03 Bounding box prediction

One anchor per object; the rest learn 0 or are ignored

Each of the 10,647 predicted boxes outputs 4 coordinates, 1 objectness (the probability that the box holds an object) and 80 classes. The coordinates move and stretch an anchor with the same formula as YOLOv2. What changed is the target each box has to hit in training. For each ground-truth object, only the box of the best-fitting of the nine anchors learns objectness 1, the coordinates and the class. Boxes of anchors that overlap the ground truth by more than 0.5 without being the best are left without a penalty. All the others learn objectness 0.

Follow the step tabs. Drag the dog or move the size slider, and the counts at each step change. Overlap is measured with the predicted boxes at the start of training, that is, each anchor at the center of its cell. Keyboard: in the picture, the arrow keys move the dog and + − resize it.

1 / 4

① Lay the ground-truth box over the nine anchors with their centers aligned. Only the shapes are compared, not the positions. The one anchor with the largest overlap (IOU) takes this dog.

Overlap with each of the nine anchors (IOU with the centers aligned)

In YOLOv2, the objectness of the box that takes an object was trained toward the IOU of the predicted box and the ground truth. YOLOv3 trains it toward 1 with logistic regression. And for each ground-truth object only one box of the nine anchors takes it. A box that overlaps the ground truth a lot without fitting best learns neither 0 nor 1.

Classes are not a single choice

YOLOv3 does not split the 80 classes with one softmax; it predicts each class separately with a logistic (a probability from 0 to 1). A softmax's class probabilities add up to 1, so it assumes each box has exactly one class. With overlapping labels (for example “Woman” and “Person” in Open Images) that assumption is wrong.

Say the truth for one box is both “Person” and “Woman”. Move the scores (logits) the network outputs and compare the probabilities of the two approaches.

Truth: Person 1 · Woman 1 · Dog 0

Softmax (YOLOv2): the three add up to 1

A logistic per class (YOLOv3): each 0–1 on its own

Softmax: Person 0.72, Woman 0.27 · logistic: Person 0.88, Woman 0.73. With a softmax the two true labels share a probability of 1, so both cannot exceed 0.5. With logistics, Person and Woman are both above 0.5, which says plainly that both are true.

Each class is trained with binary cross-entropy. The paper says a softmax was not needed for good performance on COCO. The detection score for each class is objectness × class probability.

Source: paper §2.1 Bounding Box Prediction, §2.2 Class Prediction · choosing the anchor for a ground truth among all nine, and the detection score, follow the authors' code (yolo_layer.c)

Summary

From YOLOv2 to YOLOv3

  1. One 13 × 13 grid → three grids. The 845 boxes of 13 × 13 × 5 anchors become 13 × 13 × 3 + 26 × 26 × 3 + 52 × 52 × 3 = 10,647. The nine anchors are divided by size, three at a time, and small objects are predicted on 52 × 52, whose cells are 8 pixels.
  2. Folding down → unfolding up. YOLOv2 folded 26 × 26 and attached it to 13 × 13. YOLOv3 unfolds the deep 13 × 13 features by 2× at a time, concatenates them with the earlier 26 × 26 and 52 × 52 maps, and predicts there too.
  3. Hitting the IOU, softmax → 1, ignore, 0, and a logistic per class. For each ground-truth object only one anchor's box learns objectness 1, the boxes of other anchors that overlap it by more than 0.5 are ignored, and the rest learn 0. Classes are independent probabilities.

It is still weak at fitting boxes exactly. On COCO its AP50 at IOU 0.5 is high at 57.9, but its AP averaged over IOU 0.5–0.95 is 33.0. Small objects (APS 18.3) improved, but on medium and large objects it is comparatively lower than the other detectors in the same table, and the paper says the cause needs more investigation (§3, Table 3).

MisconceptionThere is a size rule that sends large objects to 13 × 13 and small objects to 52 × 52.

In factThere is no such rule. Once the anchor whose shape fits best (IOU with the centers aligned) is chosen from the nine, its grid follows. So at the same area of 1,000 pixels, a 50 × 20 box goes to 52 × 52 and a 20 × 50 box goes to 26 × 26.

MisconceptionUpsampling restores fine detail from coarse features.

In factUpsampling only copies each value into a 2 × 2 block of four cells and creates no new information (it has no learned weights either). The fine detail comes from the earlier maps it is concatenated with (layers 61 and 36).

Source: paper §2.1, §2.2, §2.3, §3, Table 3 · the authors' Darknet config cfg/yolov3.cfg and code (yolo_layer.c, upsample_layer.c) · the grids of 50 × 20 and 20 × 50 are this page's calculation

Questions

Frequently asked questions

What is the difference between YOLOv3 and YOLOv2?

YOLOv3 predicts on grids of three sizes (13 × 13, 26 × 26, 52 × 52) and raises the anchors to nine, three per grid. Objectness is learned with logistic regression, and for each ground-truth object only one anchor's box learns 1; classes are predicted with independent logistics instead of a softmax. The backbone changed from Darknet-19 to Darknet-53, which has residual connections. On COCO, AP rose from 21.6 to 33.0 and AP on small objects from 5.0 to 18.3 (paper Table 3).

What is multi-scale prediction in YOLOv3?

It means predicting one image on three grids of different sizes. For a 416 × 416 input, the 13 × 13 (32-pixel cells), 26 × 26 (16-pixel) and 52 × 52 (8-pixel) grids each output boxes from three anchors per cell. Large anchors sit on the coarse grid and small anchors on the fine grid, so small objects are covered too.

What is YOLOv3's output size, and how many boxes does it predict?

For a 416 × 416 input it is 13 × 13 × 255, 26 × 26 × 255 and 52 × 52 × 255. 255 is 3 anchors × (4 coordinates + 1 objectness + 80 COCO classes). That makes 507 + 2,028 + 8,112 = 10,647 boxes, and 22,743 for a 608 × 608 input.

What are YOLOv3's nine anchors, and how are they divided?

(10 × 13), (16 × 30), (33 × 23), (30 × 61), (62 × 45), (59 × 119), (116 × 90), (156 × 198) and (373 × 326) pixels, chosen by running k-means on the COCO ground-truth boxes. The small three are used by the 52 × 52 grid, the middle three by 26 × 26 and the large three by 13 × 13. The paper says it chose nine anchors and three scales arbitrarily.

Which YOLOv3 grid predicts a given object?

Of the nine anchors, the one whose shape fits the object best (the largest IOU with the centers aligned) takes the object, and the cell holding the object's center on that anchor's grid predicts it. There is no separate rule that picks the grid by size.

What do upsampling and concatenation do in YOLOv3?

The features of layer 79, the last layer of the neck's 13 × 13 row, are reduced to 256 channels by a 1 × 1 convolution and upsampled 2× to 26 × 26, then concatenated along the channels with the earlier 26 × 26 × 512 map of Darknet-53 (layer 61), making 768 channels. Doing the same once more makes the 52 × 52 map (384 channels). The paper says this gets more meaningful semantic information from the upsampled features and finer-grained information from the earlier map.

How does YOLOv3 learn objectness?

With logistic regression. For each ground-truth object, only the box of the best-fitting of the nine anchors has target 1, and only that box also learns the coordinates and the class. The box of an anchor that is not the best but overlaps the ground truth by more than 0.5 is ignored (no penalty), and every other box has target 0.

Why does YOLOv3 use logistic classifiers instead of a softmax?

A softmax assumes each box has exactly one class. That does not fit data with overlapping labels such as “Woman” and “Person” in Open Images, so YOLOv3 uses an independent logistic classifier per class with binary cross-entropy. The paper says a softmax was not needed for good performance on COCO.

Source: paper §2.1, §2.2, §2.3, §2.4, Table 3 · the authors' Darknet config cfg/yolov3.cfg · the box counts are this page's calculation

Further reading

The numbers on this page come from the paper and the authors' config file; the grids and overlaps in the pictures are computed by the page itself. No photos or datasets were used.