한국어

KNOWLEDGE · OBJECT DETECTION

How does YOLOv2 find crowded and small objects better?

YOLOv2 gives each cell five anchor boxes chosen from the data, so even one cell can share several objects among them, and a passthrough layer folds fine-grained 26 × 26 features into the 13 × 13 grid.

YOLOv1 divides an image into a 7 × 7 grid; each cell outputs 2 boxes but predicts only one set of classes. So when two object centers fall in one cell it misses one of them, and because it works from coarse, heavily downsampled features, its boxes are often off. YOLOv2 fixed these weaknesses with a series of improvements. This page lets you try two of them: anchor boxes and the passthrough layer.

VOC 2007 mAP as each change is added, from YOLOv1 to YOLOv2 (paper Table 2)

  1. YOLOv163.4
  2. Batch normalization65.8+2.4
  3. High-resolution classifier69.5+3.7
  4. Convolutional + anchor boxes0169.2−0.3
  5. New network: Darknet-1969.6+0.4
  6. Anchors from data + location bound to the cell01·0274.4+4.8
  7. Passthrough0375.4+1.0
  8. Multi-scale training76.8+1.4
  9. High-resolution detector (544)78.6+1.8

The bars run from mAP 60 to 80. The switch to anchor boxes lowered mAP slightly, but recall rose from 81% to 88% (§2).

  1. 01 Anchor boxesOne cell, five box templates
  2. 02 Choosing anchorsThe data picks the shapes
  3. 03 PassthroughFold and attach

01 Anchor boxes

Five box templates per cell, one class per template

YOLOv2 takes a 416 × 416 image and divides it into a 13 × 13 grid. Each cell has five box templates of different shapes: the anchor boxes. An object is taken by one anchor in the cell that contains its center, the anchor that overlaps it most when their centers are aligned. Unlike YOLOv1, each anchor predicts its own objectness and classes, so even two objects in one cell can both be learned if their shapes differ.

Drag the person, the car, and the dog. Switch between YOLOv1 and YOLOv2 with the buttons above. With a keyboard, select the picture, move with the arrow keys (Shift moves one cell at a time), and press Space to switch objects.

Boxes 13 × 13 × 5 = 845 · class probabilities: 845 sets (1 per anchor)

Selected object

person → row 7, column 7 · anchor 2

How much the object overlaps each of this cell's 5 anchors (IOU with centers aligned). The anchor with the most overlap takes it.

Position in the cell σ(tx), σ(ty)
0.11, 0.76
Size relative to the anchor etw, eth
0.74, 0.67
What the network predicts: tw, th
−0.30, −0.40

bx = σ(tx) + cx · by = σ(ty) + cy · bw = pwetw · bh = pheth

At 0%, the anchor sits unchanged at the center of the cell. The network moves and stretches the anchor to fit the object: the sigmoid σ keeps the center inside its own cell, and the size becomes et times the anchor. The closer the anchor's shape is to the object, the closer tw and th are to 0, and the easier they are to learn.

Counted over the 12,032 objects in the VOC 2007 test set, 6.8% lose their slot to another object in the same cell with YOLOv1 (7 × 7 × 2 boxes, one set of classes per cell). With YOLOv2 (13 × 13 × 5 anchors, one set of classes per anchor) it is 1.5%. For objects smaller than 1% of the image area, it drops from 24.0% to 7.9%.

Source: paper §2 Convolutional With Anchor Boxes, Direct location prediction, Figure 3 · The anchor assignment rule follows the author's Darknet code (region_layer.c) · The anchor sizes and the 6.8% and 1.5% were computed directly from the VOC ground-truth boxes

02 Choosing anchors

Five box templates: which shapes?

The closer the anchors are to the shapes of objects, the less the network has to correct, and the easier it learns. YOLOv2 doesn't hand-pick its anchors: it clusters the ground-truth boxes of the training data with k-means. Each point picks its nearest anchor, then each anchor moves to the mean of its points, and this repeats until nothing changes. The distance is 1 − IOU: the more two boxes overlap with their centers aligned, the closer they are.

Each dot is the width and height of one VOC ground-truth box (relative to the image, on a log scale). Drag the iteration slider or the curve to follow k-means as it moves the anchors; each anchor leaves a trail (the ring is where it started). Drag a handle to restart from there. Keyboard: arrows on the iteration slider; in the picture, pick an anchor with [ ] and resize it with the arrow keys.

Distance

k-means progress: the average IOU after each move of the anchors

Iteration 0 · starting anchors: 5 squares. Each dot is colored by its nearest anchor. Average IOU 57.9%. Drag the slider to the right to watch k-means move the anchors.

Average IOU57.9%
Smaller half of boxes53.7%
Larger half of boxes62.2%

Average IOU: the mean, over all boxes, of each box's IOU with its best-matching anchor

The current anchors drawn in the center cell of a 416 × 416 input

Paper Table 1: average IOU of the VOC 2007 ground-truth boxes with their closest anchor
How the anchors were chosenCountAvg IOU
k-means, Euclidean distance558.7
k-means, 1 − IOU distance561.0
k-means, 1 − IOU distance967.2

With the same 5 anchors, 1 − IOU distance fits better than Euclidean distance. Euclidean distance weighs the errors of large boxes heavily, so the anchors get pulled toward large boxes. Switch the distance above and compare the anchors' paths and the score for the smaller half of the boxes. With 9 anchors the fit improves, but the model gets more complex; the paper chose k = 5 as a balance.

Source: paper §2 Dimension Clusters, Table 1, Figure 2 · The dots are a random 2,000 of the 40,058 ground-truth boxes in the VOC 2007+2012 training set (numbers only, no images), and the scores were computed directly from these dots

03 Passthrough

Fold the fine 26 × 26 features and attach them

One cell of the 13 × 13 feature map covers 32 × 32 pixels of the input. That is enough for large objects but coarse for locating small ones. The passthrough layer takes the 26 × 26 features still available in an earlier layer and attaches them to the 13 × 13 features. The detector predicts boxes from these widened features.

Where it takes from and where it attaches

The YOLOv2 detector has 22 convolutional layers: the 18 of Darknet-19 without its last, classification convolution, followed by three 3 × 3 × 1024 convolutions for detection and one 1 × 1 output convolution. The paper adds the passthrough “from the final 3 × 3 × 512 layer to the second to last convolutional layer.” Hover over or click a block to see what that layer is; click ①–④ to send the folding picture below to that step.

conv 13 · 3 × 3 × 512 · output 26 × 26 × 512 · the last layer at 26 × 26 resolution; the passthrough takes this output

  1. ① SourceThe 26 × 26 × 512 output of the 13th convolution (3 × 3 × 512): the last layer at 26 × 26 resolution, right before the fifth max pooling.
  2. ② Main pathThe same output becomes 13 × 13 at the fifth max pooling and passes through 7 convolutions (14–18 of Darknet-19 and detection layers 19 and 20) to become 13 × 13 × 1024.
  3. ③ ShortcutThe passthrough skips that path. It folds 26 × 26 × 512 into channels, 2 × 2 at a time, making 13 × 13 × 2048. You fold it yourself below.
  4. ④ MergeRight before the 21st convolution (second to last, 3 × 3 × 1024), the two paths are concatenated along the channels into 13 × 13 × 3072. The 21st convolution takes this and outputs 13 × 13 × 1024, and the 22nd (1 × 1 × 125) outputs 13 × 13 × 125.

How it folds and attaches

On the ③ shortcut, each neighboring 2 × 2 group of cells is stacked along the channels instead of in space. 26 × 26 × 512 becomes 13 × 13 × 2048, which at ④ is concatenated next to the main path's 13 × 13 × 1024 to make 13 × 13 × 3072.

Click the step tabs to fold. In the picture above, the place where the current step happens lights up. Click a cell on the 26 × 26 face to follow where its 2 × 2 group goes. Drag to rotate. Keyboard: arrows rotate, Shift + arrows move the group.

1 / 4

① The 26 × 26 × 512 features from the 13th convolution (the last 3 × 3 × 512 layer at 26 × 26 resolution). Each cell holds 512 numbers. The color shows where a cell sits in its 2 × 2 group.

Picked group: rows 13–14, columns 13–14 of 26 × 26 → one cell at row 7, column 7 of 13 × 13

Number of values: 26 × 26 × 512 = 346,112 → 13 × 13 × 2048 = 346,112

The 4 values of a 2 × 2 group go into different channels of one output cell. Nothing is thrown away, only moved, so the fine detail of the 26 × 26 resolution survives in the 13 × 13 grid. Max pooling would keep only the largest of the four. In the paper, the passthrough raised mAP by 1, from 74.4 to 75.4.

The author's published configuration files changed over time. The first version, from November 2016, folds all 512 channels as in this picture, making 3072 channels. From the March 2017 version on, a 1 × 1 convolution on the ③ shortcut first cuts the channels to 64, so 13 × 13 × 256 is attached (1280 channels in total). The source and the merge point stay the same.

Source: paper §2 Fine-Grained Features, §3 Training for detection, Table 2 · The layer order and connections follow cfg/yolo-voc.cfg in the author's Darknet repository (route −9, reorg 2, route −1,−3) and its history

Summary

From YOLOv1's limits to YOLOv2

  1. One class per cell → one class per anchor. YOLOv1 had 7 × 7 × 2 = 98 boxes with 49 sets of class probabilities (one per cell). YOLOv2 has 13 × 13 × 5 = 845 boxes with 845 sets (one per anchor), so one cell can take several objects of different shapes. The output is 13 × 13 × 125.
  2. Boxes from scratch → small edits to templates chosen from data. The anchor shapes come from running k-means with 1 − IOU distance on the ground-truth boxes, and the network adjusts the center within the cell and the size to et times the anchor. The corrections are small and the center can't leave its cell, so training is stable.
  3. Coarse 13 × 13 → fine 26 × 26 folded in. The passthrough stacks each 2 × 2 neighborhood into channels to make 13 × 13 × 3072, keeping the information needed to locate small objects.

Small objects are still a weak spot: on COCO, YOLOv2's AP for small objects is 5.0, lower than the 9.0 of SSD512 in the same table (paper Table 5).

MythWith anchor boxes, box sizes are limited to the 5 anchors.

ActuallyAnchors are starting points. The network scales the width and height to et times the anchor and moves the center anywhere within the cell. Anchors decide which predictor takes which object, and they keep the corrections small.

MythThe passthrough shrinks (pools) the earlier layer's features before attaching them.

ActuallyIt doesn't shrink anything. It only moves each 2 × 2 neighborhood into channels, so all 346,112 values of 26 × 26 × 512 fit into 13 × 13 × 2048. Max pooling would keep only 1 in 4.

Source: paper §2, §3, Figure 3, Table 5

Questions

Frequently asked questions

How is YOLOv2 different from YOLOv1?

It adds batch normalization, a high-resolution classifier, anchor boxes, anchor sizes chosen by k-means, location prediction bound to the cell, a passthrough layer, multi-scale training, and a new network, Darknet-19. In the paper's Table 2, VOC 2007 mAP rises from 63.4 to 78.6, and the same model run at 416 × 416 reaches 76.8 mAP at 67 frames per second.

What is an anchor box?

It is a predefined box shape. Instead of drawing a box from scratch, the network predicts how far to move and stretch the anchor. YOLOv2 has 5 anchors in each of its 13 × 13 cells and predicts objectness and classes separately for each anchor, so it can find two differently shaped objects in the same cell.

How many anchor boxes does YOLOv2 use, and how are they chosen?

Five. They are chosen by running k-means clustering on the widths and heights of the training data's ground-truth boxes, with 1 − IOU as the distance. More anchors raise the average IOU but make the model more complex, so the paper chose 5 as a balance (Table 1: 61.0 with 5, 67.2 with 9).

Why does the anchor k-means use 1 − IOU instead of Euclidean distance?

With Euclidean distance, large boxes produce larger errors, which pulls the cluster centers toward large boxes. What we want are anchors with high IOU regardless of box size, so 1 − IOU is used. In the paper's Table 1, with 5 anchors the average IOU was 58.7 with Euclidean distance and 61.0 with IOU distance.

How does YOLOv2 predict box position and size?

Relative to the cell's top-left corner (c_x, c_y) and the anchor size (p_w, p_h), it builds the box as b_x = σ(t_x) + c_x, b_y = σ(t_y) + c_y, b_w = p_w·e^(t_w), b_h = p_h·e^(t_h). The sigmoid σ keeps the center inside its own cell, which stabilizes training; combined with anchors chosen by k-means, mAP rises from 69.6 to 74.4 in the paper's Table 2.

What is the passthrough layer in YOLOv2?

It takes the 26 × 26 × 512 features of an earlier layer, stacks each neighboring 2 × 2 group of cells along the channels to get 13 × 13 × 2048, and concatenates that with the original 13 × 13 features. Nothing is thrown away, only moved, so the fine-grained features survive and help locate small objects. In the paper it raised mAP from 74.4 to 75.4.

Why is the YOLOv2 output 13 × 13 × 125?

The 416 × 416 input is downsampled 32 times to a 13 × 13 grid, and in each cell 5 anchors each output 4 coordinates, 1 objectness score, and 20 classes (PASCAL VOC): 5 × (4 + 1 + 20) = 125, or 845 boxes per image. Because 13 is odd, there is a single cell at the very center of the image.

Are YOLOv2 and YOLO9000 the same?

They come from the same paper (YOLO9000: Better, Faster, Stronger, CVPR 2017) but are different models. YOLOv2 is the improved detector from the first part; YOLO9000 uses that architecture, trained jointly on COCO detection data and ImageNet classification data, to detect more than 9000 object categories. YOLO9000 uses 3 anchors instead of 5.

Source: paper §2, §3, §4, Table 1, Table 2, Table 3 · The 845 boxes and the output size are this page's calculation

Further reading

The anchor sizes in 01, the dots and scores in 02, and the 6.8% and 1.5% in 01 were computed directly from the numbers of the PASCAL VOC ground-truth boxes. No images were used.