ResNet: Residual Connections, 152 Layers, and ImageNet

Microsoft Research’s residual shortcuts solved degradation in deep networks, trained a 152-layer model, and carried into later Transformer blocks.

Stacking more layers was making networks worse, not better

Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun posted the paper to arXiv on December 10, 2015. Their starting observation was a puzzle: adding layers to a plain convolutional network eventually made training error go up, not down. That was not overfitting, since training accuracy itself got worse. The authors called this degradation, and it meant depth had stopped being a free way to add capacity. [1]

Their fix left each stacked group of layers learning a residual, the difference between its output and its input, instead of the whole transformation. A shortcut connection carries the input forward and adds it back after the layers, so a block with nothing useful to contribute can settle toward passing its input through unchanged rather than corrupting it. [1]

From breaking past twenty layers to a hundred and fifty two

The paper's ImageNet models reached 152 layers, eight times deeper than the VGG networks it compared against, while carrying fewer parameters. A single 152 layer residual network scored 4.49 percent top five validation error. An ensemble of six models, including two at that depth, brought the ILSVRC 2015 test result to 3.57 percent, enough for first place in the classification task. [1][2]

The same team also won the ImageNet detection and localization tasks and the COCO 2015 detection and segmentation tasks that year, reporting a 28 percent relative improvement on COCO object detection over the previous best. Depth that plain networks could not use without getting worse was, with residual connections added, directly convertible into better scores across several benchmarks at once. [1][2]

A shortcut simple enough to bolt onto almost anything

Atlas interpretation: The residual block is a small change to wire into an existing design: keep a copy of the input, run it through some layers, add the copy back in. That made it easy to adopt without redesigning a network from scratch, and it removed a ceiling that had been blocking a straightforward path to more capacity since the leap forward of AlexNet three years earlier. [1]

Atlas interpretation: The same shortcut later turned up outside vision entirely. Transformer blocks wrap their attention and feed forward layers in residual connections, which is what lets those networks be stacked dozens of layers deep without the same training breakdown ResNet was built to fix. The mechanism generalized past the problem it was invented for. [3]

Sources

  1. Deep Residual Learning for Image Recognition

    arXiv · Dec 10, 2015

  2. ILSVRC2015 Results

    ImageNet

  3. Attention Is All You Need

    arXiv · Jun 12, 2017