Residual networks behave like ensembles of relatively shallow networks

Publikation: Bidrag til tidsskrift › Konferenceartikel › Forskning › fagfællebedømt

Andreas Veit
Michael Wilber
Belongie, Serge

In this work we propose a novel interpretation of residual networks showing that they can be seen as a collection of many paths of differing length. Moreover, residual networks seem to enable very deep networks by leveraging only the short paths during training. To support this observation, we rewrite residual networks as an explicit collection of paths. Unlike traditional models, paths through residual networks vary in length. Further, a lesion study reveals that these paths show ensemble-like behavior in the sense that they do not strongly depend on each other. Finally, and most surprising, most paths are shorter than one might expect, and only the short paths are needed during training, as longer paths do not contribute any gradient. For example, most of the gradient in a residual network with 110 layers comes from paths that are only 10-34 layers deep. Our results reveal one of the key characteristics that seem to enable the training of very deep networks: Residual networks avoid the vanishing gradient problem by introducing short paths which can carry gradient throughout the extent of very deep networks.

Originalsprog	Engelsk
Tidsskrift	Advances in Neural Information Processing Systems
Sider (fra-til)	550-558
Antal sider	9
ISSN	1049-5258
Status	Udgivet - 2016
Eksternt udgivet	Ja
Begivenhed	30th Annual Conference on Neural Information Processing Systems, NIPS 2016 - Barcelona, Spanien Varighed: 5 dec. 2016 → 10 dec. 2016

Konference

Konference	30th Annual Conference on Neural Information Processing Systems, NIPS 2016
Land	Spanien
By	Barcelona
Periode	05/12/2016 → 10/12/2016
Sponsor	et al., Google, Intel Corporation, KLA-Tencor, Microsoft, Winton

Bibliografisk note

ID: 301828179

Datalogisk Institut