Reference. Why do deep convolutional networks generalize so poorly to small image transformations?
Convolutional Neural Networks (CNNs) are commonly assumed to be invariant to small image transformations: either because of the convolutional architecture or because they were trained using data augmentation. Recently, several authors have shown that this is not the case: small translations or rescalings of the input image can drastically change the network’s prediction. In this paper, we quantify this phenomena and ask why neither the convolutional architecture nor data augmentation are sufficient to achieve the desired invariance. Specifically, we show that the convolutional architecture does not give invariance since architectures ignore the classical sampling theorem, and data augmentation does not give invariance because the CNNs learn to be invariant to transformations only for images that are very similar to typical images from the training set. We discuss two possible solutions to this problem: (1) antialiasing the intermediate representations and (2) increasing data augmentation and show that they provide only a partial solution at best. Taken together, our results indicate that the problem of insuring invariance to small image transformations in neural networks while preserving high accuracy remains unsolved.
Cite
Cited by (1)
Stride and Translation Invariance in CNNs moutonStrideTranslationInvariance2020
Convolutional Neural Networks have become the standard for image classification tasks, however, these architectures are not invariant to translations of the input image. This lack of invariance is attributed to the use of stride which ignores the sampling theorem, and fully connected layers which lack spatial reasoning. We show that stride can greatly benefit translation invariance given that it is combined with sufficient similarity between neighbouring pixels, a characteristic which we refer to as local homogeneity. We also observe that this characteristic is dataset-specific and dictates the relationship between pooling kernel size and stride required for translation invariance. Furthermore we find that a trade-off exists between generalization and translation invariance in the case of pooling kernel size, as larger kernel sizes lead to better invariance but poorer generalization. Finally we explore the efficacy of other solutions proposed, namely global average pooling, anti-aliasing, and data augmentation, both empirically and through the lens of local homogeneity.
Cites 62 works (0 here)
External (62)
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations (2019)
- Making Convolutional Networks Shift-Invariant Again (2019)
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs (2018)
- The Elephant in the Room (2018)
- Learned Deformation Stability in Convolutional Neural Networks (2018)
- Studying Invariances of Trained Convolutional Neural Networks (2018)
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric (2018)
- Adversarial Examples that Fool both Computer Vision and Time-Limited\n Humans (2018)
- Why do deep convolutional networks generalize so poorly to small image transformations? (self-citation of the arXiv version) (2018)
- Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning (2017)
- Harmonic Networks: Deep Translation and Rotation Equivariance (2017)
- Synthesizing Robust Adversarial Examples (2017)
- Robust Physical-World Attacks on Deep Learning Models (2017)
- A Rotation and a Translation Suffice: Fooling CNNs with Simple Transformations (2017)
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities (2017)
- Quantifying Translation-Invariance in Convolutional Neural Networks (2017)
- Exploring the Space of Black-box Attacks on Deep Neural Networks (2017)
- Polar Transformer Networks (2017)
- Query-limited Black-box Attacks to Classifiers (2017)
- Densely Connected Convolutional Networks (2017)
- Dilated Residual Networks (2017)
- One pixel attack for fooling deep neural networks (cited in v1) (2017)
- Deep Residual Learning for Image Recognition (2016)
- Adversarial examples in the physical world (2016)
- Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images (2016)
- Accessorize to a Crime (2016)
- RIFD-CNN: Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection (2016)
- Fine-grained Recognition in the Noisy Wild: Sensitivity Analysis of Convolutional Neural Networks Approaches (2016)
- Group Equivariant Convolutional Networks (2016)
- Exploiting cyclic symmetry in convolutional neural networks (2016)
- Going deeper with convolutions (2015)
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification (2015)
- Rotation-invariant convolutional neural networks for galaxy morphology prediction (2015)
- Understanding image representations by measuring their equivariance and equivalence (2015)
- Manitest: Are classifiers really invariant? (2015)
- Multi-Scale Context Aggregation by Dilated Convolutions (2015)
- Geodesics of learned representations (2015)
- Object Detectors Emerge in Deep Scene CNNs (2015)
- Visualizing and Understanding Convolutional Networks (2014)
- Deep Symmetry Networks (2014)
- Very Deep Convolutional Networks for Large-Scale Image Recognition (2014)
- Scale-Invariant Convolutional Neural Networks (2014)
- Transformation Properties of Learned Visual Representations (2014)
- Semantic image segmentation with deep convolutional nets and fully connected CRFs (2014)
- Rotation, Scaling and Deformation Invariant Scattering for Texture Discrimination (2013)
- Intriguing properties of neural networks (2013)
- Learning about Canonical Views from Internet Image Collections (2012)
- ImageNet classification with deep convolutional neural networks (cited in v1/v2) (2012)
- Unbiased look at dataset bias (2011)
- Discovering favorite views of popular places with iconoid shift (2011)
- ImageNet: A large-scale hierarchical image database (2009)
- Finding iconic images (2009)
- Computing iconic summaries of general visual concepts (2008)
- Scene Summarization for Online Image Collections (2007)
- Distinctive Image Features from Scale-Invariant Keypoints (2004)
- Local Scale Selection for Gaussian Based Description Techniques (2000)
- Object recognition from local scale-invariant features (1999)
- Scale-space theory: a basic tool for analyzing structures at different scales (1994)
- Shiftable multiscale transforms (1992)
- Backpropagation Applied to Handwritten Zip Code Recognition (1989)
- Neocognitron: A hierarchical neural network capable of visual pattern recognition (1988)
- Neocognitron: A Self-Organizing Neural Network Model for a Mechanism of Visual Pattern Recognition (1982)