Reference. Stride and Translation Invariance in CNNs

Convolutional Neural Networks have become the standard for image classification tasks, however, these architectures are not invariant to translations of the input image. This lack of invariance is attributed to the use of stride which ignores the sampling theorem, and fully connected layers which lack spatial reasoning. We show that stride can greatly benefit translation invariance given that it is combined with sufficient similarity between neighbouring pixels, a characteristic which we refer to as local homogeneity. We also observe that this characteristic is dataset-specific and dictates the relationship between pooling kernel size and stride required for translation invariance. Furthermore we find that a trade-off exists between generalization and translation invariance in the case of pooling kernel size, as larger kernel sizes lead to better invariance but poorer generalization. Finally we explore the efficacy of other solutions proposed, namely global average pooling, anti-aliasing, and data augmentation, both empirically and through the lens of local homogeneity.

Cite

Cite as @moutonStrideTranslationInvariance2020 (helia, typst) · \cite{moutonStrideTranslationInvariance2020} (LaTeX)
BibTeX
bibtex · 17 lines
@incollection{moutonStrideTranslationInvariance2020,
 title = {Stride and {{Translation Invariance}} in {{CNNs}}},
 author = {Mouton, Coenraad and Myburgh, Johannes C. and Davel, Marelie H.},
 date = {2020},
 doi = {10.1007/978-3-030-66151-9_17},
 url = {http://arxiv.org/abs/2103.10097},
 urldate = {2024-01-15},
 volume = {1342},
 pages = {267--281},
 keywords = {Computer Science - Machine Learning},
 abstract = {Convolutional Neural Networks have become the standard for image classification tasks, however, these architectures are not invariant to translations of the input image. This lack of invariance is attributed to the use of stride which ignores the sampling theorem, and fully connected layers which lack spatial reasoning. We show that stride can greatly benefit translation invariance given that it is combined with sufficient similarity between neighbouring pixels, a characteristic which we refer to as local homogeneity. We also observe that this characteristic is dataset-specific and dictates the relationship between pooling kernel size and stride required for translation invariance. Furthermore we find that a trade-off exists between generalization and translation invariance in the case of pooling kernel size, as larger kernel sizes lead to better invariance but poorer generalization. Finally we explore the efficacy of other solutions proposed, namely global average pooling, anti-aliasing, and data augmentation, both empirically and through the lens of local homogeneity.},
 eprintclass = {cs},
 eprinttype = {arXiv},
 eprint = {2103.10097},
 booktitle = {Artificial Intelligence Research (SACAIR 2020)},
 series = {Communications in Computer and Information Science}
}
hayagriva YAML (typst)
yaml · 23 lines
moutonStrideTranslationInvariance2020:
  type: anthos
  title: Stride and {Translation Invariance} in {CNNs}
  author:
  - Mouton, Coenraad
  - Myburgh, Johannes C.
  - Davel, Marelie H.
  date: 2020
  page-range: 267-281
  url:
    value: http://arxiv.org/abs/2103.10097
    date: 2024-01-15
  serial-number:
    arxiv: '2103.10097'
    doi: 10.1007/978-3-030-66151-9_17
  abstract: Convolutional Neural Networks have become the standard for image classification tasks, however, these architectures are not invariant to translations of the input image. This lack of invariance is attributed to the use of stride which ignores the sampling theorem, and fully connected layers which lack spatial reasoning. We show that stride can greatly benefit translation invariance given that it is combined with sufficient similarity between neighbouring pixels, a characteristic which we refer to as local homogeneity. We also observe that this characteristic is dataset-specific and dictates the relationship between pooling kernel size and stride required for translation invariance. Furthermore we find that a trade-off exists between generalization and translation invariance in the case of pooling kernel size, as larger kernel sizes lead to better invariance but poorer generalization. Finally we explore the efficacy of other solutions proposed, namely global average pooling, anti-aliasing, and data augmentation, both empirically and through the lens of local homogeneity.
  parent:
    type: anthology
    title: Artificial Intelligence Research (SACAIR 2020)
    volume: 1342
    parent:
      type: anthology
      title: Communications in Computer and Information Science
Cites 15 works (1 here)
With notes (1)

Why do deep convolutional networks generalize so poorly to small image transformations? azulayWhyDeepConvolutional

Convolutional Neural Networks (CNNs) are commonly assumed to be invariant to small image transformations: either because of the convolutional architecture or because they were trained using data augmentation. Recently, several authors have shown that this is not the case: small translations or rescalings of the input image can drastically change the network’s prediction. In this paper, we quantify this phenomena and ask why neither the convolutional architecture nor data augmentation are sufficient to achieve the desired invariance. Specifically, we show that the convolutional architecture does not give invariance since architectures ignore the classical sampling theorem, and data augmentation does not give invariance because the CNNs learn to be invariant to transformations only for images that are very similar to typical images from the training set. We discuss two possible solutions to this problem: (1) antialiasing the intermediate representations and (2) increasing data augmentation and show that they provide only a partial solution at best. Taken together, our results indicate that the problem of insuring invariance to small image transformations in neural networks while preserving high accuracy remains unsolved.
DOI
moutonStrideTranslationInvariance2020 reference entries/refs/moutonStrideTranslationInvariance2020/moutonStrideTranslationInvariance2020.hel