Reference. Interpreting Neural Networks through the Polytope Lens
Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Previous mechanistic descriptions have used individual neurons or their linear combinations to understand the representations a network has learned. But there are clues that neurons and their linear combinations are not the correct fundamental units of description: directions cannot describe how neural networks use nonlinearities to structure their representations. Moreover, many instances of individual neurons and their combinations are polysemantic (i.e. they have multiple unrelated meanings). Polysemanticity makes interpreting the network in terms of neurons or directions challenging since we can no longer assign a specific feature to a neural unit. In order to find a basic unit of description that does not suffer from these problems, we zoom in beyond just directions to study the way that piecewise linear activation functions (such as ReLU) partition the activation space into numerous discrete polytopes. We call this perspective the polytope lens. The polytope lens makes concrete predictions about the behavior of neural networks, which we evaluate through experiments on both convolutional image classifiers and language models. Specifically, we show that polytopes can be used to identify monosemantic regions of activation space (while directions are not in general monosemantic) and that the density of polytope boundaries reflect semantic boundaries. We also outline a vision for what mechanistic interpretability might look like through the polytope lens.
Cite
Cites 44 works (0 here)
External (44)
- Toy Models of Superposition (2022)
- Origami in N dimensions: How feed-forward networks manufacture linear separability (2022)
- Traversing the Local Polytopes of ReLU Neural Networks: A Unified Approach for Network Verification (2021)
- Knowledge Neurons in Pretrained Transformers (2021)
- An Interpretability Illusion for BERT (2021)
- Multimodal Neurons in Artificial Neural Networks (2021)
- A mathematical framework for transformer circuits (2021)
- Transformer Feed-Forward Layers Are Key-Value Memories (2020)
- Analyzing Individual Neurons in Pre-trained Language Models (2020)
- Curve detectors (2020)
- Thread: Circuits (2020)
- Reverse-engineering deep ReLU networks (2019)
- Deep ReLU Networks Have Surprisingly Few Activation Patterns (2019)
- The Geometry of Deep Networks: Power Diagram Subdivision (2019)
- Towards the neural population doctrine (2019)
- Complexity of Linear Regions in Deep Networks (2019)
- Approximation Theory and Approximation Practice, Extended Edition (2019)
- From Hard to Soft: Understanding Deep Network Nonlinearities via Vector Quantization and Statistical Inference (2018)
- A Spline Theory of Deep Learning (2018)
- Mad Max: Affine Spline Insights Into Deep Learning (2018)
- Sensitivity and Generalization in Neural Networks: an Empirical Study (2018)
- The building blocks of interpretability (2018)
- Network Dissection: Quantifying Interpretability of Deep Visual Representations (2017)
- Feature visualization (2017)
- Why neurons mix: high dimensionality for higher cognition (2016)
- Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks (2016)
- From the neuron doctrine to neural networks (2015)
- Object Detectors Emerge in Deep Scene CNNs (2014)
- Understanding Locally Competitive Networks (2014)
- Intriguing properties of neural networks (2013)
- On the number of response regions of deep feed forward networks with piece-wise linear activations (2013)
- Context-dependent computation by recurrent dynamics in prefrontal cortex (2013)
- The importance of mixed selectivity in complex cognitive tasks (2013)
- Population codes in the visual cortex (2013)
- Rectified Linear Units Improve Restricted Boltzmann Machines (2010)
- Explicit Encoding of Multimodal Percepts by Single Neurons in the Human Brain (2009)
- Temporal complexity and heterogeneity of single-neuron activity in premotor and motor cortex (2007)
- Invariant visual representation by single neurons in the human brain (2005)
- Functional neuroanatomy of face and object processing. A positron emission tomography study (1992)
- Information capacity of the Hopfield model (1985)
- Neural networks and physical systems with emergent collective computational abilities (1982)
- Shape Representation in Parallel Systems (1981)
- Receptive fields, binocular interaction and functional architecture in the cat's visual cortex (1962)
- What the Frog's Eye Tells the Frog's Brain (1959)