Reference. Structured World Representations in Maze-Solving Transformers
Transformer models underpin many recent advances in practical machine learning applications, yet understanding their internal behavior continues to elude researchers. Given the size and complexity of these models, forming a comprehensive picture of their inner workings remains a significant challenge. To this end, we set out to understand small transformer models in a more tractable setting: that of solving mazes. In this work, we focus on the abstractions formed by these models and find evidence for the consistent emergence of structured internal representations of maze topology and valid paths. We demonstrate this by showing that the residual stream of only a single token can be linearly decoded to faithfully reconstruct the entire maze. We also find that the learned embeddings of individual tokens have spatial structure. Furthermore, we take steps towards deciphering the circuity of path-following by identifying attention heads (dubbed ), which are implicated in finding valid subsequent tokens.
Cite
Cited by (1)
maze-dataset: Maze Generation with Algorithmic Variety and Representational Flexibility ivanitskiy-2025-maze
Cites 20 works (1 here)
With notes (1)
A Configurable Library for Generating and Manipulating Maze Datasets ivanitskiy-2023-a
Understanding how machine learning models respond to distributional shifts is a key research challenge. Mazes serve as an excellent testbed due to varied generation algorithms offering a nuanced platform to simulate both subtle and pronounced distributional shifts. To enable systematic investigations of model behavior on out-of-distribution data, we present , a comprehensive library for generating, processing, and visualizing datasets consisting of maze-solving tasks. With this library, researchers can easily create datasets, having extensive control over the generation algorithm used, the parameters fed to the algorithm of choice, and the filters that generated mazes must satisfy. Furthermore, it supports multiple output formats, including rasterized and text-based, catering to convolutional neural networks and autoregressive transformer models. These formats, along with tools for visualizing and converting between them, ensure versatility and adaptability in research applications.
External (19)
- Understanding and Controlling a Maze-Solving Policy Network (2023)
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEval (2023)
- Uncovering mesa-optimization algorithms in Transformers (2023)
- Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis (2023)
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla (2023)
- Can Transformers Learn to Solve Problems Recursively? (2023)
- Finding Neurons in a Haystack: Case Studies with Sparse Probing (2023)
- Towards Automated Circuit Discovery for Mechanistic Interpretability (2023)
- Eliciting Latent Predictions from Transformers with the Tuned Lens (2023)
- Progress measures for grokking via mechanistic interpretability (2023)
- Actually, othello-gpt has a linear emergent world model (2023)
- What learning algorithm is in-context learning? Investigations with linear models (2022)
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small (2022)
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task (2022)
- Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks (2022)
- Towards Understanding Grokking: An Effective Theory of Representation Learning (2022)
- In-context learning and induction heads (2022)
- A Mechanistic Interpretability Analysis of a GridWorld Agent-Simulator (Part 1 of N)
- Towards monosemanticity: Decomposing language models with dictionary learning