Reference. Structured World Representations in Maze-Solving Transformers

Transformer models underpin many recent advances in practical machine learning applications, yet understanding their internal behavior continues to elude researchers. Given the size and complexity of these models, forming a comprehensive picture of their inner workings remains a significant challenge. To this end, we set out to understand small transformer models in a more tractable setting: that of solving mazes. In this work, we focus on the abstractions formed by these models and find evidence for the consistent emergence of structured internal representations of maze topology and valid paths. We demonstrate this by showing that the residual stream of only a single token can be linearly decoded to faithfully reconstruct the entire maze. We also find that the learned embeddings of individual tokens have spatial structure. Furthermore, we take steps towards deciphering the circuity of path-following by identifying attention heads (dubbed 𝑎𝑑𝑗𝑎𝑐𝑒𝑛𝑐𝑦 ℎ𝑒𝑎𝑑𝑠), which are implicated in finding valid subsequent tokens.

Cite

Cite as @ivanitskiy-2023-structured (helia, typst) · \cite{ivanitskiy-2023-structured} (LaTeX)
BibTeX
bibtex · 8 lines
@misc{ivanitskiy-2023-structured,
  author = {Michael Ivanitskiy and Alexander F. Spies and Tilman Räuker and Guillaume Corlouer and Chris Mathwin and Lucia Quirke and Can Rager and Rusheb Shah and Dan Valentine and Cecilia Diniz Behn and Katsumi Inoue and Samy Wu Fung},
  title = {Structured World Representations in Maze-Solving Transformers},
  year = {2023},
  month = {12},
  eprint = {2312.02566},
  archiveprefix = {arXiv}
}
hayagriva YAML (typst)
yaml · 19 lines
ivanitskiy-2023-structured:
  type: misc
  title: Structured World Representations in Maze-Solving Transformers
  author:
  - Ivanitskiy, Michael
  - Spies, Alexander F.
  - Räuker, Tilman
  - Corlouer, Guillaume
  - Mathwin, Chris
  - Quirke, Lucia
  - Rager, Can
  - Shah, Rusheb
  - Valentine, Dan
  - Behn, Cecilia Diniz
  - Inoue, Katsumi
  - Fung, Samy Wu
  date: 2023-12
  serial-number:
    arxiv: '2312.02566'
Cited by (1)

maze-dataset: Maze Generation with Algorithmic Variety and Representational Flexibility ivanitskiy-2025-maze

DOI
Cites 20 works (1 here)
With notes (1)

A Configurable Library for Generating and Manipulating Maze Datasets ivanitskiy-2023-a

Understanding how machine learning models respond to distributional shifts is a key research challenge. Mazes serve as an excellent testbed due to varied generation algorithms offering a nuanced platform to simulate both subtle and pronounced distributional shifts. To enable systematic investigations of model behavior on out-of-distribution data, we present 𝚖𝚊𝚣𝚎-𝚍𝚊𝚝𝚊𝚜𝚎𝚝, a comprehensive library for generating, processing, and visualizing datasets consisting of maze-solving tasks. With this library, researchers can easily create datasets, having extensive control over the generation algorithm used, the parameters fed to the algorithm of choice, and the filters that generated mazes must satisfy. Furthermore, it supports multiple output formats, including rasterized and text-based, catering to convolutional neural networks and autoregressive transformer models. These formats, along with tools for visualizing and converting between them, ensure versatility and adaptability in research applications.
arXiv
External (19)
ivanitskiy-2023-structured reference entries/refs/ivanitskiy-2023-structured/ivanitskiy-2023-structured.hel