Reference. On Logical Extrapolation for Mazes with Recurrent and Implicit Networks
Recent work suggests that certain neural network architectures – particularly recurrent neural networks (RNNs) and implicit neural networks (INNs) – are capable of logical extrapolation. When trained on easy instances of a task, these networks (henceforth: logical extrapolators) can generalize to more difficult instances. Previous research has hypothesized that logical extrapolators do so by learning a scalable, iterative algorithm for the given task which converges to the solution. We examine this idea more closely in the context of a single task: maze solving. By varying test data along multiple axes – not just maze size – we show that models introduced in prior work fail in a variety of ways, some expected and others less so. It remains uncertain whether any of these models has truly learned an algorithm. However, we provide evidence that a certain RNN has approximately learned a form of ‘deadend-filling’. We show that training these models on more diverse data addresses some failure modes but, paradoxically, does not improve logical extrapolation. We also analyze convergence behavior, and show that models explicitly trained to converge to a fixed point are likely to do so when extrapolating, while models that are not may exhibit more exotic limiting behavior such as limit cycles, even when they correctly solve the problem. Our results (i) show that logical extrapolation is not immune to the problem of goal misgeneralization, and (ii) suggest that analyzing the dynamics of extrapolation may yield insights into designing better logical extrapolators.
Cite
Cited by (1)
maze-dataset: Maze Generation with Algorithmic Variety and Representational Flexibility ivanitskiy-2025-maze
Cites 43 works (1 here)
With notes (1)
A Configurable Library for Generating and Manipulating Maze Datasets ivanitskiy-2023-a
Understanding how machine learning models respond to distributional shifts is a key research challenge. Mazes serve as an excellent testbed due to varied generation algorithms offering a nuanced platform to simulate both subtle and pronounced distributional shifts. To enable systematic investigations of model behavior on out-of-distribution data, we present , a comprehensive library for generating, processing, and visualizing datasets consisting of maze-solving tasks. With this library, researchers can easily create datasets, having extensive control over the generation algorithm used, the parameters fed to the algorithm of choice, and the filters that generated mazes must satisfy. Furthermore, it supports multiple output formats, including rasterized and text-based, catering to convolutional neural networks and autoregressive transformer models. These formats, along with tools for visualizing and converting between them, ensure versatility and adaptability in research applications.
External (42)
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach (2025)
- A Generalization Bound for a Family of Implicit Networks (2024)
- Three-Operator Splitting for Learning to Predict Equilibria in Convex Games (2024)
- Deep learning for accelerated and robust MRI reconstruction (2024)
- Learning to Reason with LLMs (OpenAI) (2024)
- Differentiating Through Integer Linear Programs with Quadratic Regularization and Davis-Yin Splitting (2023)
- Maze Dataset (software) (2023)
- Path Independent Equilibrium Models Can Better Exploit Test-Time Computation (2022)
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals (2022)
- Online Deep Equilibrium Learning for Regularization by Denoising (2022)
- Explainable AI via learning to optimize (2022)
- Deep Equilibrium Optical Flow Estimation (2022)
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without Overthinking (2022)
- The Evolution of Out-of-Distribution Robustness Throughout Fine-Tuning (2022)
- Is Attention Better Than Matrix Decomposition? (2021)
- Datasets for Studying Generalization from Easy to Hard Examples (2021)
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent Networks (2021)
- SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit models (2021)
- Feasibility-based fixed point networks (2021)
- JFB: Jacobian-Free Backpropagation for Implicit Networks (2021)
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training Data (2021)
- Deep Equilibrium Architectures for Inverse Problems in Imaging (2021)
- Measuring Robustness in Deep Learning Based Compressive Sensing (2021)
- Thinking Deeply with Recurrence: Generalizing from Easy to Hard Sequential Reasoning Problems (2021)
- Comparison of Hand Follower and Dead-End Filler Algorithm in Solving Perfect Mazes (2020)
- Monotone operator equilibrium networks (2020)
- Multiscale Deep Equilibrium Models (2020)
- Scaling Laws for Neural Language Models (2020)
- The Troublesome Kernel: On Hallucinations, No Free Lunches, and the Accuracy-Stability Tradeoff in Inverse Problems (2020)
- Deep Equilibrium Models (2019)
- Implicit Deep Learning (2019)
- Universal Transformers (2018)
- A User's Guide to Topological Data Analysis (2017)
- (Quasi)Periodicity Quantification in Video Data, Using Topology (2017)
- Sliding Windows and Persistence: An Application of Topological Methods to Signal Analysis (2013)
- Topological Analysis of Recurrent Systems (2012)
- Detecting strange attractors in turbulence (1981)
- Goal Misgeneralization in Deep Reinforcement Learning
- Algorithm Design for Learned Algorithms
- Learning to Optimize: Where Deep Learning Meets Optimization and Inverse Problems
- On Training Implicit Models
- The clrs algorithmic reasoning benchmark