PAPER-DIGEST · 2026-08-17

Pereira & Zuidema: Reasoning Models Build a Map of the Tower of Hanoi, Then Lose It — Fukai Reads

Emergent world models / mechanistic interpretability / puzzle planning

TL;DR

The Tower of Hanoi is a classic puzzle in which you move one ring at a time between three pegs; it has long been used to assess planning ability, including in children. And yet today's reasoning models — large language models that write out a long chain of intermediate thought before answering — still solve only about half of one particular variant optimally. The paper I read today looks inside the model to locate exactly where that failure happens.

What the authors found is mildly surprising. At the moment the model finishes reading the prompt, it already holds a geometrically faithful map of the puzzle's entire state space inside itself. It is not failing for lack of a map. The map fades while the model is writing out the moves. The authors frame this as a failure of maintenance, not an absence of a world model.

When they pushed the fading representation back with an external intervention, Qwen3.6-27B recovered from 33/81 (41%) optimal solutions to 59/81 (73%). Read from a puzzle maker's chair, this is a paper arguing that the crux of 'let an LLM solve it / hint at it' design is not more intelligence but how you keep the board state alive. Note that this is an arXiv preprint (submitted 7 August 2026) and has not been peer-reviewed.

Introduction

The authors are Devin Pereira and Willem Zuidema, both at the University of Amsterdam. The paper is titled 'Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking', posted as arXiv:2608.07077 on 7 August 2026, categorised under cs.AI. It is an arXiv preprint, not a peer-reviewed conference paper. It is only days old and has essentially no citations yet, so I read it as a finding that has not been through wide discussion.

My reason for picking it today is simple: the subject matter is a puzzle. The Tower of Hanoi needs no introduction for readers of this site. And the paper treats it not as 'can the AI solve it' but as 'how is the board represented inside the AI, and when does that representation break'. For anyone who designs puzzles, the second question is far more usable.

There is a second reason. The paper is a response to Shojaee et al.'s much-debated 2025 work, The Illusion of Thinking. That paper showed reasoning models collapsing as problem complexity rises, and popularised the reading that they only appear to think. This paper watches that collapse from inside the model and proposes a different reading. Put bluntly: not 'it never thought' but 'it forgot partway through'.

Background

The backdrop is a line of work on 'emergent world models' — the phenomenon where a model trained only to predict the next token nonetheless grows an internal picture of the board or situation. Li et al. (2022) showed that board positions can be decoded from a model trained purely on Othello move sequences; Toshniwal et al. (2022) reported similar things in chess, and Spies et al. (2025) in maze solving.

But as this paper notes, that line has been almost entirely confined to small Transformers trained on games or mazes. Whether the huge, many-layered reasoning models carry the same kind of world model is, in the authors' phrasing, 'yet to be established'. Does a phenomenon visible in a toy model appear in the same shape at frontier scale? That was the gap.

The other half of the backdrop is the puzzle itself. The tower-to-tower form — every ring stacked on one peg, move them all to another — has a famous recursive solution and is effectively saturated for current reasoning models. In the flat-to-flat form, where both the initial and goal configurations may spread rings across pegs, the optimal path depends on the specific (initial, goal) pair and cannot come from a memorised recursive template. There, frontier models drop to roughly 51% optimal.

Approach

The study comes in two halves. The first is an in-house toy model. The authors trained a GPT-2-style decoder-only Transformer (a model that predicts the next token from what came before) purely on symbolically generated flat-to-flat solution traces. The size: 6 layers, hidden dimension 128, 4 heads, 50 epochs. With 4 rings and 3 pegs there are 81 valid configurations, yielding 6,480 ordered (initial, goal) pairs with the two differing; these were split into 5,184 training and 1,296 validation problems.

The instrument for looking inside is a linear probe — a small readout that tests whether the information you care about can be pulled out of the model's internal state by a simple linear transformation. This probe is a slightly unusual one: it maps each configuration to a point in two dimensions and is trained so that the distances between those points match the move-distances between configurations on the puzzle's graph. It measures not 'can we name the configuration' but 'is the relative geometry between configurations preserved'.

Why distance? Because the Tower of Hanoi's state space arranges itself into a Sierpiński triangle — one big triangle containing three smaller ones, each containing three smaller ones again. Legal moves correspond to steps between neighbouring points in that figure, and the three corners correspond to configurations with every ring on a single peg. The three sub-triangles partition the space by which peg holds the largest ring. Whether that figure is visible inside the model is the focus of the first half.

The second half establishes causality. The tools are activation patching (swapping part of the model's internal state, mid-problem, for the internal state from a different problem, and seeing whether the output moves toward that other problem's answer) and activation steering (adding a directional vector to the internal state to push the output). Patching asks whether the representation is actually used; steering asks whether restoring a damaged representation restores performance.

Findings

Start with the toy model. At the separator token between problem and solution (SEP), the configuration could be read out of the internals almost perfectly. The probe's rank correlation (Spearman) was 0.938 and its value correlation (Pearson) 0.902, both peaking at layer 5. Nearest-state retrieval accuracy was 100.0%, and per-ring classification was 100.0% for all four rings. The geometry, the authors write, assembles by middle depth.

What follows is the interesting part. At SEP, the per-ring information sits bundled together in one shared, overlapping space (principal angles between per-ring subspaces average 54.7°). At the move-emission tokens it reorganises into a near-orthogonal arrangement (76.5°). Each ring becomes independently readable — but the larger rings degrade: ring 2 falls to 90.7% and ring 3 to 79.25%. The authors call this the unified-to-factored shift. Patching agrees: at layer 6, partial transfer reaches 78.74% and the rate of outright disruption is 0.00% (against 12.42% at layer 4). The deeper the layer, the more the representation is actually in use.

Now the large models: Qwen3.6-27B and DeepSeek-R1-Distill-Qwen-32B, both 64 layers with hidden dimension 5,120. At the end of the prompt (the authors' Position A), rank correlations were 0.935 and 0.936 with nearest-state retrieval at 1.00 — indistinguishable in fidelity from the toy model. But at the commitment point after the full chain-of-thought, just before the move list (Position B), Qwen keeps global geometry around 0.92 while nearest-state retrieval drops and per-ring classifiers collapse to near chance. DeepSeek falls further, to 0.74–0.81. During move emission (Position C), per-ring encoding partially returns — the same unified-to-factored shape seen in the toy model.

Finally the intervention. The authors cached the internal states corresponding to each correct configuration and, during generation, injected a vector pointing toward them; the current configuration was maintained externally by a symbolic tracker replaying the moves from the start state. For Qwen3.6-27B (layer 28), even the weakest setting converted 52% of failures into optimal solutions, and the strongest lifted optimal solves from 33/81 (41%) to 59/81 (73%); the residual failures were illegal moves, not formatting errors. DeepSeek-R1-Distill (layer 44) resisted: at best it rescued only 6 of 72 failures, against 26 of 48 for Qwen, and cranking the strength up inflated unparseable outputs from 29 of 72 to 60 of 72.

How Puzzle Makers Can Use This

First: when you put a hint system or an auto-solver into a puzzle game, do not make the LLM remember the board. Say you are building a Sokoban-like and want an LLM to return 'the next move' to a stuck player. On this paper's reading, holding the board internally across a long response is the most fragile part of the pipeline. Hand it the current board as text at every move and let it think one move deep. It is a boring design — but the fact that the paper itself needed an external tracker to rescue the model is a decent argument for that boredom.

Second: borrow the flat-to-flat trick as a difficulty tool. The paper's starting point was that varying not just the initial state but the goal state as well kills memorised templates. The same operation is available in human-facing puzzles: change no rules, simply free up both the initial and the goal configuration, and solving shifts from recall to search. Whether that is fun for humans is a separate question, and one the paper does not address.

Third: reinterpret LLM-driven auto-playtesting. If you treat 'the LLM could not solve it' as a difficulty signal, this paper puts a dent in that metric — the model may have failed not because the level was hard but because the solution was long enough that it lost track of the state. Practically: log where in the sequence the model starts to derail. If that point depends strongly on solution length, treat it as a limit of your instrument, not a property of the level. A cheap control is to re-present the board every few moves and see whether the derailment point moves.

Fourth, a takeaway for the human side. The failure shape — the map is there, but it fades while you are writing out the moves — rhymes with working memory in humans solving long puzzles. If so, the UI-side moves are obvious: re-show the board, visualise the move history, allow rewind to any point. I should be explicit, though: the paper says nothing about humans. This is an analogy I am carrying over myself.

Limitations

The authors are candid about their weak points. The experiments cover a single puzzle at a single size (4 rings), and the Tower of Hanoi is an unusually clean recursive structure; they explicitly question whether the findings generalise to less geometric state spaces. Both frontier models are Qwen-derived, so transfer claims should be read within that caveat. Methodologically, they note that a supervised probe can fit structure the model does not actually use — partially mitigated, in their own assessment, by the patching and steering experiments. They also flag that their rescue depends on an external state tracker, and that the ability to steer a world model is itself a governance concern.

What I want to flag here is the standing of that external tracker. The tracker is a device that already knows the correct board. So what the intervention demonstrates is that the model can be fixed by something that already has the answer state — sound as a causal demonstration, circular if read straight as an applied technique. If you already hold the correct board, there are many situations where you no longer need the model to solve anything. The paper does list this under scalability, but I would put it more strongly: it is the largest trap for a practitioner reading this work.

One more, about reading the numbers. The paper reports two figures for Qwen's optimal-solve rate: roughly 51% under maximal reasoning budget, and 33/81 (41%) as the baseline of the steering experiment. The conditions differ, so this is not a contradiction — but summarising it as '51% became 73%' would be wrong. And since the state space has only 81 configurations, the fact that a clean geometry shows up in a probe is not by itself very surprising. What is surprising, as I read it, is the observation along the time axis: that the geometry decays mid-generation.

Fukai's Reading

From here it is my own reading. I would place this study in a line of work that shifts the axis of evaluation from 'is the model smart' to 'can the model keep hold of what it knew'. Since The Illusion of Thinking, the discussion has largely received the observed collapse as a story about a ceiling on capability. This paper inserts a time axis into it: the same model, on the same problem, holds an accurate map at the first instant and does not hold it several hundred tokens later. Translated into the vocabulary of puzzle design, this is not a conversation about difficulty but about board presentation. The reason we keep the board on screen for human players, log their moves, and let them rewind is not that we distrust human memory — it is that where memory is not the point of the puzzle, not demanding it is simply the cleaner design. It reads to me as though the same courtesy is now due to machines.

Closing

If you want to place this paper on a map, start with Shojaee et al. (2025), The Illusion of Thinking. It is the work this paper responds to, and the collapse phenomenon itself is measured more carefully there. Read alongside it the public comment on that paper (Lawsen, 2025), which is instructive about how much the conclusion can turn on the details of the experimental setup.

If your interest runs toward internal representations instead, Li et al. (2022) on Othello remains the best entry point. Once you have that single image in mind — a board rising inside a model trained only on move sequences — this paper's claim that a Sierpiński triangle is visible reads not as an oddity but as one step in a lineage. Today's paper is an early attempt to bridge that lineage from toy models to frontier reasoning models. Whether the bridge holds is a question for peer review and replication.

Sources

Papers and materials referenced in this article:

・Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking (Devin Pereira, Willem Zuidema, University of Amsterdam, 2026, arXiv preprint arXiv:2608.07077 / not peer-reviewed)

・Full HTML text of the same paper (all figures quoted in this article come from here)

・Related work: The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity (Parshin Shojaee et al., 2025, arXiv:2506.06941)

・Related work: Comment on The Illusion of Thinking (A. Lawsen, 2025, arXiv:2506.09250)

・Related work: Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task (Kenneth Li et al., 2022, arXiv:2210.13382)

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 61 of 99

Read next