PAPER-DIGEST · 2026-08-22
Zhao et al.: People Build Their Own Reusable Parts While Solving Puzzles — Fukai Reads
Cognitive science / learning abstractions / the ruler for difficulty
TL;DR
What I read today is a paper that looks straight at the moment a puzzle solver builds a reusable part of their own. The stage is a task where people recreate a picture on a 10×10 grid. Participants can save any half-finished shape with one button, then place that shape as a single move in the next puzzle. In other words, they grow their own toolbox while playing. Anyone who has played a construction puzzle such as Opus Magnum knows the pleasure of saving a contraption and reusing it.
Opus Magnum (Zachtronics, 2017; from its Steam store page)
Here is the result. The share of moves that used a saved part rose from 21% on the first puzzle to 87% on the fourteenth. On the later puzzles, around nine in ten participants saved the same shape. And the tastiest finding for designers is this. What tracked human time and number of moves was not the length of the shortest program, but the number of candidates searched before the answer was found (r=.82 and r=.79). The length of the shortest program predicted neither time nor moves (r=-.20, p>.05).
The same picture takes fewer moves once you build parts (conceptual diagram, AI-made)
Introduction: Who Wrote It
The paper is titled "Online library learning in human visual puzzle solving." The authors are Pinzhe Zhao (University of Edinburgh), Emanuele Sansone (MIT CSAIL / KU Leuven), Marta Kryven (Dalhousie University) and Bonan Zhao (University of Edinburgh). The text states that Kryven and Bonan Zhao contributed equally. It is posted as arXiv:2603.23244v1; the ID tells us it went up in March 2026, and the licence is CC BY 4.0.
Bibliographic details of today's paper (conceptual diagram, AI-made)
Why did I pick it today? We are on the side that ships a puzzle every day, and we trip over the same question every day: how do you measure difficulty? This paper proposes moving that ruler somewhere else. But let me state the caveats plainly. This is a preprint that has not been peer reviewed. There were 30 participants and a single condition; the authors call it exploratory. The study is preregistered (the procedure and exclusion criteria were posted before running it), but no replication exists yet.
What to keep in mind while reading this paper (conceptual diagram, AI-made)
Background: The Gap
That people cut hard work into parts has been known for a long time. Gobet et al. (2001) organised learning around chunking. Lake et al. (2015) showed that people learn concepts as procedures for drawing them. On the machine side, DreamCoder by Ellis et al. (2023) grows its own dictionary of parts out of the problems it solves.
Where this study stands (conceptual diagram, AI-made)
The gap sits one step earlier. When, and which part, does a person decide to build? The decision is hard because the future is hidden. You choose what to keep without knowing what problem comes next. The authors call this online library learning: writing the dictionary while you solve.
When you pick a part, the future is not visible (conceptual diagram, AI-made)
Building parts is not always a win. It costs time and moves, and it adds to what you must remember. Save a shape you use only once and only the overhead remains. The authors list three reasons this balance is hard to observe: the space of possible parts is wide open, expressivity and efficiency pull against each other, and the task must be rich yet still structured.
Building parts is not always a win (conceptual diagram, AI-made)
The Task They Built
The task is called the Pattern Builder Task. The screen has three parts. On the left, the target shape and the running score. In the middle, a 10×10 working grid. On the right, your program, one line per move. Participants rebuild the target on the grid. There are 14 puzzles, always in the same order, growing harder toward the end.
The Pattern Builder screen (schematic, AI-made)
There are only twelve tools. Five shapes: horizontal line, vertical line, diagonal, square border, triangle. Three operations that combine two things: add, subtract, keep only the overlap. Four that change one thing: invert, reflect left-right, reflect up-down, reflect along the diagonal. All 14 targets are reachable with those twelve. No formulas appear; only buttons you can press.
Only twelve tools are available (conceptual diagram, AI-made)
The heart of it is a plus button. Press it on any half-finished shape and it is saved as a part — a "helper" in the paper's words. Helpers sit on a shelf as thumbnails and persist into later puzzles until you delete them. In the next puzzle, a helper can be placed as a single move. After the 14 puzzles comes a free play phase of at least five minutes, with no target to match.
A part is born with one button (conceptual diagram, AI-made)
The authors also have machines solve the task. Programs are enumerated from short to long, and programs that behave identically are pruned (only one program per resulting picture is kept). In the library versions, a solved pattern is registered as a new part. Four variants are compared, and the key quantity is the number of candidates expanded before the answer is reached.
What the paper's model actually counts (conceptual diagram, AI-made)
What They Found
First, nearly everyone solved the puzzles. Across 30 participants × 14 puzzles = 420 trials, accuracy was 92.4% (SD=14.3%, 95% CI [87.3%, 97.5%]). A total of 470 helpers were created, 15.7 per person on average, with individuals ranging from 0 to 35. Twenty-eight of the 30 created at least one. Time per puzzle averaged 84 seconds (median 44.1s, SD 136.1s), and the median number of steps was 3.0 (M=3.8, SD=3.2).
Nearly everyone solved them (figures quoted from the paper's Results)
Now the main point. The share of solution steps that used a saved helper was 21% on puzzle 1, 80% on puzzle 9 and 87% on puzzle 14. The slope is +3.54% per trial, with r=.79, p<.001. Helpers were also used more on puzzles whose ground-truth program was longer (r=.77, p=.001). When the puzzle got harder, people moved to the side of making tools.
Share of solution steps that used a saved part (values from the paper; the curve shape is illustrative)
The choice of parts did not scatter. Over the first seven puzzles, the most popular shape was saved by only 53–64% of participants. Later, the same shape was saved by 81% on puzzle 8, 94% on puzzle 9, 82% on puzzle 10 and 93% on puzzle 12. Saving the target picture itself also rose, from 50.8% in the first seven trials to 78.5% in the last seven (+3.25% per trial).
Later on, almost everyone saved the same part (figures from the paper's Results)
And then the ruler for difficulty. The number of candidates the model expanded correlated with mean solution time at r=.82 (p<.001) and with mean number of steps at r=.79 (p<.001). Mixed-effects models that account for per-participant variation point the same way (steps β=0.86, z=6.73; time β=0.41, z=9.94; both p<.0001). The length of the shortest program, by contrast, predicted neither time nor steps (r=-.20, p>.05) and correlated only with success rate, negatively (r=-.67, p<.01).
Which ruler predicted difficulty (point positions illustrative; values from the paper's Figure 6A)
How Designers Can Use This
One. Give a daily puzzle a shelf of parts. A line of play you found in today's grid can be saved, and placed as one move in tomorrow's grid. The shelf can be cleared by hand, because hoarding unused parts only adds to the cost of choosing. In the paper, helper usage climbed as the puzzles got harder (21% → 87%). Translated into design: a shelf pays off right before a hard stretch.
Use case 1: a daily puzzle whose parts carry over (design sketch, AI-made)
Two. Replace the difficulty label. The usual "5 moves minimum" or "three stars" speaks only about the board. Leaning on this paper's ruler instead means counting, with your own solver, how wide the search is before the answer appears. In the paper that quantity tracked human effort at r=.82 for time and r=.79 for steps. Puzzles you had filed as easy on move count may well reorder.
Use case 2: swap the ruler behind the difficulty label (design sketch, AI-made)
Three. Reread an existing classic with these eyes. The Witness (Thekla, Inc., 2016) places over 500 puzzles on an island and adds the meaning of one symbol at a time. Players reuse what they have learned inside their heads, but they cannot save it onto the board. This paper's task can be read as pulling that mental shelf out onto the screen. Which also means the design space of a visible shelf is still open.
The Witness (Thekla, Inc., 2016; from its Steam store page)
Four. Let players name the parts they build. The free play observations are the ground for this. In an average of 6.32 minutes, 27 of 30 participants submitted a creation, and 64 of the 80 creations (80.0%) carried a name the participant chose — names like "Fireflower" and "city sky line." Seventeen of 30 (56.7%) built new helpers even with no target to match, and those who had built more during the task were likelier to keep building (β=0.29, p=.005).
Use case 3: let players name their parts (figures from the paper's Results)
Limitations
The authors raise two limits themselves. One is that they only observed people solving alone. Parts passing from person to person, or being named and taught, sit outside this study. The other is the model. Theirs can only turn a finished pattern into a part. People, though, kept a fragment of the middle because it looked useful later. The authors write that extending the model to such partial abstraction would give a fuller account.
Limits the authors acknowledge (conceptual diagram, AI-made)
What Fukai would point at here is the order of the puzzles. All 14 came in the same sequence, harder toward the end. So the rise in helper use splits into two possible causes: participants got better, or the puzzles started demanding parts. With no reordered condition, this design cannot separate them. The authors do stress that participants never saw the whole task space, but the order effect itself is not addressed.
What Fukai noticed: the order is fixed (conceptual diagram, AI-made)
One more. The correlation of r=.82 rests on 14 points. It is a correlation over per-puzzle averages, not over participants. The ordinary caveat applies: high correlations come easily with few points. There were 30 participants, recruited on Prolific at £7.27/hr, in a single condition. It is a preprint with no replication yet. And "candidates expanded" is a number their model emits; swap the model and the number moves.
One more layer of caution when reading the numbers (source: the paper's Experiment / Results)
How Fukai Reads It
From here on this is my reading, not what the paper says. I would place this study in the current that unsettles the assumption that difficulty lives in the board. On the same board, how wide the search must be depends on which parts you happen to have accumulated. If so, difficulty reads less as a property of the puzzle and more as a relation between the puzzle and that person's toolbox. Speaking as someone who ships a daily puzzle, that means one more knob besides "make the board easier."
From here on it is Fukai's interpretation (conceptual diagram, AI-made)
Closing
If you want to go deeper, read the machine side and the human side against each other and a map appears. On the machine side, DreamCoder by Ellis et al. (2023). On the human side, the conceptual bootstrapping work of Zhao et al. (2024, Nature Human Behaviour) and the symbolic search work of Rule et al. (2024, Nature Communications). On the same arXiv sits "Prospective Compression in Human Abstraction Learning" (arXiv:2605.09985), about compressing while looking ahead. I have not read it yet, so I will not summarise it.
Papers that make the map legible (conceptual diagram, AI-made)
Four things to take away. Put a place to build parts inside the board. Try measuring difficulty by the width of the search. Let players name what they build. And the fourth is homework for the research side: measure the effect of order apart from the effect of getting better. The demo the paper points to is still public. Solve the 14 puzzles yourself and you will feel, in your hands, the moment you want to save a shape.
Four things to try tomorrow (conceptual diagram, AI-made)
References
Papers and materials referenced in this article:
・Full HTML text (task design, the 14 targets, the Results figures, the Figure 6 scatter plots)
・The authors' preregistration (AsPredicted)
・The Pattern Builder demo the paper points to (solve the 14 puzzles yourself)
・Prior work this paper builds on: DreamCoder (Ellis et al., 2023), concept learning (Lake et al., 2015, Science), chunking (Gobet et al., 2001), conceptual bootstrapping (Zhao et al., 2024, Nature Human Behaviour), symbolic metaprogram search (Rule et al., 2024, Nature Communications). I name them only as citations inside this paper; I did not reread the originals.
・Related, unread: Prospective Compression in Human Abstraction Learning (arXiv:2605.09985)
・Sources of the game images: the Steam store page for Opus Magnum (Zachtronics, 2017) and the Steam store page for The Witness (Thekla, Inc., 2016). Every other figure is the author's own conceptual diagram (AI-made), not a reproduction of the paper's figures.
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 65 of 103
