PAPER-DIGEST · 2026-09-23

von Gugelberg: How people solve differed more by 'which item' than by how hard it was — Fukai Reads

Cognitive psychology — eye movements on figural puzzles and the effect of sequence

TL;DR

Three hundred people solved figural matrix problems (find the missing cell in a 3-by-3 grid of shapes) while an eye tracker recorded, item by item, where they looked.

People with stronger reasoning spent more time on the problem itself and built an answer in their head before checking the options. When items got harder, stronger reasoners also changed their looking behaviour more.

The most interesting finding comes last: people differed more in how their behaviour shifted as the test went on than in how they reacted to harder items. For anyone who orders puzzles in a sequence, that is hard to ignore.

Who wrote it, and where?

The author is Helene M. von Gugelberg of the Institute of Education at the University of Zurich, Switzerland. It is a single-author paper.

It appears in the Journal of Intelligence, volume 14, issue 9, article 206. It is peer-reviewed: received 26 June 2026, accepted 21 August, published 1 September. It is open access, and the analysis data are posted on OSF (a platform for sharing research data).

I did not find any mention of preregistration (publicly stating hypotheses and analyses before running the study) in what I read. The author herself describes the study as exploratory. I come back to this under limitations.

Why this one today? Because data on how a solver's eyes move between the board and the answer options, tracked item by item across hundreds of people, is exactly what puzzle makers rarely get to see.

Builders and eliminators: what was already known

The task is the figural matrix: a 3-by-3 grid of shapes with the bottom-right cell missing. You work out the hidden rules along rows and columns and pick the answer from eight options. It is a staple of intelligence testing.

Two ways of solving these have long been described, going back to work by Snow in 1978 and 1980. One is constructive matching: study the grid, build the answer in your head, then compare it with the options. The other is response elimination: flick back and forth between grid and options, crossing out candidates that don't fit.

Stronger reasoners tend to use the first. Meta-analyses (statistical syntheses of many studies) broadly support this. Later work found a third strategy (Jarosz et al., 2019) and used eye movements to sort people into three groups (Li et al., 2022). A Baba Is You board with word blocks spelling out rules placed on the gridBaba Is You (Hempuli Oy, 2019). A classic of reading the rules off the board and building the answer yourself. Image: Steam store page

What was less clear was change. Do people use more or less constructive matching as items get harder? Earlier findings disagreed; one report (Liu et al., 2023) found strong reasoners use more and weaker reasoners less. There is also a known item-position effect, where performance behaves differently later in a test. Few studies had looked at difficulty, position and individual differences together.

How eye movements were measured

Participants were university students and community members in Switzerland. Of 319 who took part, 300 were analysed after excluding, for example, those whose eye-tracker calibration failed. Mean age was 26.26 (SD 10.81).

There were two sessions at least a day apart. The first was the Corsi block-tapping task: nine squares light up in sequence and you reproduce the order. It measures visuospatial working memory, the ability to hold locations in mind briefly. Across 21 trials, the mean number correct was 11.96 (SD 2.87).

The second session was the matrix test itself: Form 1 of the Figural Matrices from the Mitres battery (Kyllonen et al., 2019). There were two practice items, a 30-minute overall limit and 2 minutes per item. Difficulty did not rise steadily from front to back. Mean completion time was 22.93 minutes (SD 5.99).

Eye movements were recorded at 500 samples per second with an EyeLink 1000 Plus. Three indices were computed: (1) toggle rate, switches between grid and options divided by time on the item; (2) the proportion of time spent looking at the grid; (3) the proportion of time before the first look at the options. Fewer switches, more grid time and a later first look at the options all point towards building the answer.

The analysis used Bayesian multilevel models, a statistical framework that handles differences between people and between items at the same time. Reasoning ability, working memory, item difficulty and item position went in together, along with their interactions.

What was found

First, stronger reasoners leaned towards building the answer on all three indices: fewer switches, a larger share of time on the grid and a later first look at the options. That matches earlier work, though the author notes in her conclusion that the effects were 'generally smaller' than previously reported.

Second, harder items came with lower toggle rates and a larger share of time on the grid, and this shift was larger for stronger reasoners. Faced with a hard item, strong reasoners settle in and read the grid more carefully.

Third, as the test went on, the share of time on the grid fell and the share of time before the first look at the options rose. Time per item dropped slightly, yet accuracy did not decline, which the author takes as making growing disengagement unlikely. First-person view walking the uninhabited island of The WitnessThe Witness (Thekla, Inc., 2016). Solving panel after panel, the way you read the rules grows along the way. Image: Steam store page

Fourth, and central: between-person variability in how behaviour changed with item position exceeded the variability in how it changed with difficulty. In the author's words, individuals differ more strongly in how their behaviour evolves across the test than in their responses to increasing difficulty.

Also, people who on average spent less time on the grid showed stronger position effects. The author reads this as some solvers gradually realising that success depends on identifying the underlying rules. Working memory, by contrast, showed no consistent effect, appearing only in certain interactions.

The specific estimates are in Tables 3 to 5 of the paper. I could not transcribe those numbers with confidence from what I was able to read, so this article reports only the direction of each effect.

How puzzle makers can use this

What follows is my own extrapolation, not something the paper tested.

(1) In multiple-choice puzzles, treat when the options appear as a design lever. Show the board first and reveal candidates a moment later, and you may nudge players towards building the answer. If matching against candidates is the fun you want, show them from the start. The same puzzle can play differently either way.

(2) Log back-and-forth as a sign of struggle. A sudden rise in trips between board and hint panel, or board and candidate list, may signal a switch to eliminating. But toggle rates varied a lot between people in this study, so watch each player's change from their own baseline rather than using one threshold for everyone. A screenshot from Islands of InsightIslands of Insight (Lunarch Studios / Behaviour Interactive, 2024). Many puzzle types share one world, and players pick the order. Image: Steam store page

(3) Don't order levels by a difficulty curve alone. The largest individual differences here were in change across positions. If you build daily puzzles or long level sequences, watching whether a player's approach is shifting from the early levels may matter as much as per-level difficulty estimates.

(4) Give struggling players a way of looking, not an answer. Strong reasoners changed approach more on hard items; weaker ones may not make that switch as readily. A hint like 'look at one row only and put into words what changes' prompts a reading procedure instead.

(5) Use early levels to show that reading the rules pays off. The author suggests some solvers come to see the importance of the rules as they go. Tutorial levels arranged to speed up that realisation might change how the rest of the game is played.

How far to trust it

The author acknowledges three weaknesses. The study was exploratory, not designed around a target effect, so the interactions found should be read with caution. Working memory was measured with a single, fairly simple task. And although the sample was larger and wider in age than earlier studies, it is far from representative of the general population.

Fukai's own points are these. First, the material is an intelligence-test matrix, not a game. Solving under a time limit with your head on a chin rest in a lab is very different from playing a puzzle on the sofa. My use cases above leap across that gap.

Second, this is observational. It does not show that making people build the answer first would raise accuracy; that needs an experiment that manipulates the approach. Third, from what I read, I could not tell clearly how the reasoning score was computed. If it comes from the same test, the link to the eye-movement indices needs careful interpretation.

The author also notes that effects were smaller than earlier reports and that the raw number of toggles has been inconsistent across studies. I would treat this not as a general law of the mind but as what was observed under these conditions.

Fukai's reading

This part is my opinion. I read the study as a quiet objection to measuring difficulty on a single scale. When we build levels we instinctively draw a difficulty curve. But here, what separated people was not their reaction to that curve; it was how their way of looking grew as they moved through the sequence. In design-critique terms, a level sequence is a curriculum for ways of seeing before it is a staircase of difficulty. That is probably exactly what The Witness was doing with its island of panels.

What to read next

On insight and switching how you look, the paper by Dygert & Jarosz I covered earlier connects directly: it linked recovering from misread sentences to solving insight problems.

On how presentation changes solvability, read Macchi et al. alongside this. On sequencing, Ponnock & Ho on the order of Mario 1-1 and Han et al. on when learning order matters widen the map.

Sources

Papers and related material referenced in this article:

・Individual Differences in Problem-Solving Behavior in Reasoning Tests: The Role of Item Difficulty and Item Position (Helene M. von Gugelberg, 2026, Journal of Intelligence 14(9), 206, peer-reviewed)

・Open-access full text of the same paper (MDPI)

・Analysis data (OSF)

Reactions (no login)

Anonymous • one of each per visitor per day

関連シリーズ

Paper Digest第93回 / 全98回

Read next