PAPER-DIGEST · 2026-09-14

Dygert & Jarosz: People who repair a misread sentence also solve insight puzzles — Fukai Reads

Cognitive psychology — the restructuring shared by problem solving and reading

TL;DR

People who can recover on their own after misreading a sentence also tended to solve insight puzzles well. That is the conclusion of a paper published on 1 July 2026 in the peer-reviewed journal Journal of Intelligence. Re-reading and insight seem to rest on the same ground.

There were two experiments, with 85 and 97 undergraduates. The link survived after statistically removing working memory capacity and figural-analogy intelligence. It did not show up at all for problems that only require following a procedure.

In numbers: the partial correlation in Experiment I was r = 0.33 (p = .002), about 11% shared variance. In Experiment II, performance on ambiguous sentences added 5–9% of explanatory power to anagram scores. The samples are small, and neither study was preregistered.A screen from Baba Is You, where word blocks on the floor spell out the rulesBaba Is You (Hempuli Oy, 2019). The words on the board are the rules; push them around and the meaning changes. Image: Steam store page

Who wrote it, and where?

The authors are Sarah K. C. Dygert and Andrew F. Jarosz, both in the Department of Psychology at Mississippi State University. The paper appears in Journal of Intelligence, volume 14, issue 7, article 128, published on 1 July 2026.

One point up front. This is not a preprint sitting on arXiv; it is a peer-reviewed, open-access article. The authors' own keywords are creativity, problem-solving, ambiguity and comprehension.

I picked it for a simple reason. People who make puzzles are trying to design the "aha" every single day, and this paper explains that moment using a completely different drawer: re-reading a sentence. As of September 2026 it has almost no citations yet, so it has not been widely debated.

What is the "restructuring" said to be behind insight?

Restructuring — rebuilding your representation of a problem when the first approach stalls — sits at the centre of insight research. The paper puts it this way: "Restructuring then leads to a solution and 'insight,' in which the solution appears suddenly without prior awareness."

The other lead character is the garden path sentence. In the paper's definition, these are sentences where "structurally- and/or semantically ambiguous units are initially misinterpreted." You read on, hit a contradiction, and have to go back and rebuild the sentence.

Here is the textbook example, offered as my own gloss rather than a quote: "The horse raced past the barn fell." Most readers take "raced" as the main verb and stall at "fell". The correct reading is that the horse which was raced past the barn fell down.

These two topics have lived in separate literatures. But they share an oddity: both are reported to depend only weakly on working memory capacity, the amount of information you can hold at once. The authors took that as a hint that a shared process might be at work.

How did they test whether the process is shared?

Experiment I ran with 85 undergraduates (25 male, 60 female) from General Psychology at Mississippi State University, all native English speakers. The session took about an hour.

The language task had 50 sentences, of which 12 were garden path sentences and 38 were unambiguous. Comprehension was measured by truth-value judgements. On the insight side there were 29 usable anagrams (rearrange the letters into a word) and 24 rebus problems (a phrase expressed by the layout of words and pictures), each with a 30-second limit.

The design hinges on a subtraction. Ordinary reading skill — accuracy on the unambiguous sentences — was removed statistically before comparing garden path accuracy with insight performance. That kills the boring explanation that good readers are simply good at everything.

Experiment II used a separate group of 97 students (32 male, 65 female) who had not taken part before. Working memory was measured with the Symmetry Span task and fluid intelligence with 25 figural analogies.A screen from Understand, drawing a line on a dotted board to guess the hidden ruleUnderstand (Artless Games, 2020). The rule is never stated; you draw lines and read the feedback. Image: Steam store page

Problems were then split in two. The creative type was the anagram task; the analytic type was modular arithmetic, 48 items, where a fixed procedure produces the answer and nothing needs revising. The sentence set added 12 ambiguous relative-clause sentences and 24 controls, answered as multiple choice with a "Both could be correct" option.

What did they find?

The headline number from Experiment I: with comprehension of unambiguous sentences removed, the partial correlation was r = 0.33, p = .002. The paper says this "suggests 11% of the variance is shared between these constructs (R2 = 0.11)".

The reverse pairing shows nothing. Insight scores against accuracy on the unambiguous sentences, controlling for garden path accuracy, gave r = 0.14, p = .215. A Bayesian comparison favoured the model including ambiguous sentences by a factor of 19.59. Reading times showed the same weak pattern (r = 0.23, p = .04).

In Experiment II, entering working memory first still left ambiguous-sentence performance adding 9% (ΔR² = 0.09) to creative problem solving. Entering fluid intelligence first left 5% (ΔR² = 0.05). Split by sentence type: garden path ΔR² = 0.06 (B = 0.26, t(96) = 2.64, p = .010); relative clause ΔR² = 0.10 (B = 0.90, t(96) = 3.48, p = .001).

The control condition does real work. The same sentence scores did not predict analytic modular arithmetic at all: ΔR² < 0.01, with a Bayes factor of 0.30, meaning the data were 3.31 times more likely under the null.A screen from Return of the Obra Dinn, a frozen shipboard scene in two-tone ditheringReturn of the Obra Dinn (Lucas Pope (3909 LLC), 2018). The real deduction starts the moment you notice your first reading was wrong. Image: Steam store page

The nicest result is a finer split of the relative-clause items. Only the sentences that genuinely require revision (RC-True) predicted creative problem solving (ΔR² = 0.05, B = 0.23, t(96) = 2.42, p = .017). The ones that do not require it (RC-NP1) predicted nothing (B = −0.007, p = .941). What matters is not meeting ambiguity but repairing it.

The authors' own wording is careful: the two studies are "consistent with restructuring during creative problem solving drawing upon the same processes as the revisions necessary to understand ambiguous language." Consistent with — not proof of.

How can puzzle makers use this?

One line to take away first. If this framework holds, puzzle difficulty can be designed along two separate axes: how much the player must hold in mind, and how many assumptions they must throw away. The evidence is Experiment II: after removing working memory and fluid intelligence, revision ability still explained 5–9% of creative problem solving.

Use one: treat misdirection as an explicit level-design material. Baba Is You is a meta-puzzle game released by Hempuli Oy in 2019 in which you push the rule sentences themselves (review). Most of its hard moments are not about memory load; they are about breaking the role you first assigned to an object — structurally the same move a garden path sentence demands.

Use two: rethink hints. The default hint adds information. If this result holds, the effective hint may be the one line that makes the player doubt their first reading: is that really a wall? does that number really count objects? The fact that only RC-True items mattered points that way.

Use three: build the tutorial curve on two axes. A level that adds elements to remember and a level that betrays an assumption you taught earlier probably tax different abilities. Filing both under "difficulty 5" will make you misread where players get stuck.

Use four: treat the answer format as a design surface. Experiment II offered a "Both could be correct" option, which the authors flag as a possible confound. From a designer's side it is the opposite: adding one slot for ambiguity to the answer menu makes players start doubting their own reading.

Use five: treat unexplained rules as their own kind of difficulty. Understand is a logic-deduction puzzle game released by Artless Games in 2020 in which you probe a hidden rule by drawing lines (review). What it demands is less memory than speed at dropping your own hypothesis.

For the record: this is correlational work. "Add misdirection and creativity grows" is not a claim you can take from this paper. What you can take is a ruler for separating kinds of difficulty.

How far should we trust this?

Start with the weaknesses the authors state. Measurement reliability was low: α = 0.60 for the ambiguous garden path sentences and α = 0.40 for the controls, which they note can both suppress correlations and destabilise effect-size estimates.

Second, restructuring itself was never measured directly. In their words, "there are no direct measures of restructuring in this study", and they call the findings "an initial step".

Third, the exploratory response-time analyses were inconsistent in Experiment II. The authors also concede that the "Both could be correct" option may have pushed participants into revising rather than sticking with a first interpretation.

Now the points I noticed as a reader. First, neither experiment was preregistered, as the paper states outright. Second, the samples are 85 and 97 undergraduates from a single university, with a skewed sex ratio. That is thin ground for statements about people in general.

Third, Experiment II represents creative problem solving with anagrams alone; the rebus task from Experiment I is gone, so the construct differs slightly between studies. Fourth, garden path items lean hard on English word order, and this paper says nothing about whether the same link appears in other languages.

How Fukai reads it

This paragraph is my own reading. I want to place this study as an attempt to slip a second ruler into a field that has measured difficulty as one solid thing. Puzzle difficulty is usually discussed in quantities on the board: move counts, branching, how much you must remember. What this paper adds is a quantity on the player's side — the work of letting go of your own first interpretation. In the vocabulary of design criticism, it comes close to counting misdirection as a mechanic rather than as staging. The authors themselves only claim consistency; I am taking one step past that, and I want to say so plainly.

What to read next for the wider map

If you want to widen the map from the insight side, read it alongside our earlier piece on Chao et al., where insight looks like searching far afield. Today's paper views insight from the act of dropping a first reading; that one views it from how far the search ranges.

From the language side, follow the "good enough" processing literature in sentence comprehension: readers often fail to fully revise and keep dragging their first interpretation along. Today's paper can be read as measuring individual differences in the ability to cut that drag.

The data are open. The authors posted both studies to the Open Science Framework (the paper records an access date of 4 April 2026). Letting a sceptical reader check for themselves is one of this paper's better qualities.

References

Papers and material referenced in this article:

Navigating the Garden Path: Evidence for Restructuring as a Shared Process (Sarah K. C. Dygert & Andrew F. Jarosz, 2026, Journal of Intelligence 14(7):128, peer-reviewed)

DOI: 10.3390/jintelligence14070128

Data for both studies (Open Science Framework)

・Related: Chao et al.: Insight Is About Searching Far — Fukai Reads

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 84 of 89

Read next