PAPER-DIGEST · 2026-08-01
Jeong et al.: Same Puzzle, Different Answer Buttons, Different Difficulty — Fukai Reads
Cognitive load / interaction design / pediatric attention assessment
TL;DR
Change only the buttons a player answers with, and performance changes. This paper shows it on a tablet Stroop task (the classic attention test where a color word is printed in a conflicting ink color and you must report the ink, not the word). Two answer formats were compared: picking from written labels reading "red," "blue," "yellow" (text), versus picking from colored patches (color). The stimulus, font size, button size, layout and touch mechanics were identical; only the look of the response options differed.
127 children aged 6 to 12 did both. The color format produced higher accuracy (0.91 vs 0.86; Cohen d=0.34, P<.001) and faster responses (1183.21 ms vs 1268.56 ms; difference −85.4 ms, d=−0.69, P<.001). Only color-format performance correlated significantly with parent-reported attention problems (r=−0.20, P=.03; text was r=−0.05, P=.56). Prefrontal blood-flow measures, by contrast, showed no significant difference between conditions.
The authors conclude that response format affects both task demand and measurement validity. Restated for people who build puzzles: difficulty comes partly from content and partly from how input is constructed, and the second part is something design can remove.
Introduction
The paper is by Harim Jeong, Yong Jeon Cheong, Jihyeong Ro, Jihyun Cha, Jongkwan Choi and Minyoung Jung, published in JMIR Serious Games (2026, volume 14, e93468). It is a peer-reviewed original paper (DOI: 10.2196/93468, published 2026-05-26). A preprint of the same manuscript appeared on JMIR's preprint server on 2026-02-22, so anyone who wants to compare before and after review can follow that. The first and corresponding authors are at the Korea Brain Research Institute in Daegu; two co-authors are at OBELAB Inc, which makes the measuring device.
My reason for picking it today is simple. Anyone tuning puzzle difficulty eventually hits the question of whether the difficulty is in the problem or in the controls. This paper separates those two experimentally, and not crudely: it holds the stimulus fixed and swaps only the look of the response options, a design an implementer can copy.
It is written in a clinical, digital-health frame, so games are not discussed directly. But the skeleton — change the shape of the input and you change what you measure — carries straight over to timing a daily puzzle or driving adaptive difficulty. I read it less as a clinical paper and more as a measurement of how much a UI contaminates a score.
Background
The foundation is cognitive load theory, the framework for the mental effort needed to process information and carry out a task. It splits load in two: what the task itself intrinsically demands (intrinsic load) and the surplus created by how things are presented and operated (extraneous load). In education and e-learning, tidying up presentation to cut extraneous load has repeatedly been shown to improve performance.
In digital health, the authors note, this line of work has mostly targeted clinician workflows, leaving patient-facing assessment tools comparatively neglected. That matters because added load in an assessment produces errors and slower responses, and those enter the data as an effect of the screen rather than of the child's actual attention. The authors call it noise that masks the intended construct.
The Stroop task was chosen because it is a standard measure of selective attention and inhibitory control. Trials where word and ink conflict are slower and more error-prone than congruent ones, and that gap reflects the effort of resolving interference, making the paradigm sensitive to design-induced changes in demand. Additionally, in primary-school children reading is not yet fully automatic, so answering by reading words risks mixing reading fluency into a measure of inhibition — that is where this study's hypothesis starts.
Approach
Participants were typically developing children recruited from local communities in the Republic of Korea. 137 enrolled between January and September 2023; 7 were dropped for technical failures and 3 withdrew, leaving 127 (55 girls, 43.3%; 72 boys, 56.7%; mean age 9.15, SD 1.56, range 6-12). Everyone did both formats in a within-subject design, and a sensitivity analysis indicated power to detect effects of d=0.25 or larger.
The task ran on a tablet (Samsung Galaxy Tab Advanced2). In both conditions a stimulus such as the word "blue" printed in yellow ink appeared, and the child reported the ink while ignoring the word. In the text condition the answer was chosen from written labels; in the color condition, from color patches. The core word-color conflict was preserved in both, so intrinsic load was held constant and only extraneous load at the response stage was manipulated. Five practice trials preceded 45 main trials per condition, self-paced, with a 30-second rest between conditions.
Load was measured with fNIRS (functional near-infrared spectroscopy, a non-invasive method that tracks blood-flow changes near the cortical surface), recording the prefrontal cortex over 15 channels. Two indices were used: functional connectivity built from channel-to-channel correlations, expressed as the change from a 5-minute resting baseline (delta-FC), and global efficiency, an index of how economically information can be exchanged across the network. On the behavioral side: accuracy, reaction time, and a composite score integrating the two (equivalent to the Rate Correct Score, absorbing speed-accuracy trade-offs).
For validity, the authors used the Attention Problems subscale of BASC-2, a parent-report behavioral rating scale, and compared its correlation with task performance in each condition. They also controlled general cognitive ability with an intelligence test (K-WISC-V), regressing age and full-scale IQ out of performance and fitting random forests (an ensemble of decision trees that also reports which variables matter for prediction) to the residuals, to compare the contributions of neural and behavioral measures.
Findings
Behaviorally the difference was clear. Accuracy: color 0.91 (SD 0.09) vs text 0.86 (SD 0.16), mean difference 0.05 (95% CI 0.03-0.08), t126=3.81, Cohen d=0.34, P<.001. Reaction time: color 1183.21 ms (SD 146.78) vs text 1268.56 ms (SD 145.18), difference −85.4 ms (95% CI −107.2 to −63.5), t126=−7.72, d=−0.69, P<.001. Composite: color 0.036 (SD 0.007) vs text 0.031 (SD 0.008), t126=6.81, d=0.62, P<.001 (Table 1 of the paper). Effect sizes ranged from medium to large.
The neural indices, by contrast, did not differ. Delta-FC was 0.12 (SD 0.33) for color vs 0.07 (SD 0.29) for text, t126=1.65, d=0.15, P=.10; prefrontal global efficiency was 0.55 (SD 0.17) vs 0.53 (SD 0.17), t126=1.52, d=0.14, P=.13. Both were numerically higher in the color condition but neither reached significance. The authors attribute this to effect sizes of this magnitude being typical in developmental fNIRS work, to prefrontal connectivity still maturing at these ages, and to both conditions sharing the core conflict so that the difference was subtle.
On validity, color-condition performance correlated significantly with parent-reported attention problems (r=−0.20, 95% CI −0.37 to −0.03, P=.03), while the text condition showed essentially none (r=−0.05, 95% CI −0.23 to 0.12, P=.56). However, the test of the difference between the two correlations, via Fisher r-to-z, gave z=1.54, P=.12, and the authors themselves write that it "approached but did not reach conventional significance." That is worth holding onto as something they did not claim.
The random forest results are a further turn. With age and IQ entered directly, most of the contribution went to age (58.86% in color, 73.21% in text). But in the residualized models, with age and full-scale IQ regressed out of performance, the neural indices took over: global efficiency 39.79% (color) / 43.60% (text), delta-FC 35.75% / 33.02%, and attention problems about 24% in both. The abstract states that individual variation in prefrontal efficiency accounted for 40% to 44% of residual performance variance. The relative contribution of age was also smaller in the color format (58.86% vs 73.21%).
Where you can use it
First: when A/B testing difficulty, build two versions with identical board logic and only the input construction swapped, then compare accuracy and completion time. That is exactly this paper's procedure, and it moved 85 ms and five accuracy points. If you run a daily Sudoku-like or Slitherlink-like puzzle and players call it heavy, suspect the number-entry or line-drawing mechanics before the board logic. The paper's control discipline — hold font, button size, layout and touch mechanics fixed, vary only the shape of the options — transfers directly.
Second: if you rank players by time, design on the assumption that input construction feeds straight into the standings. 85 ms is small in absolute terms, but d=−0.69 is not small at all. If your leaderboard treats a few hundred milliseconds as meaningful, part of that gap may be measuring tap-ability rather than skill. My concrete suggestion is a simple operating rule: never let the input UI differ between a recorded run and a practice run.
Third: if you feed performance into skill estimation or adaptive difficulty, choose the lower-extraneous-load format for measurement. Here, only the color format correlated with the external yardstick (parent ratings). Translated to games: if you want to estimate real ability but feed in numbers polluted by control friction, adaptation points the wrong way. In systems that time hint delivery off performance, that pollution becomes a misread of the player.
Fourth: in puzzles for children or non-native speakers, do not make people read where they need not. The design idea here was to drop language processing out of the response stage entirely by replacing text labels with color patches. Icons, color coding and shapes can function as load reduction, not decoration. Fifth, a lesson for player models: control for age and experience first. In the raw models here, age took 60% to 70% of the contribution, and the same structure can appear in your own logs.
Limitations
The authors list four weaknesses. The condition order was fixed as color then text, so practice or fatigue effects cannot be ruled out (though they argue practice should have favored the later text condition, and color still won, so the format effect appears robust to order). The sample was limited to typically developing 6-to-12-year-olds, so generalization to other developmental stages and clinical groups is not assured. The neural feature set was coarse — only global efficiency and condition-level delta-FC — and event-locked or channel-specific analyses might be more sensitive. And the random forest models had modest predictive performance (cross-validated R-squared of 9% to 23%), so they should be read as exploratory estimates of relative importance, not established predictive models.
What follows is what I noticed on reading. What I, Fukai, would point out first is that reading fluency itself was never measured. The central argument is that reading is not yet automatic at these ages, so answering by reading words mixes in reading ability — yet no reading measure was collected, with age and the IQ test standing in as proxies. The tastiest part of the argument is not tested directly.
Second, the validity argument is inevitably weak. The contrast between a significant color correlation and a non-significant text one catches the eye, but the direct test of the difference between the two correlations was not significant (z=1.54, P=.12). One being significant and the other not does not mean the two differ significantly. The authors say so honestly, though the abstract's concluding sentence reads somewhat more strongly. I do not think one can write that the color format was shown to be more valid. The correlation itself is r=−0.20, with the confidence interval reaching up to −0.03.
Third, two researchers from OBELAB Inc, the maker of the NIRSIT Lite device used, are co-authors and supplied both the hardware and the analysis software. This is stated in the author-contributions section, so nothing is hidden, but it is worth remembering while reading the neural results. Fourth, more a caution than a limitation: it remains open whether the two conditions are the same task at different difficulty or two different tasks. Removing verbal competition at the response stage may change what is measured, not merely lower the load. The authors have a rebuttal ready, but they have not settled it.
Fukai's reading
This section alone is my own interpretation. I would place this study at the most practical edge of the movement carrying cognitive load theory out of the classroom and onto the screen. In the vocabulary of design criticism, what it measured is the size of the input tax. Of the effort a player spends, the part going to the puzzle's content and the part going to the construction of input are separate ledgers, and the second is a rate the designer sets. What makes this paper interesting is that lowering that rate by one notch — the smallest possible intervention — moved performance and also changed how well the score meshed with an external yardstick. For a puzzle author, not confusing those two ledgers is, I think, the core of the craft. Making input fiddly when you want to add difficulty is close to raising a tax and calling it a challenge.
Closing
On its own this is an observation under limited conditions: 127 typically developing children in Korea, one Stroop paradigm, a fixed order. The right reading is not "people are like this" but "under these conditions this was observed." Replication and preregistered analyses are listed by the authors themselves as future work (counterbalancing order, retest reliability, wider age range, clinical subgroups, time-resolved connectivity analysis, and mediation or structural equation testing), which makes this a starting-point paper. It is also newly published and not yet widely discussed.
For a wider map, read it alongside Seyderhelm and Blackmore's work on assessing game-task difficulty through real-time measures of performance and cognitive load (2023, Simulation & Gaming). That one takes games themselves as its subject and asks how to measure difficulty, which pairs neatly with this paper's entry point of how the measurement contaminates the score. As for cognitive load theory itself, grasping the distinction between intrinsic and extraneous load is enough to read this paper.
References
Papers and related material referenced in this article:
・Preprint version of the same paper (first posted 2026-02-22, before peer review)
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 46 of 47