PAPER-DIGEST · 2026-08-17

Lu et al.: Is Flow Made of Difficulty, or of the Effort You Spend? — Fukai Reads

flow experience / mental effort / difficulty design (psychology)

TL;DR

Match the difficulty to the player's skill and they enter flow — the state of absorbed, time-forgetting focus. It is one of the most frequently cited claims in game design. The paper I read today tries to cut a seam into that claim. Is flow driven by the difficulty of the task itself, or by the amount of effort the player actually invests? Until now the two have been treated as one bundle.

The authors used a visual discrimination task (two grid images shown side by side; the participant judges whether they are identical or differ by one shifted square). Grid density manipulated 'difficulty'; a prior announcement of what percentage of trials contained a difference manipulated the motivation to invest effort. EEG P300 was recorded alongside. The study appeared in the peer-reviewed journal Psychological Research (Springer, 2025).

The results do not fit a simple picture. The difficulty manipulation landed hard (effect size partial eta-squared = 0.64), while the expectancy manipulation had no effect at the individual level and only a weak one at the trial level (d = 0.04 in the difficult condition). Flow showed an inverted-U shape against difficulty, but only marginally (p = 0.053), and P300 showed no relation to flow at all (b = -0.003, p = 0.97). What to take away is not a conclusion but a map: the single word 'difficulty' splits into at least three separate dials.

Introduction

The paper is titled 'Disentangling the effects of task difficulty and effort on flow experience.' The authors are Hairong Lu, Dimitri Van der Linden and Arnold B. Bakker, of the Department of Industrial Psychology and People Management at the University of Johannesburg. It appeared in Psychological Research (Springer), volume 89, issue 4, article 113, published online 26 June 2025, DOI 10.1007/s00426-025-02128-x. This is not an arXiv preprint; it went through peer review.

I chose it because for anyone shipping a daily puzzle, where to place difficulty is the single largest design variable — and I would rather not take the psychology folklore behind that decision on faith. If I am going to lean on a received claim, I want to have read the study that went and tested the claim itself. This paper does exactly that: it inserts a second explanatory variable into the 'challenge-skill match' account at the centre of flow research.

One caveat up front. This is a single laboratory experiment with 37 participants, and the authors themselves frame it as exploratory. There is no mention of preregistration (publishing hypotheses and analysis plans before running the study, so that interpretations cannot be swapped in after seeing the results). Throughout this article I keep what was demonstrated separate from what was not.

Background

The received account in flow research runs like this. Absorption arises when the challenge of the task matches the skill of the player; too easy and you are bored, too hard and you are anxious. Flow therefore traces an inverted U against difficulty — high in the middle, low at both ends. This is the curve most game design textbooks reproduce.

But a separate literature draws almost the same curve: research on mental effort, the amount of cognitive force invested in a task. Under motivational intensity theory (the framework holding that effort keeps rising with difficulty as long as success still seems attainable, then collapses the moment the task is judged impossible), effort also traces an inverted U against difficulty. The two even share physiological markers such as pupil dilation.

That creates a problem. If flow and effort peak at the same point on the difficulty axis and take the same shape, then what we have been explaining as 'flow because challenge matched skill' might really be 'flow because that was the easiest place to pour effort into.' The authors' question is whether the two can be experimentally pulled apart. If you can move effort without moving difficulty, the separation becomes possible.

Approach

The design was within-subjects (every participant experiences every condition): 3 difficulty levels x 2 expectancy conditions, six cells in all. The task was visual discrimination — two grid images of white and light grey squares presented simultaneously, with the participant judging whether they were identical or differed by one shifted square. Difficulty was manipulated through grid density: 2x2 (easy), 8x8 (intermediate), 14x14 (difficult). The finer the grid, the harder a one-square shift is to spot.

Effort, by contrast, had to be moved without changing the task itself. The authors used EPDD (expected probability of detecting differences — the probability announced in advance for how often a difference would appear). The control condition was told 50% of trials contained a difference; the high-expectancy condition was told 80%. The manipulation follows the expected-value-of-control idea that higher perceived odds of success make effort seem worth spending.

Four things were measured. Flow: a single-item 10-point scale ('to what extent do you think you were in a flow state?'), used after a preliminary interview confirming participants understood the concept. Perceived difficulty: a 7-point scale from -3 to +3 asking whether the task matched their skill. Effort: reaction time and accuracy. And physiology: the EEG P300 (a peak in the brainwave appearing roughly 300 milliseconds after a stimulus, associated with the allocation of attentional resources), taken at the Pz electrode in a 300-400 ms window.

Participants were 37 undergraduates (29 female, mean age 19.87, SD 1.78) — the count after excluding one for EEG data loss and one for misunderstanding flow. Six five-minute blocks were run in randomised order. Analysis used mixed-effects models (estimating the overall pattern while absorbing between-participant differences as a random factor), with correction applied for the post-hoc comparisons.

Findings

First, the manipulations did separate as intended. Perceived difficulty moved strongly with grid density: F(180,2) = 162.80, p < 0.001, partial eta-squared = 0.64. The easy condition felt easier than the intermediate one by d = 1.24 and than the difficult one by d = 2.30. The expectancy manipulation, meanwhile, did not move perceived difficulty at all (F(180,1) = 0.85, p = 0.36). Difficulty and expectancy behaved as separate dials. Worth noting: the intermediate x high-expectancy cell landed almost exactly on the challenge-skill match point (mean -0.11, not significantly different from zero).

Effort. At the individual level, reaction time moved dramatically with difficulty (F(180,2) = 1202.45, p < 0.001, partial eta-squared = 0.93), as did accuracy (F(180,2) = 533.53, p < 0.001, partial eta-squared = 0.86), but expectancy produced no main effect (reaction time: F(180,1) = 1.69, p = 0.195). Zoom in to the trial level, however — 8,785 trials — and the expectancy effect appears: F(1,8743.1) = 11.93, p < 0.001. The breakdown matters: no difference in the easy condition (p = 0.436, d = 0.01), with high expectancy raising reaction time only in the intermediate (p = 0.005, d = 0.05) and difficult (p = 0.003, d = 0.04) conditions. The effect sizes are very small.

Flow itself averaged 6.14 (SD 1.977) on the 10-point scale. Under control expectancy an inverted-U against difficulty was visible but marginal (quadratic coefficient b = -0.66, p = 0.053). Under high expectancy the inverted U disappeared (b = -0.35, p = 0.292). The main effect of difficulty fell short of significance (F(2,180) = 2.63, p = 0.074), and comparing high expectancy against control within the difficult condition gave t(36) = 1.69, p = 0.099, d = 0.28 — not significant. The authors write that flow appeared to trace the movement of effort, while stating explicitly that the interaction was not significant.

The physiological measure was a clear miss. P300 showed no inverted-U against difficulty under either expectancy condition (p = 0.706 and p = 0.140) and no relationship with flow (b = -0.003, p = 0.97). The authors note that similar null findings have appeared in prior work and treat the result as one that challenges assumptions about the neural correlates of flow.

Where This Is Useful

One: how to think about difficulty indicators. A line like 'today's solve rate: 68%' moves the perceived odds of success without touching the board at all. What this study suggests is that such expectancy manipulations barely registered at the block level and showed up only weakly trial by trial. If so, a single pre-game banner is likely a weaker lever than feedback that updates mid-solve — squares becoming certain, candidates shrinking, moves remaining made visible.

Two: hint design. If I were building a Sokoban-like, I would place the strategy-confirming hint ('this crate moves last') ahead of the labour-replacing hint ('put this crate here'). The latter removes the effort; the former preserves the effort while raising the sense that success is attainable. In this paper's framing, only the former corresponds to raising EPDD. Given the small effect sizes, treat it as a direction worth testing rather than a settled rule.

Three: the granularity of playtest measurement. The most practical lesson here may be that the same manipulation was invisible at the individual level and visible at the trial level. If I were evaluating a change to a daily puzzle, I would make fine-grained behavioural logs the primary measure — time per move, number of backtracks, where players stop — and demote the single 'did you enjoy it' item to a supporting role.

Four: where to place the difficulty curve. The cell closest to a genuine challenge-skill match in this experiment was intermediate difficulty combined with high expectancy (mean perceived difficulty -0.11). For a hypercasual PCG (Procedural Content Generation — automatic creation of game content) system tuning difficulty on the fly, the two-stage arrangement of aiming near the middle of the solve-rate distribution and then layering a 'this is solvable' signal on top was, under these conditions at least, the least strained placement.

Limitations

The authors acknowledge plenty. Implementing the expectancy manipulation at block level may have diluted its trial-by-trial impact, and the effect sizes were small. The control value of 50% was already relatively high, so a ceiling may have hidden differences that a lower-expectancy condition would reveal. The sample of 37 was set by practical considerations rather than a power analysis, as they state. Pupil diameter proved unusable because the task involved long processing times and frequent saccades, leaving P300 as the only physiological measure. And they describe the findings overall as initial, if novel, insights.

What I would point out here is, first, the row of p-values. The headline results cluster at 0.053, 0.074 and 0.099. One such value can be shrugged off; a row of them is more accurately read as 'there may be an effect, but this sample could not establish it.' It is certainly not material for declaring that flow is determined by effort rather than difficulty. Even the expectancy effect that did reach significance at the trial level carries effect sizes of d = 0.04 to 0.05 — hard to call perceptible in practice.

Second, the nature of the task. This is abstract visual discrimination: no narrative, no goal, no audience. Flow in a puzzle game rests substantially on what solving something means, and a task that strips that away yields a measure of flow I would be cautious about importing directly into daily puzzle design. The participants were also 37 undergraduates from a single university with a mean age of 19.87 — not a population that overlaps with a general solver base.

On balance, the value of this paper lies in showing that an experiment separating flow from effort can be built, not in establishing what the separation yields. With no preregistration mentioned and the authors themselves calling it exploratory, treating the findings as a hypothesis until replication arrives is the appropriate stance. Given the replication problems in psychology, that is not excessive caution.

Fukai Reads

From here on this is my own reading. I would place this study in the broader movement of taking apart a design variable that has been compressed into the single word 'difficulty.' The hardness of the task, the effort the player actually invests, and the sense that the thing is solvable — in practice these three have long been treated as one dial. What the paper offers, I think, is not a conclusion but a map showing that there are at least three dials. In the vocabulary of design criticism, difficulty tuning is less about making the task easier and closer to engineering the conditions under which a player judges the effort worth spending. That reading, I should say, is not something the paper itself demonstrates.

Closing

For readers who want to go deeper, I recommend pairing this with Strojny and colleagues' 2023 paper in PLOS ONE. It tested motivational intensity theory in an actual game (Icy Tower), with 39 participants playing across four difficulty levels. Its conclusion: involvement rises with difficulty while the task remains feasible, then drops sharply once it becomes unachievable. It views from the side of a real game the same framework today's paper handled with an abstract laboratory task. Read together, the two sketch a map of how far the relationship between difficulty, effort and involvement is understood, and where the understanding stops.

Flow research itself spans decades, and today's paper only drives a test into one corner of it. Which is why reading this one study and concluding that the nature of flow has been settled would be a mistake. The accurate reading, I think, is that what we have been bundling under the word 'flow' is beginning to be taken apart by experimental hands. What those of us building puzzles can do is check the intermediate results of that dismantling against our own boards.

Sources

Papers and materials referenced in this article:

・Disentangling the effects of task difficulty and effort on flow experience (Hairong Lu, Dimitri Van der Linden, Arnold B. Bakker, 2025, Psychological Research 89(4), Article 113 / peer-reviewed)

・DOI: 10.1007/s00426-025-02128-x

・Open-access version of the same paper (PubMed Central, PMC12202627)

・Related work: Player involvement as a result of difficulty: An introductory study to test the suitability of the motivational intensity approach to video game research (Paweł Strojny, Agnieszka Strojny, Krzysztof Rębilas, 2023, PLOS ONE)

Reactions (no login)

Anonymous • one of each per visitor per day

Learn — Curriculum

LearnPart 4 Difficulty — Designing the Learning Curve and FailureChapter 12 Measuring Difficulty6 / 10

Part of these series

Paper DigestEpisode 60 of 98

Read next