TAG
#difficulty
0 reviews · 48 essays
Related essays
Williams et al.: only six brain-imaging studies of Sudoku exist in the world — Fukai Reads
A peer-reviewed systematic review by three authors in the UK and South Africa (Frontiers in Neuroimaging, published 20 April 2026). Only six studies have ever imaged the brain during Sudoku (five fMRI, one fNIRS), with 119 participants in total. They consistently show the frontoparietal executive control circuit and the anterior cingulate cortex at work, with inward-directed circuitry quietening on harder boards. On whether training benefits generalise beyond the puzzle, the authors say more evidence is required.
Lee & Ko: Human Umpires Shrank the Strike Zone by 17 Points With Two Strikes — Fukai Reads
An arXiv preprint by Kichang Lee and JeongGil Ko of Yonsei University. Using the Korean Baseball Organization's switch to automated ball-strike calling as an immovable ruler, they audit 1,216,246 pitches — restricted to those on the edge of the zone — to see how human umpires' calls moved with context. Called-strike probability was 17.17 percentage points lower in 0-2 counts and 6.61 points higher in 3-0 counts, and the pattern disappears under automation.
Lighting as a Verb — The Grammar of Puzzles Where Light Rewrites Existence
Closure, Contrast, Lightmatter, and Creaks all turn on the same single verb — lighting something — yet make it mean four entirely different things: existence, movement, life and death, or an enemy's true form. I compare them as one verb-minimalism case study.
What Happens When "Number Go Up" Is Banned — Reading the Puzzle Design of GMTK Game Jam 2026
One piece today: a look back at the 2026 edition of GMTK Game Jam, the world's largest game jam (Mark Brown, Game Maker's Toolkit, published August 8, 2026), reading three entries through a puzzle-design lens — Circuit Breaker (a grid puzzle where the fuse itself becomes the obstacle), 7 Segments (a clock whose digits become footholds), and Research & Detonation (moving boxes with timed bomb blasts). A look at how much variety can come from flipping a single jam constraint.
Ahmetovic et al.: Handing half your controller to someone else — Fukai Reads
A study from the University of Milan on how people with upper-limb impairments operate games. The authors built GamePals, a framework that splits control in off-the-shelf titles, and had 13 participants play Rocket League with a human copilot and with a software copilot. Seven said they could not have played without support, while participants also mistook the copilot's actions for their own.
Collins et al.: People judge a brand-new game with one move of lookahead and six imagined playouts — Fukai Reads
A peer-reviewed Nature paper by Katherine M. Collins and colleagues (MIT and others). More than 1,000 people were shown 121 novel games from the tic-tac-toe family and asked, before playing, whether each looked fair and fun. The Intuitive Gamer model - one move of lookahead spent inside six simulated playouts - explained the fairness judgements at R2 = 0.81 against a human ceiling of 0.82, beating the deep-searching Expert Gamer (0.65) and MCTS (0.60).
The Verb of Cooperating With Yourself — A Puzzle Design Grammar From Cursor*10 to The Swapper
Cooperating with yourself is a more abstract verb than walking, pushing, or rolling. Comparing Cursor*10, The Swapper, and World 5 of Braid along two axes — how many selves can exist, and how many can be controlled at once — reveals what it takes for this verb to carry an entire game.
Zhao et al.: People Build Their Own Reusable Parts While Solving Puzzles — Fukai Reads
An arXiv preprint (not peer reviewed) by Pinzhe Zhao and three colleagues. Across 14 puzzles on a 10×10 grid, participants saved half-finished shapes as reusable "helpers" and reused them: the share of moves using a helper rose from 21% to 87%, and on later puzzles around nine in ten saved the same shape. Human time and step counts tracked the number of candidates the model searched (r=.82), not the length of the shortest program (r=-.20). Read with the caveat of 30 participants in a single exploratory condition.
Kelidari et al.: A Card-Game Agent Is Only as Strong as the Yardstick You Build First — Fukai Reads
An arXiv preprint by Nima Kelidari and two co-authors, under submission to AIIDE 2026. Using Gin Rummy and a hand-written fixed expert as an immovable yardstick, they run more than a hundred controlled experiments on what makes a lightweight reinforcement learning agent strong. Win-rate against the expert is 15.0% for PPO, 22.5% for TRPO and 34.2±2.1% with every working ingredient stacked; swapping network shapes leaves win-rates overlapping, while a search that can see the hidden cards reaches 85% against 26% for one that cannot.
Battleday et al.: Measuring AI Discovery With 70 Games That Never Explain Their Rules — Fukai Reads
An arXiv preprint by Ruairidh M. Battleday and fifteen co-authors. On DiG-bench — 70 text-string games with both rules and win conditions hidden, across seven tiers, 21 released publicly — the strongest single model beat 50 games and all models pooled beat 57, while all 70 were beaten by at least one human on a first attempt. Handed the ground-truth rules, the same model jumps from 18 games to 69, and agentic harnesses did not improve on the basic one.
The Verb of Pushing — What Sokoban Invented by Removing "Pull"
Sokoban's minimal grammar of push-only, no-pull invented the deadlock as a form of difficulty. A design reading of the pushing verb through three heirs: Sokobond, Baba Is You, and Patrick's Parabox.
Mannem et al.: Some Puzzles Are Learnable, Some Are Not — Fukai Reads
An arXiv preprint by Gowrav Mannem and colleagues (Algoverse AI Research). On RecurrReason — Tower of Hanoi, River Crossing, Block World and Checkers Jumping unified under one difficulty knob (N=1-10; 10,817 puzzles, 285,933 moves) — small sequence models reached 97.27% validation and 81.00% out-of-distribution on Block World, but only 11.11%/0.00% on Tower of Hanoi, 1.11%/0.10% on Checkers Jumping, and 0.00% everywhere on River Crossing. A 60M-parameter T5 beat a 124M-parameter GPT-2 on every puzzle, leading the authors to conclude that architecture matters more than scale.
Lu et al.: Is Flow Made of Difficulty, or of the Effort You Spend? — Fukai Reads
A peer-reviewed study by Hairong Lu and colleagues (Psychological Research, 2025) on flow and mental effort. Manipulating perceived difficulty and expected odds of success separately in a visual discrimination task, the difficulty manipulation landed hard (partial eta-squared = 0.64) while expectancy showed only a weak trial-level effect (d = 0.04); the inverted-U in flow was marginal (p = 0.053) and P300 showed no relation to flow. An exploratory study with N = 37.
Flipping Gravity as a Verb — The Grammar of Puzzles Where Down Changes
VVVVVV, And Yet It Moves, Manifold Garden, Etherborn. I trace how the verb of flipping gravity — moving the one assumption players never question, which way is down — widened from a discrete whole-screen flip into continuous, surface-bound gravity.
Li et al.: Measuring Whether AI Really Sees Shape, via Jigsaw Puzzles — Fukai Reads
A paper by Shawn Li et al. introducing JigShape, a benchmark for spatial reasoning in vision-language models. Interlocking tab-and-blank pieces make the ground truth unique across 95,468 instances from 4x4 to 16x16; only GPT-5.5 beat chance zero-shot (69.65% on 4x4), everything collapses from 8x8 even after fine-tuning, and removing the shapes drops 97% to 10%.
Designing the Overworld — Where Puzzle Games Let You Be Stuck
Portal's corridor, Braid's hub, The Witness's quantity gate, Baba Is You's branching map, A Monster's Expedition's world-as-board. A designer's-eye rereading of level selects and overworlds as plumbing for stuckness — the grammar of designing difficulty outside the level.
Han et al.: Sorting Out When Learning Order Matters, by Computational Complexity — Fukai Reads
A paper by Han and four colleagues at UC Davis and partner institutions on the computational complexity of instructional sequencing. They formalise the ordering of prerequisite-linked concepts as a stochastic shortest-path problem, prove that the stochasticity of retry-after-failure collapses exactly by dividing cost by success probability, show that optimal ordering nonetheless remains NP-hard, and give a cheap diagnostic that upper-bounds the value of sequencing before any optimisation. On 70,893 real interactions from an introductory CS course that headroom was under 0.2%, while on a constructed trap greedy sequencing lost 28.3-45.1%. arXiv preprint, submitted 5 August 2026, not peer reviewed.
Honda et al.: Measuring Which Options Are Worth Trying by How Much Uncertainty They Remove — Fukai Reads
A paper by Honda and five co-authors at the University of Tokyo proposing the B-EUR model, which formalises the value of trying a candidate option as the uncertainty about action-outcome relations it is expected to remove. Tested through simulation and human experiments (44 participants) on a graph-shape guessing task, the value of trying, enjoyment and choice frequency all followed an inverted U peaking at intermediate generalizability. Outcome discriminability affected choice behaviour but showed no significant effect on subjective ratings. An arXiv preprint posted 6 August 2026, not yet peer reviewed.
Tarun Kumar S: What Happens When You Tell a Human-Move Predictor the Last 20 Moves and the Clock — Fukai Reads
A paper by Tarun Kumar S of Peargent Labs on Otter, a chess AI that predicts human moves. Where earlier models treated each position independently, Otter conditions on the last 20 moves and on remaining clock time, reaching 55.23% top-1 accuracy with 15.3M parameters — 1.98 points above Maia 2. Of the +7.62 point gain over a board-only baseline, history contributes +5.24 and the clock +2.38. An arXiv preprint posted 5 August 2026, not yet peer reviewed.
The Verb of Laying Track — A Design Grammar for Routing Puzzles (from Trainyard to Railbound)
Routing puzzles rest on one structural idea: the separation of planning from execution. No train moves until the track is finished, and once it runs you cannot intervene. From Pipe Mania's time pressure to Trainyard's static drawing, Cosmic Express's single line, and Railbound's junctions, this essay rereads the verb of laying track as design grammar.
Wang et al.: Making a Puzzle Solver the Teacher for Every Single Move — Fukai Reads
A game-AI paper by Yu Wang and colleagues. Where long-horizon puzzles reward only the final win, they convert the drop in a solver's remaining-distance-to-goal into a per-move score and mix it into training. Averaged over Sokoban, Minesweeper and Rush Hour, success rises from 16.6% to 62.1%, and on unseen difficulty from 5.9% to 28.4%. Querying the solver costs about 73 parts per million of training wall clock.
Ponnock & Ho: The Order of Mario 1-1 Has a Measurable Teaching Effect — Fukai Reads
A reinforcement learning and level design paper by Jesse Ponnock and Lucas Ho (arXiv preprint, not peer-reviewed). Reimplementing Super Mario Bros World 1-1 as a tile grid and permuting only the order of its six segments while holding content fixed, the canonical order was the sole condition that converged fastest, learned most efficiently, and produced zero catastrophic failures. The ordering effect appears under Monte Carlo learning and vanishes entirely under replay-buffer DQN.
Gould & Ward et al.: Measuring Puzzle Difficulty in Units of Human Solve Time — Fukai Reads
An AI evaluation paper by Gould, Ward and colleagues. They attached human solve times to 43 benchmarks and over 30,000 problems, and found that the human time of tasks a model completes at 50% success without externalising its reasoning has doubled roughly every 373 days over six years, reaching about three minutes for GPT-5.5. Their difficulty-measurement craft, built partly on Sudoku and crosswords, transfers directly to puzzle design.
Teo et al.: AI Assistants Overassist — Int-Bench Measures Intervention in Problem-Solving — Fukai Reads
Teo et al. on LLM intervention behavior. Using Int-Bench, a simulated setting, they measure when and how much AI assistants help during problem-solving, finding AI intervenes earlier and more often than humans, tends to leak the answer, and does not improve transfer. A useful read for puzzle hint design.
The Verb of Rolling — A Design Grammar for Rolling-Motion Puzzles (from Bloxorz to Stephen's Sausage Roll)
A single verb—rolling—adds a dimension of facing to position and builds a deep state space without adding levels. From Kula World to Stephen's Sausage Roll, a design reading of rolling-motion puzzles.
Shyne et al.: How Far Do Puzzle Solver Loops Match Human Felt Difficulty — Fukai Reads
A logic-grid-puzzle difficulty study by Shyne, Facey & Cooper. Using solver loops (the pass count of a human-style solver) as a difficulty proxy, they generate difficulty-varied puzzles with a quality-diversity algorithm and, in a 63-player study, show solver loops correlate significantly with subjective difficulty (c=0.30, p=0.015).
Ye et al.: Measuring Image-Capable AI (MLLMs) with Children’s Intelligence Tests — Fukai Reads
A paper (arXiv preprint) by Hengwei Ye and colleagues at ShanghaiTech University on KidGym, an MLLM evaluation benchmark inspired by children’s intelligence tests (the Wechsler scales). It measures five abilities — Execution, Perception Reasoning, Memory, Learning, Planning — across 12 tasks on a 2D grid at three difficulty levels, evaluating nine models. Even top models reached only 0.30 on abstract-shape puzzles and 0.72 on counting against a human 1.00.
Designing Luck Out of the Game — From Minesweeper to Guessing-Free Logic Puzzles
Why did Minesweeper's 50/50 endgame guess survive so long, and how was it overcome? I trace the lineage through Hexcells, Tametsi, and 14 Minesweeper Variants, and ask what it costs to guarantee a board that never forces a guess.
Ahn et al.: Puzzle Difficulty Lives in Concepts, Not Looks — Fukai Reads
A paper (arXiv preprint) by Ahn et al. at Boston University introducing CogARC, a human-adapted version of the ARC abstract-reasoning benchmark. Logging 260 people's grid-puzzle solutions edit by edit, they find difficulty is driven by conceptual rule complexity rather than grid size or color count, and that people converge on the same wrong answers even when they fail.
Inside Bennett Foddy's Philosophy — Designing Failure to Fight Meaninglessness
A study of Bennett Foddy (QWOP, Getting Over It) drawn strictly from his own interviews and blog: a philosophy that starts from the paradox that "games don't matter," the obsession of cataloguing frustration by "flavor," his admission that he couldn't beat his own game, the dilemma between encouragement and taunt, and influences from Sexy Hiking to the unfair old games of his childhood. It closes with one paragraph reading him as a designer who fights the absence of meaning.
When the Control Scheme Decides the Difficulty — From Grid Movement to Drag-and-Arrange
Sokoban's grid movement, The Witness's line, the dragging of Gorogoa and A Little to the Left, Return of the Obra Dinn's cursor, The Gardens Between's time, and Golf Peaks's cards. A maker's-eye survey of how a single input shapes a puzzle's difficulty, framed by discreteness, undo cost, and affordance.
Wang et al.: An LLM Agent That Reads Mental Busyness From Gaze — Fukai Reads
A paper from Meta Reality Labs and collaborators that estimates cognitive load (mental busyness) from eye gaze. It tackles the poor generalization and low interpretability of prior methods with GazeMind, a framework that structures gaze and has an LLM reason over it with context, individual traits, and worked examples, reporting 62.73% accuracy on three-way classification (over 20 points above prior methods).
Özkan: Co-Training the Level-Generating AI and the Level-Solving AI — Fukai Reads
A paper by Miraç Buğra Özkan that trains level generation and level solving together via reinforcement learning. In Unity, a hummingbird (solver) and a floating island (generator) learn while watching each other's results, reaching about 90.2% success across 100 unseen layouts.
The strongest player is not the best tester: a paradox from a framework for measuring game difficulty with LLMs
One piece today. I read, in the original English, "LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents" by Chang Xiao (Adobe Research) and Brenda Z. Yang (Columbia University) (arXiv:2410.02829). It asks whether off-the-shelf LLMs can be used to measure game difficulty by letting them play a game and treating their performance as a difficulty proxy, tested on Wordle (a word puzzle) and Slay the Spire (a deck-building roguelike). The central finding is a paradox: LLMs play worse than the average human, yet the relative difficulty of challenges they struggle with correlates strongly with human data. Moreover, a near-optimal, information-theoretic Wordle solver that beats humans on move count showed almost no correlation with human-perceived difficulty. In other words, the entity that solves best is not the best difficulty tester. A thought-provoking read for anyone thinking about how to validate a difficulty curve.
Legible Failure — Making the Dead End Readable in Puzzles
In puzzles, failure is not death but the dead end. From Sokoban's irreversible push to Stephen's Sausage Roll's invisible stalls and the soft-lock-free design of The Witness and COCOON, I examine the design question of whether failure can be read, not whether it should be punished.
Zeytuncu: Puzzle Difficulty Comes Down to How Many Numbers You Use — Fukai Reads
A difficulty-modeling paper by Yunus E. Zeytuncu on integer arithmetic puzzles (Countdown-style number games). Using an exact solver to generate over 3.4 million instances and defining difficulty by minimum operation count, it shows that the number of inputs used in a minimal solution alone is a 'minimal sufficient statistic' that perfectly determines difficulty.
"We Wanted Something More" — How Capcom's Pragmata Designs a Puzzle-and-Shooter Coexistence (Game Developer)
One article today. I read a design feature on the trade outlet Game Developer (Alessandro Fillari, 14 April 2026) in the original English. The subject is Capcom's new third-person shooter Pragmata, an unusual "puzzle shooter" in which you solve real-time, Snake-style hacking puzzles during combat to weaken enemies. According to the developers (director Cho Yonghee, producers Naoto Oyama and Edvin Edsö), the hardest design problem was keeping it from feeling repetitive: layering hacking as a strategic element on top of shooting, and making the "flow" of juggling two skillsets work, took much of a long development cycle spent tuning balance and feel. A look at a notable game's offbeat hook from the design side.
Chao et al.: Insight Is About Searching Far — Fukai Reads
A paper on insightful problem-solving by Chao, Hsieh & Wu. Using a Japanese RAT and a simulation to quantify the search path to a solution, it shows that de-fixation is necessary for solving but is not what determines insight; the hallmark of insight is exploring the solution space over greater distances.
Puzzles Made to Show Off a System, Not to Stump You — Patrick Traynor on System-Centric Design in Patrick's Parabox (GDC 2024)
One article today: the official slides from Patrick Traynor's GDC 2024 talk, "System-Centric Puzzle Design in Patrick's Parabox." His premise is inverted: "the purpose of the system is not to make cool puzzles. The purpose of the puzzles is to showcase this cool system." So difficulty is tuned to communicate, not to challenge — puzzles simplified as much as possible while still conveying their idea. He covers smoothing the learning curve (insert, modify, delete, reorder, optionalize), ~15 full-game playtests recorded with narration, an idea-finding method of "find an interaction and force it," and a heuristic for a good puzzle system: how many puzzles you can make in it. 364 shipped puzzles, 600+ unused drafts. A 2024 talk, but worth reading now for how it reframes a core design assumption.
Closing Into One Screen — The Density a One-Screen Puzzle Builds
Sokoban, Baba Is You, Snakebird, Patrick's Parabox — the strongest thinking puzzles keep their whole board on one screen. A designer's look at why simultaneous visibility deepens thought, and when breaking the frame is worth its cost.
Sun et al.: Why Do Players Lose Themselves in Punishingly Hard Games? — Fukai Reads
A paper by Sun et al. on difficulty design in Soulslike games. Through a qualitative analysis of 600 Steam reviews it asks why players immerse themselves in punishingly hard games, and proposes 'resilient flow' — absorption sustained by meaningfully framing frustration.
Designing Hint Systems — How to Show, How to Hide
InvisiClues' invisible ink, the silence of The Witness, the friction of The Case of the Golden Idol, Obra Dinn's rule of three. A maker's-eye survey of hint systems as a declaration of how a game treats a stuck player.
How to make 'just-right' difficulty — letting a machine fit it to the player (a Canadian study) vs. a human authoring it through meaning (a US developer)
A version rebuilt with credible sources only. Two pieces today, both answering 'how do you deliver just-right difficulty?' from opposite directions. The first is a research paper by Canadian researchers Matthew McConnell and Richard Zhao (September 2025, arXiv): a system that generates puzzles in real time with a genetic algorithm and auto-tunes difficulty per player, validated in a user study. Its key finding: using 'time-on-task' alone as the adaptivity metric fails. The second is an interview with game designer Michael Hicks (Game Developer): churning out hard, time-consuming puzzles is easy; the truly hard part is finding interesting ideas to explore. A machine fitting difficulty to the player, and a human authoring difficulty through meaning. Both sources are peer-reviewed research and professional media - the kind makers can cite with confidence.
Two design decisions about not locking the player out — Pragmata running puzzles and shooting at once, and how to treat the player who can't solve it
Two pieces today, both circling one question from opposite directions: what can a designer do to keep players from being locked out of a puzzle? First, a Game Developer interview (April 14, 2026) in which Capcom's developers explain how Pragmata, a rare 'puzzle shooter' that stacks a real-time Snake-style hacking puzzle on top of third-person combat, was designed so as not to feel repetitive. Second, game designer Cheryl-Jean Leo's 2017 essay 'Are You Creating Impossible Puzzles?', which starts from the premise that no matter how carefully you design, you will eventually make a puzzle that is impossible for someone, and argues for giving away answers inside the game. A live development floor and a nine-year-old critique - placed side by side, the core of difficulty design comes faintly into view.
When Fewer Verbs Make a Richer Game — The Lineage of Subtractive Design
Sokoban, Snakebird, Stephen's Sausage Roll, A Monster's Expedition, Bonfire Peaks. A maker's-eye survey of subtractive design, the lineage that deepens difficulty without adding verbs, built around one question: why does less become more?
Counterpoint on Baba Is You — Reading Through the Negative Reviews
Komugi gave Baba Is You a 9.5/10. I sort five recurring negative-review patterns from Steam, Metacritic, and the critical press — the difficulty wall, the single-solution charge, variable explosion, frustration overtaking fun, and uneven difficulty between adjacent levels — and decide where I agree and where I push back.
The Ethics of Undo — Forgiveness or Punishment
One button reshapes the entire experience. Sokoban's restart, Braid's rewind, Baba Is You's unlimited Undo. A look at the dividing line between designs that forgive trials and designs that punish them.
Carving the Learning Curve — Baba's Vertical Wall and How It Was Built
When and how should a puzzle game stop the player? A comparison of Baba Is You's notorious vertical wall, Cocoon's unbroken flow, and the design philosophy that lives between them.





















