TAG
#difficulty
0 reviews · 28 essays
Related essays
Wang et al.: Making a Puzzle Solver the Teacher for Every Single Move — Fukai Reads
A game-AI paper by Yu Wang and colleagues. Where long-horizon puzzles reward only the final win, they convert the drop in a solver's remaining-distance-to-goal into a per-move score and mix it into training. Averaged over Sokoban, Minesweeper and Rush Hour, success rises from 16.6% to 62.1%, and on unseen difficulty from 5.9% to 28.4%. Querying the solver costs about 73 parts per million of training wall clock.
Ponnock & Ho: The Order of Mario 1-1 Has a Measurable Teaching Effect — Fukai Reads
A reinforcement learning and level design paper by Jesse Ponnock and Lucas Ho (arXiv preprint, not peer-reviewed). Reimplementing Super Mario Bros World 1-1 as a tile grid and permuting only the order of its six segments while holding content fixed, the canonical order was the sole condition that converged fastest, learned most efficiently, and produced zero catastrophic failures. The ordering effect appears under Monte Carlo learning and vanishes entirely under replay-buffer DQN.
Gould & Ward et al.: Measuring Puzzle Difficulty in Units of Human Solve Time — Fukai Reads
An AI evaluation paper by Gould, Ward and colleagues. They attached human solve times to 43 benchmarks and over 30,000 problems, and found that the human time of tasks a model completes at 50% success without externalising its reasoning has doubled roughly every 373 days over six years, reaching about three minutes for GPT-5.5. Their difficulty-measurement craft, built partly on Sudoku and crosswords, transfers directly to puzzle design.
Teo et al.: AI Assistants Overassist — Int-Bench Measures Intervention in Problem-Solving — Fukai Reads
Teo et al. on LLM intervention behavior. Using Int-Bench, a simulated setting, they measure when and how much AI assistants help during problem-solving, finding AI intervenes earlier and more often than humans, tends to leak the answer, and does not improve transfer. A useful read for puzzle hint design.
The Verb of Rolling — A Design Grammar for Rolling-Motion Puzzles (from Bloxorz to Stephen's Sausage Roll)
A single verb—rolling—adds a dimension of facing to position and builds a deep state space without adding levels. From Kula World to Stephen's Sausage Roll, a design reading of rolling-motion puzzles.
Shyne et al.: How Far Do Puzzle Solver Loops Match Human Felt Difficulty — Fukai Reads
A logic-grid-puzzle difficulty study by Shyne, Facey & Cooper. Using solver loops (the pass count of a human-style solver) as a difficulty proxy, they generate difficulty-varied puzzles with a quality-diversity algorithm and, in a 63-player study, show solver loops correlate significantly with subjective difficulty (c=0.30, p=0.015).
Ye et al.: Measuring Image-Capable AI (MLLMs) with Children’s Intelligence Tests — Fukai Reads
A paper (arXiv preprint) by Hengwei Ye and colleagues at ShanghaiTech University on KidGym, an MLLM evaluation benchmark inspired by children’s intelligence tests (the Wechsler scales). It measures five abilities — Execution, Perception Reasoning, Memory, Learning, Planning — across 12 tasks on a 2D grid at three difficulty levels, evaluating nine models. Even top models reached only 0.30 on abstract-shape puzzles and 0.72 on counting against a human 1.00.
Designing Luck Out of the Game — From Minesweeper to Guessing-Free Logic Puzzles
Why did Minesweeper's 50/50 endgame guess survive so long, and how was it overcome? I trace the lineage through Hexcells, Tametsi, and 14 Minesweeper Variants, and ask what it costs to guarantee a board that never forces a guess.
Ahn et al.: Puzzle Difficulty Lives in Concepts, Not Looks — Fukai Reads
A paper (arXiv preprint) by Ahn et al. at Boston University introducing CogARC, a human-adapted version of the ARC abstract-reasoning benchmark. Logging 260 people's grid-puzzle solutions edit by edit, they find difficulty is driven by conceptual rule complexity rather than grid size or color count, and that people converge on the same wrong answers even when they fail.
Inside Bennett Foddy's Philosophy — Designing Failure to Fight Meaninglessness
A study of Bennett Foddy (QWOP, Getting Over It) drawn strictly from his own interviews and blog: a philosophy that starts from the paradox that "games don't matter," the obsession of cataloguing frustration by "flavor," his admission that he couldn't beat his own game, the dilemma between encouragement and taunt, and influences from Sexy Hiking to the unfair old games of his childhood. It closes with one paragraph reading him as a designer who fights the absence of meaning.
When the Control Scheme Decides the Difficulty — From Grid Movement to Drag-and-Arrange
Sokoban's grid movement, The Witness's line, the dragging of Gorogoa and A Little to the Left, Return of the Obra Dinn's cursor, The Gardens Between's time, and Golf Peaks's cards. A maker's-eye survey of how a single input shapes a puzzle's difficulty, framed by discreteness, undo cost, and affordance.
Wang et al.: An LLM Agent That Reads Mental Busyness From Gaze — Fukai Reads
A paper from Meta Reality Labs and collaborators that estimates cognitive load (mental busyness) from eye gaze. It tackles the poor generalization and low interpretability of prior methods with GazeMind, a framework that structures gaze and has an LLM reason over it with context, individual traits, and worked examples, reporting 62.73% accuracy on three-way classification (over 20 points above prior methods).
Özkan: Co-Training the Level-Generating AI and the Level-Solving AI — Fukai Reads
A paper by Miraç Buğra Özkan that trains level generation and level solving together via reinforcement learning. In Unity, a hummingbird (solver) and a floating island (generator) learn while watching each other's results, reaching about 90.2% success across 100 unseen layouts.
The strongest player is not the best tester: a paradox from a framework for measuring game difficulty with LLMs
One piece today. I read, in the original English, "LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents" by Chang Xiao (Adobe Research) and Brenda Z. Yang (Columbia University) (arXiv:2410.02829). It asks whether off-the-shelf LLMs can be used to measure game difficulty by letting them play a game and treating their performance as a difficulty proxy, tested on Wordle (a word puzzle) and Slay the Spire (a deck-building roguelike). The central finding is a paradox: LLMs play worse than the average human, yet the relative difficulty of challenges they struggle with correlates strongly with human data. Moreover, a near-optimal, information-theoretic Wordle solver that beats humans on move count showed almost no correlation with human-perceived difficulty. In other words, the entity that solves best is not the best difficulty tester. A thought-provoking read for anyone thinking about how to validate a difficulty curve.
Legible Failure — Making the Dead End Readable in Puzzles
In puzzles, failure is not death but the dead end. From Sokoban's irreversible push to Stephen's Sausage Roll's invisible stalls and the soft-lock-free design of The Witness and COCOON, I examine the design question of whether failure can be read, not whether it should be punished.
Zeytuncu: Puzzle Difficulty Comes Down to How Many Numbers You Use — Fukai Reads
A difficulty-modeling paper by Yunus E. Zeytuncu on integer arithmetic puzzles (Countdown-style number games). Using an exact solver to generate over 3.4 million instances and defining difficulty by minimum operation count, it shows that the number of inputs used in a minimal solution alone is a 'minimal sufficient statistic' that perfectly determines difficulty.
"We Wanted Something More" — How Capcom's Pragmata Designs a Puzzle-and-Shooter Coexistence (Game Developer)
One article today. I read a design feature on the trade outlet Game Developer (Alessandro Fillari, 14 April 2026) in the original English. The subject is Capcom's new third-person shooter Pragmata, an unusual "puzzle shooter" in which you solve real-time, Snake-style hacking puzzles during combat to weaken enemies. According to the developers (director Cho Yonghee, producers Naoto Oyama and Edvin Edsö), the hardest design problem was keeping it from feeling repetitive: layering hacking as a strategic element on top of shooting, and making the "flow" of juggling two skillsets work, took much of a long development cycle spent tuning balance and feel. A look at a notable game's offbeat hook from the design side.
Chao et al.: Insight Is About Searching Far — Fukai Reads
A paper on insightful problem-solving by Chao, Hsieh & Wu. Using a Japanese RAT and a simulation to quantify the search path to a solution, it shows that de-fixation is necessary for solving but is not what determines insight; the hallmark of insight is exploring the solution space over greater distances.
Puzzles Made to Show Off a System, Not to Stump You — Patrick Traynor on System-Centric Design in Patrick's Parabox (GDC 2024)
One article today: the official slides from Patrick Traynor's GDC 2024 talk, "System-Centric Puzzle Design in Patrick's Parabox." His premise is inverted: "the purpose of the system is not to make cool puzzles. The purpose of the puzzles is to showcase this cool system." So difficulty is tuned to communicate, not to challenge — puzzles simplified as much as possible while still conveying their idea. He covers smoothing the learning curve (insert, modify, delete, reorder, optionalize), ~15 full-game playtests recorded with narration, an idea-finding method of "find an interaction and force it," and a heuristic for a good puzzle system: how many puzzles you can make in it. 364 shipped puzzles, 600+ unused drafts. A 2024 talk, but worth reading now for how it reframes a core design assumption.
Closing Into One Screen — The Density a One-Screen Puzzle Builds
Sokoban, Baba Is You, Snakebird, Patrick's Parabox — the strongest thinking puzzles keep their whole board on one screen. A designer's look at why simultaneous visibility deepens thought, and when breaking the frame is worth its cost.
Sun et al.: Why Do Players Lose Themselves in Punishingly Hard Games? — Fukai Reads
A paper by Sun et al. on difficulty design in Soulslike games. Through a qualitative analysis of 600 Steam reviews it asks why players immerse themselves in punishingly hard games, and proposes 'resilient flow' — absorption sustained by meaningfully framing frustration.
Designing Hint Systems — How to Show, How to Hide
InvisiClues' invisible ink, the silence of The Witness, the friction of The Case of the Golden Idol, Obra Dinn's rule of three. A maker's-eye survey of hint systems as a declaration of how a game treats a stuck player.
How to make 'just-right' difficulty — letting a machine fit it to the player (a Canadian study) vs. a human authoring it through meaning (a US developer)
A version rebuilt with credible sources only. Two pieces today, both answering 'how do you deliver just-right difficulty?' from opposite directions. The first is a research paper by Canadian researchers Matthew McConnell and Richard Zhao (September 2025, arXiv): a system that generates puzzles in real time with a genetic algorithm and auto-tunes difficulty per player, validated in a user study. Its key finding: using 'time-on-task' alone as the adaptivity metric fails. The second is an interview with game designer Michael Hicks (Game Developer): churning out hard, time-consuming puzzles is easy; the truly hard part is finding interesting ideas to explore. A machine fitting difficulty to the player, and a human authoring difficulty through meaning. Both sources are peer-reviewed research and professional media - the kind makers can cite with confidence.
Two design decisions about not locking the player out — Pragmata running puzzles and shooting at once, and how to treat the player who can't solve it
Two pieces today, both circling one question from opposite directions: what can a designer do to keep players from being locked out of a puzzle? First, a Game Developer interview (April 14, 2026) in which Capcom's developers explain how Pragmata, a rare 'puzzle shooter' that stacks a real-time Snake-style hacking puzzle on top of third-person combat, was designed so as not to feel repetitive. Second, game designer Cheryl-Jean Leo's 2017 essay 'Are You Creating Impossible Puzzles?', which starts from the premise that no matter how carefully you design, you will eventually make a puzzle that is impossible for someone, and argues for giving away answers inside the game. A live development floor and a nine-year-old critique - placed side by side, the core of difficulty design comes faintly into view.
When Fewer Verbs Make a Richer Game — The Lineage of Subtractive Design
Sokoban, Snakebird, Stephen's Sausage Roll, A Monster's Expedition, Bonfire Peaks. A maker's-eye survey of subtractive design, the lineage that deepens difficulty without adding verbs, built around one question: why does less become more?
Counterpoint on Baba Is You — Reading Through the Negative Reviews
Komugi gave Baba Is You a 9.5/10. I sort five recurring negative-review patterns from Steam, Metacritic, and the critical press — the difficulty wall, the single-solution charge, variable explosion, frustration overtaking fun, and uneven difficulty between adjacent levels — and decide where I agree and where I push back.
The Ethics of Undo — Forgiveness or Punishment
One button reshapes the entire experience. Sokoban's restart, Braid's rewind, Baba Is You's unlimited Undo. A look at the dividing line between designs that forgive trials and designs that punish them.
Carving the Learning Curve — Baba's Vertical Wall and How It Was Built
When and how should a puzzle game stop the player? A comparison of Baba Is You's notorious vertical wall, Cocoon's unbroken flow, and the design philosophy that lives between them.










