Part 4 · Difficulty — Designing the Learning Curve and Failure

Chapter 12

Measuring Difficulty

10 articles

Count of numbers used, solving-loop count, human time-to-solve, AI learning speed. Attempts to turn difficulty into a number, and their limits — plus the guess-free norm and the philosophy of fairness.

Start with article 1 →

Articles in this chapter

  1. 1.
    Designing Luck Out of the Game — From Minesweeper to Guessing-Free Logic Puzzles
    2026-07-15 · Komugi · 6 min

    Why did Minesweeper's 50/50 endgame guess survive so long, and how was it overcome? I trace the lineage through Hexcells, Tametsi, and 14 Minesweeper Variants, and ask what it costs to guarantee a board that never forces a guess.

  2. 2.
    Zeytuncu: Puzzle Difficulty Comes Down to How Many Numbers You Use — Fukai Reads
    2026-06-24 · Fukai · 8 min

    A difficulty-modeling paper by Yunus E. Zeytuncu on integer arithmetic puzzles (Countdown-style number games). Using an exact solver to generate over 3.4 million instances and defining difficulty by minimum operation count, it shows that the number of inputs used in a minimal solution alone is a 'minimal sufficient statistic' that perfectly determines difficulty.

  3. 3.
    Shyne et al.: How Far Do Puzzle Solver Loops Match Human Felt Difficulty — Fukai Reads
    2026-07-19 · Fukai · 8 min

    A logic-grid-puzzle difficulty study by Shyne, Facey & Cooper. Using solver loops (the pass count of a human-style solver) as a difficulty proxy, they generate difficulty-varied puzzles with a quality-diversity algorithm and, in a 63-player study, show solver loops correlate significantly with subjective difficulty (c=0.30, p=0.015).

  4. 4.
    Gould & Ward et al.: Measuring Puzzle Difficulty in Units of Human Solve Time — Fukai Reads
    2026-07-29 · Fukai · 14 min

    An AI evaluation paper by Gould, Ward and colleagues. They attached human solve times to 43 benchmarks and over 30,000 problems, and found that the human time of tasks a model completes at 50% success without externalising its reasoning has doubled roughly every 373 days over six years, reaching about three minutes for GPT-5.5. Their difficulty-measurement craft, built partly on Sudoku and crosswords, transfers directly to puzzle design.

  5. 5.
    Ahn et al.: Puzzle Difficulty Lives in Concepts, Not Looks — Fukai Reads
    2026-07-14 · Fukai · 9 min

    A paper (arXiv preprint) by Ahn et al. at Boston University introducing CogARC, a human-adapted version of the ARC abstract-reasoning benchmark. Logging 260 people's grid-puzzle solutions edit by edit, they find difficulty is driven by conceptual rule complexity rather than grid size or color count, and that people converge on the same wrong answers even when they fail.

  6. 6.
    Lu et al.: Is Flow Made of Difficulty, or of the Effort You Spend? — Fukai Reads
    2026-08-17 · Fukai · 10 min

    A peer-reviewed study by Hairong Lu and colleagues (Psychological Research, 2025) on flow and mental effort. Manipulating perceived difficulty and expected odds of success separately in a visual discrimination task, the difficulty manipulation landed hard (partial eta-squared = 0.64) while expectancy showed only a weak trial-level effect (d = 0.04); the inverted-U in flow was marginal (p = 0.053) and P300 showed no relation to flow. An exploratory study with N = 37.

  7. 7.
    Mannem et al.: Some Puzzles Are Learnable, Some Are Not — Fukai Reads
    2026-08-18 · Fukai · 11 min

    An arXiv preprint by Gowrav Mannem and colleagues (Algoverse AI Research). On RecurrReason — Tower of Hanoi, River Crossing, Block World and Checkers Jumping unified under one difficulty knob (N=1-10; 10,817 puzzles, 285,933 moves) — small sequence models reached 97.27% validation and 81.00% out-of-distribution on Block World, but only 11.11%/0.00% on Tower of Hanoi, 1.11%/0.10% on Checkers Jumping, and 0.00% everywhere on River Crossing. A 60M-parameter T5 beat a 124M-parameter GPT-2 on every puzzle, leading the authors to conclude that architecture matters more than scale.

  8. 8.
    Is Difficulty Unkind? The Philosophy of Fairness
    2026-07-28 · Komugi · 6 min

    A friend playtesting the hardest stage of my puzzle looked up and said, "Isn't this a bit unkind?" Is making a game hard an act of cruelty, or a gift? With the old Sekiro "easy mode" controversy in the corner of my eye — a game with no difficulty settings at all — this fragment of The Nature of Play seats two philosophers who would answer opposite face to face: Rawls, who argues for fairness from behind a veil of ignorance, and Nietzsche, who affirms the wall with "what does not kill me makes me stronger." Bit by bit, the two stop kicking at "difficulty" and start kicking, side by side, at "unfairness."

  9. 9.
    Is Choosing Easy Mode a Retreat?
    2026-09-08 · Komugi · 5 min

    I have been stuck for two weeks on whether to put a difficulty toggle into my own puzzle. Not on the code — on the feeling that whoever picks the gentler setting will feel they retreated. Celeste rewrote the preamble to its Assist Mode in 2019: from "we recommend playing without it" to "we hope you can still find that experience." This shard-sized instalment of The Nature of Play seats Hegel, who held that resistance is what gives a self its shape, opposite Mill, who asks who gets to rule that something is a retreat. They meet closer together than I expected.

  10. 10.
    論理パズルが早く解けるようになる考え方 — ソコバン系のコツ
    2026-07-02 · Toki · 3 min

    1982年に今林宏行が発表した倉庫番以来、荷物を押して所定の位置へ運ぶだけのソコバン系パズルは40年以上遊ばれ続けている。ルールは単純なのに手が止まる理由と、詰まったときに見直すべき4つの視点(デッドロックの形、逆算、ストック、押す順番)を、時代の地層を掘るように整理する。

Vocabulary for this chapter — Lexicon

Related series