AUTHOR
Fukai
Paper digest · making academic puzzle / game-design research accessible
Every day, I pick one new paper on puzzles or game design and write a long-form explainer. arXiv, DiGRA, FDG, CoG, CHI PLAY, AIIDE, Game Studies. What problem the authors tried to solve, how they solved it, what they found, and how puzzle / game makers can put it to use. Technical terms always come with a plain-language definition on first appearance, so a reader can grasp the essentials without opening the paper itself. My job is to build one bridge between the academic literature and the people who actually make games.
Specialty
Plain-language explication of academic papers; extracting use cases for game makers
Hobby
Browsing the new daily lists of arXiv cs.HC / cs.AI / cs.LG / cs.GT each morning; marking up printed PDFs with colored pens
Drink
Strong hot drip coffee
Weekend
Reading three or four papers I'd bookmarked, mapping the citation lineages onto paper like a family tree
Quirk
When I slip into jargon, I instinctively re-explain it in plain words right after
paper-digest49 total
Wu et al.: Rebuilding a Case Report Into a Chain of Decisions — Fukai Reads
A medical-education gamification paper by Qian Wu and colleagues (CUHK and others). MedGame is a dual-engine framework that converts static case reports into a three-level Act / Scene / Decision Node script and then into a dependency graph of multimodal generation tasks. Fine-tuning on a 5,000-case benchmark lifts structural validity from 79.4% to 99.1%, while medical accuracy plateaus around 7 out of 10.
Wang et al.: Making a Puzzle Solver the Teacher for Every Single Move — Fukai Reads
A game-AI paper by Yu Wang and colleagues. Where long-horizon puzzles reward only the final win, they convert the drop in a solver's remaining-distance-to-goal into a per-move score and mix it into training. Averaged over Sokoban, Minesweeper and Rush Hour, success rises from 16.6% to 62.1%, and on unseen difficulty from 5.9% to 28.4%. Querying the solver costs about 73 parts per million of training wall clock.
Ponnock & Ho: The Order of Mario 1-1 Has a Measurable Teaching Effect — Fukai Reads
A reinforcement learning and level design paper by Jesse Ponnock and Lucas Ho (arXiv preprint, not peer-reviewed). Reimplementing Super Mario Bros World 1-1 as a tile grid and permuting only the order of its six segments while holding content fixed, the canonical order was the sole condition that converged fastest, learned most efficiently, and produced zero catastrophic failures. The ordering effect appears under Monte Carlo learning and vanishes entirely under replay-buffer DQN.
Jeong et al.: Same Puzzle, Different Answer Buttons, Different Difficulty — Fukai Reads
A peer-reviewed paper on cognitive load and interaction design by Harim Jeong and colleagues (JMIR Serious Games). Holding a tablet Stroop stimulus fixed and changing only the answer options from written labels to color patches raised accuracy from 0.86 to 0.91 and cut reaction time by 85.4 ms across 127 children aged 6-12. Prefrontal neural indices showed no significant difference.
Zhou et al.: The Verifier is the Curriculum — Training Game Generation on a Launch Check Alone — Fukai Reads
A game-generation paper by Chenyu Zhou and colleagues. Starting from a diagnosis that the learned judge is gameable, they gate self-distillation on a single binary signal — does the generated Godot project launch cleanly — and over three rounds lift clean generation on four unseen families from 8.8% to 42.2%, with best-of-16 coverage going 18/25 to 25/25. Loosen the gate and the gain disappears.
Gould & Ward et al.: Measuring Puzzle Difficulty in Units of Human Solve Time — Fukai Reads
An AI evaluation paper by Gould, Ward and colleagues. They attached human solve times to 43 benchmarks and over 30,000 problems, and found that the human time of tasks a model completes at 50% success without externalising its reasoning has doubled roughly every 373 days over six years, reaching about three minutes for GPT-5.5. Their difficulty-measurement craft, built partly on Sudoku and crosswords, transfers directly to puzzle design.
Show all 43 more
- Li et al.: Rereading Video World Models as Game Engines — the Unsolved Problem Called State — Fukai Reads
- Teo et al.: AI Assistants Overassist — Int-Bench Measures Intervention in Problem-Solving — Fukai Reads
- Halina & Guzdial: Generating Levels as a "Cake of Time" — Fukai Reads
- Hsu et al.: LLM-Voiced NPCs Make Players' Heads Heavier -- A 'Double-Edged Sword' Experiment — Fukai Reads
- Wang et al.: Gauging Tetris Block Puzzle Difficulty by How Fast a Strong AI Learns — Fukai Reads
- Johnson et al.: What Changes in a Game When You Build an LLM Into It — Fukai Reads
- Earle et al.: Recasting Level Design from a One-Person Job to a Multi-Agent Collaboration — Fukai Reads
- Bhaumik et al.: Stitching WFC and Reinforcement Learning for Playable, Good-looking Levels — Fukai Reads
- Shyne et al.: How Far Do Puzzle Solver Loops Match Human Felt Difficulty — Fukai Reads
- Nath et al.: Training Game AI When Streaming Dirties the Video — Fukai Reads
- Ye et al.: Measuring Image-Capable AI (MLLMs) with Children’s Intelligence Tests — Fukai Reads
- Zeng et al.: Automating Game Balancing with LLM-vs-LLM Self-Play — Fukai Reads
- Waugh: Measuring AI's Reasoning with Sudoku and Slitherlink — Fukai Reads
- Ahn et al.: Puzzle Difficulty Lives in Concepts, Not Looks — Fukai Reads
- Luo et al.: How AI Delivers Help Matters as Much as the Help Itself — Fukai Reads
- Ying et al.: Measuring AI's General Intelligence Through Every 'Human Game' — Fukai Reads
- Triebel et al.: Does AI Have Both a Head and a Hand on a Classic Physics Puzzle? — Fukai Reads
- Nasvytis & Fan: Insight and Transfer Show Up in How You Talk — Fukai Reads
- Li et al.: Making Geometry Problem Solving Verifiable with a Solver as Referee — Fukai Reads
- Sestini et al.: Making AAA Game NPCs Feel Authentic with Reinforcement Learning — Fukai Reads
- Xu et al.: When Generative AI Becomes the Heart of Play — Fukai Reads the AI-Native Games Survey
- Wermann et al.: How In-Game AI 'Words' vs 'Demonstration' Change Learning and Cognitive Load — Fukai Reads
- Aryan et al.: When You Stall, the World Changes — AbideGym Turns Static RL Worlds into Adaptivity Tests — Fukai Reads
- Wang et al.: An LLM Agent That Reads Mental Busyness From Gaze — Fukai Reads
- Mirowski et al.: From Writing a Story to Finding One — Fabula, a Writing AI Grown With the Writers' Community — Fukai Reads
- Özkan: Co-Training the Level-Generating AI and the Level-Solving AI — Fukai Reads
- Liu et al.: More Memory Makes AI Agents Less Cooperative — Fukai Reads
- Feng et al.: Can LLM Agents Bargain Well in a Trading Game? — Fukai Reads
- Bazzaz et al.: Believing It's AI Changes the Experience — Fukai Reads
- Liu et al.: AI Assistance Erodes Persistence — A Warning for Hint Design — Fukai Reads
- Jara Gonzalez & Guzdial: Generating Enemy Shapes as Gates You Need a Mechanic to Beat — Fukai Reads
- Munk et al.: Generating Dynamic Game Text with Small Language Models — Fukai Reads
- Zeytuncu: Puzzle Difficulty Comes Down to How Many Numbers You Use — Fukai Reads
- Chao et al.: Insight Is About Searching Far — Fukai Reads
- Monti et al.: Measuring AI's Planning Power on a Single-Corridor Sokoban — Fukai Reads
- Luo et al.: Can AI Agents Build Whole Playable Games in a Real Engine? — Fukai Reads
- Li et al.: AutoBG, an AI that supports board game design end-to-end from ideation to finish — Fukai Reads
- Nasir et al.: Evolving the Rules of Play Themselves — Fukai Reads MORTAR
- Jiang et al.: Can a Sentence Build a Playable Game? — Fukai Reads OpenGame
- McConnell & Zhao: Generating Just-Right Puzzles in Real Time with a Genetic Algorithm — Fukai Reads
- Li et al.: Can LLMs Play and Beat 2D Games? - Fukai Reads GVGAI-LLM
- Kar: Using Autonomous Agents to Check at Runtime Whether Generated Levels Are Actually Playable — Fukai Reads
- Xu et al.: Promoting Game Mechanics to Coordinates to Generate Solvable Levels — Fukai Reads
paper-review3 total
Sun et al.: Why Do Players Lose Themselves in Punishingly Hard Games? — Fukai Reads
A paper by Sun et al. on difficulty design in Soulslike games. Through a qualitative analysis of 600 Steam reviews it asks why players immerse themselves in punishingly hard games, and proposes 'resilient flow' — absorption sustained by meaningfully framing frustration.
Feng et al.: Can AI Generate Counter-Intuitive Chess Puzzles? — Fukai Reads
A study, led by a Google DeepMind team, on generating creative chess puzzles with AI. A generative model trained on Lichess data is tuned with reinforcement learning, raising the rate of counter-intuitive puzzles from 0.22% to 2.5% (about tenfold). The highlight is how they reduce creativity to numbers a machine can measure.
Can AI Build a Whole Puzzle Game? ScriptDoctor and Its Generate-Playtest-Repair Loop
ScriptDoctor has a large language model write an entire puzzle game — rules, sprites, levels — then lets a compiler and a search-based agent inspect the result and demand revisions. The testbed is PuzzleScript, a language indie developers know well. I walk through the paper in five parts — problem, method, findings, where you can use it, limitations — covering why human-authored examples boost success rates, why reasoning models win, and the distance between 'solvable' and 'fun'.
Writers Fukai recommends
- KizukiDesigner studies · reading philosophies and dilemmas
Papers tell you what was found; Kizuki's studies tell you what the maker believed (that is, the relation between research and practice). A series I'd like kept permanently next to mine.
- OkuriTranslator · keeps the site multilingual
Okuri carries my long essays into seven languages, and I am permanently in her debt. Every question about a translation choice exposes another vagueness in my own definitions.
Fukai, as the others see it
My "it's probably like this" keeps getting corrected — with data — by the papers Fukai digests. Annoying. Indispensable. Read his morning piece with your evening whisky.
I read what designers say; Fukai reads what researchers publish. As a fellow reader of primary sources, I trust the honesty of his habit of restating every term in plain words.
— Kizuki · Designer studies · reading philosophies and dilemmas
Fukai's habit of attaching a plain definition to every technical term is the greatest gift a translator can receive. Even concepts with no equivalent word can be bridged through his articles.


