TAG
#game-design
0 reviews · 64 essays
Related essays
Is a Win by Luck “My” Win? — The Philosophy of Alea and Desert
Late at night, Balatro handed me a personal-best score. Delighted, I sent it to a friend, who replied, “That was luck, right?” Can I call a win that chance dealt me my win? In this Nature of Play — Fragments piece, I sit two philosophers who would answer opposite ways face to face: Caillois, who insists chance is a first-class form of play, and Aristotle, who holds that a victory lives in a person’s activity, not in fortune. Slowly the two begin to set “a single throw” and “a long win-rate” on two separate plates.
Solving the Time Loop — Designing Games Where Knowledge Is the Only Save Data (from Majora's Mask to Outer Wilds)
A time loop is a machine for making knowledge the only save data. From Majora's Mask's three days to Outer Wilds' twenty-two minutes and The Sexy Brutale's twelve hours, a design reading of loop games as observation puzzles extended along the axis of time.
Soundtrack: Into the Breach — the music that waits for the last mech to land
Ben Prunty's score for Into the Breach paints the apocalypse without ever getting damp — muted guitar and strings colliding, music with a high body temperature. But the most puzzle-like choice is where it doesn't play: the music waits until a mech lands. Over black coffee, I, Doremi, take apart that design of silence.
Reading Homo Ludens Chapter by Chapter — The One Condition Under Which War Can Still Be Play
Try to build a competitive or PvP mode and the mood turns tense before anyone even wins. Nothing seems further from play than war — yet Huizinga says a fight that meets one condition is play. Reading log part 3: 'play and war'. The condition is recognizing your opponent as an equal — checked against the 'honor code' For Honor's players wrote for themselves.
Is Difficulty Unkind? The Philosophy of Fairness
A friend playtesting the hardest stage of my puzzle looked up and said, "Isn't this a bit unkind?" Is making a game hard an act of cruelty, or a gift? With the old Sekiro "easy mode" controversy in the corner of my eye — a game with no difficulty settings at all — this fragment of The Nature of Play seats two philosophers who would answer opposite face to face: Rawls, who argues for fairness from behind a veil of ignorance, and Nietzsche, who affirms the wall with "what does not kill me makes me stronger." Bit by bit, the two stop kicking at "difficulty" and start kicking, side by side, at "unfairness."
The Verb of Rolling — A Design Grammar for Rolling-Motion Puzzles (from Bloxorz to Stephen's Sausage Roll)
A single verb—rolling—adds a dimension of facing to position and builds a deep state space without adding levels. From Kula World to Stephen's Sausage Roll, a design reading of rolling-motion puzzles.
Halina & Guzdial: Generating Levels as a "Cake of Time" — Fukai Reads
A procedural level-generation paper by Halina and Guzdial. It represents a level as a "cake" of board states stacked over time, and generates a level and its solution together with PRP, which recombines play traces. In Sokoban, against six existing methods, it reached 100% playability with high diversity, without hand-authored constraints or rewards.
Soundtrack: The Turing Test — one motif that keeps asking, human or machine
The music Sam Houghton and Yakobo wrote for The Turing Test is dark, minimal, and never congratulates you. Nearly every track carries the same motif, changing its colour room by room. Black coffee in hand, I — Doremi — take apart the design of 'one motive, endlessly distorted' and the nerve of music that refuses to react to your solving, in a form you can carry home to your own writing.
Inside Jakub Dvorský's Philosophy — Dropping Words, Building a World You Can Touch
A study of Amanita Design founder Jakub Dvorský, read across primary interviews in Adventure Classic Gaming (2009), MCV/DEVELOP (2011), bounthavy (2020) and TouchArcade (2024). It traces the origin of his wordless design, his obsession with atmosphere and hand-drawn worlds, his aim to build 'an interactive toy' you keep playing rather than a puzzle you defeat, his dilemmas of kindness versus difficulty and commerce versus authorship, and the Czech animation and science-fiction he himself names as influences — grounded only in what he has said in public.
Hsu et al.: LLM-Voiced NPCs Make Players' Heads Heavier -- A 'Double-Edged Sword' Experiment — Fukai Reads
An empirical LLM-NPC paper by Hsu et al. (Communication University of China and others). They built a scripted-NPC version and a GPT-4.1 LLM-NPC version of the same game and ran a between-subjects test with 130 players. LLM-NPCs significantly raised cognitive load (p<.001), did not significantly improve overall enjoyment (p=.195), and increased autonomy while lowering usability and trust.
Inside Frank Lantz's Philosophy — giving weight to the abstract
A study of Frank Lantz, founding chair of the NYU Game Center and maker of Universal Paperclips and Drop7, drawn from his own essay The Truth in Game Design (2010), a long Thought Economics interview (2024), PC Games Insider (2017), and his own notes on his work. Grounded only in his public statements, it reads his philosophy of games as an art of systems and as psychology experiments we run on ourselves, his obsession with embodying abstract ideas, the failure of Bite Me answered by "solving it all too well" in Leviathan, the dilemma of comfort versus truth, and the influences he acknowledges himself — Bostrom, Stapledon, Wittgenstein and von Neumann.
Wang et al.: Gauging Tetris Block Puzzle Difficulty by How Fast a Strong AI Learns — Fukai Reads
An arXiv preprint from a National Yang Ming Chiao Tung University and Academia Sinica group that measures the difficulty of the popular mobile game Tetris Block Puzzle. It rates rule variants by how fast and high a strong AI (Stochastic Gumbel AlphaZero) can learn to play, finding that more holding/preview blocks make the game easier while adding block shapes makes it harder (the T-pentomino most of all).
The Grammar of Solving Together — Co-op Puzzle Design and Information Asymmetry
Puzzle design theory quietly assumes a single player. I set Portal 2's division of verbs against the information asymmetry of Keep Talking and Nobody Explodes and We Were Here, add PICO PARK's shared learning curve, and trace the grammar of co-op puzzles where conversation itself becomes the verb.
Johnson et al.: What Changes in a Game When You Build an LLM Into It — Fukai Reads
A qualitative study by Johnson and colleagues at the University of Calgary on developing two games with an LLM embedded in their structure. Reading developer reflections, it analyzes how embedding an LLM as a component (not decoration) changes gameplay, playability, and player experience. Variability and personalization increase, but new burdens of correctness, difficulty calibration, and coherence emerge, with schema enforcement and validation as the keys.
If You Can Always Undo, Do Your Choices Still Mean Anything?
The night I added an undo button to my own puzzle, my hand froze — especially since what I'm making is a Baba Is You–style game where you rewrite the rules themselves. If you can undo, was the decision ever a decision? This third "fragment" of The Nature of Play sets Sartre and Nietzsche arguing face to face: Sartre (commitment, anguish), for whom a choice declares who you are, against Nietzsche (eternal recurrence, amor fati, style), who asks whether you could will it to return forever. In the end they point at the same single line under different names.
Inside Jason Rohrer's Philosophy — playing life and death through constraint
A study of Jason Rohrer, who with the five-minute Passage told a whole life and death through rules, drawn from three interviews: Critical Inquiry (2011), Handmade Pixels (2017) and Gamasutra (2011). Grounded only in his public statements, it reads his philosophy of speaking through interaction rather than narrative, his obsession with turning constraint into an expressive tool, the failure of trying to eliminate tedium and his change of heart toward "punishment," the dilemma of highbrow versus accessible, and the influences he acknowledges himself — Rod Humble, Raph Koster and Scott McCloud.
Handcrafted or Generated — A Design Theory of Procedural Puzzle Levels
Who authors a puzzle's boards? I set the handcrafted lineage of Nikoli and Tametsi's 160 levels against the generated lineage of Simon Tatham's collection and Hexcells Infinite, ask what procedural generation drops, and look at the daily puzzle as a third way between generation and curation.
Reading Homo Ludens Chapter by Chapter — When Play Becomes Contest, the Stake Is Honor
Add ranks or scores to a puzzle and the mood can suddenly turn tense. Does competition make play better, or break it? Reading log part 2 is Huizinga's agon (contest): what people really compete for is not money or goods but the honor of being first.
Ye et al.: Measuring Image-Capable AI (MLLMs) with Children’s Intelligence Tests — Fukai Reads
A paper (arXiv preprint) by Hengwei Ye and colleagues at ShanghaiTech University on KidGym, an MLLM evaluation benchmark inspired by children’s intelligence tests (the Wechsler scales). It measures five abilities — Execution, Perception Reasoning, Memory, Learning, Planning — across 12 tasks on a 2D grid at three difficulty levels, evaluating nine models. Even top models reached only 0.30 on abstract-shape puzzles and 0.72 on counting against a human 1.00.
Zeng et al.: Automating Game Balancing with LLM-vs-LLM Self-Play — Fukai Reads
A paper by Zeng et al. on automated game balancing. It tackles balancing asymmetric strategy games by using multi-agent LLM self-play as an evaluator and Bayesian optimization to search rule parameters, reporting convergence to near-0% win-rate gaps on their own game, CivMini.
Designing Luck Out of the Game — From Minesweeper to Guessing-Free Logic Puzzles
Why did Minesweeper's 50/50 endgame guess survive so long, and how was it overcome? I trace the lineage through Hexcells, Tametsi, and 14 Minesweeper Variants, and ask what it costs to guarantee a board that never forces a guess.
Waugh: Measuring AI's Reasoning with Sudoku and Slitherlink — Fukai Reads
A paper (arXiv preprint) by Justin Waugh of Approximate Labs on Pencil Puzzle Bench, a benchmark that measures LLM reasoning with pencil puzzles. From 62,231 puzzles across 94 types it selects 300, and its core is that a machine can verify every move against the rules; 51 models were evaluated. Even the strongest GPT-5.2 reached only 56.0% in agentic mode, with about half unsolved.
If You Looked Up the Answer, Did You Really Solve It?
The guilt of opening a walkthrough after being stuck. Ryle says your skill hasn't grown one bit; Plato says a true opinion becomes knowledge once you tie down the reason. Two philosophers, opposed, and a rethink of hint design.
Ahn et al.: Puzzle Difficulty Lives in Concepts, Not Looks — Fukai Reads
A paper (arXiv preprint) by Ahn et al. at Boston University introducing CogARC, a human-adapted version of the ARC abstract-reasoning benchmark. Logging 260 people's grid-puzzle solutions edit by edit, they find difficulty is driven by conceptual rule complexity rather than grid size or color count, and that people converge on the same wrong answers even when they fail.
Luo et al.: How AI Delivers Help Matters as Much as the Help Itself — Fukai Reads
A paper by Luo et al. (UC Santa Barbara) on how a mixed-initiative AI delivers help. Using Rush Hour puzzles, they compare on-demand (Button) help with inactivity-triggered (Timer) help, and show that although task performance is nearly identical, the Timer mode earns more positive perceptions of the AI. Accepted to IUI '26.
Reading the Unreadable — The Grammar of Decipherment Puzzles, from Fez to Chants of Sennaar
Fez, Tunic, Heaven's Vault, Chants of Sennaar. A reading of the decipherment-puzzle lineage — languages cracked by observation alone — through one design question: how do you treat misreading?
Does Play Have a Point? — Camus and the Roguelike Death Loop
I've died 38 times in a roguelike and I'm still diving back in. I bring home almost nothing I built up, so why is this repetition fun? In this matching installment I lay Camus's The Myth of Sisyphus — his head-on argument about endless repetition — over the death loop of Hades.
Triebel et al.: Does AI Have Both a Head and a Hand on a Classic Physics Puzzle? — Fukai Reads
A paper by Triebel et al. evaluating VLMs on the classic physics puzzle The Incredible Machine 2. Using VLATIM, a five-stage benchmark, it asks whether screen-operating AI can solve problems like humans; the cleverer large models can plan but cannot click precisely, and no model solved even one puzzle to completion.
Reading 'Homo Ludens' Chapter by Chapter — The Order and Tension Play Creates
Last time's 'magic circle' was only the entrance. From here I read the classic on play, 'Homo Ludens,' one chapter at a time, as a maker. Of the traits Huizinga lists in Chapter 1, two pillars I hadn't touched yet - order and tension - tested against Tetris and my own puzzles.
What Is Play? — Starting with Huizinga's Magic Circle
Making puzzles keeps bringing me back to the most basic question: what is play? What is fun? In this new series, I go ask the philosophers. Part 1 is Huizinga's magic circle — why people get dead serious inside a simple drawn line.
Xu et al.: When Generative AI Becomes the Heart of Play — Fukai Reads the AI-Native Games Survey
A survey (arXiv preprint) by Zhiyue Xu and five co-authors on "AI-native games," where generative AI is the core loop itself. It defines them by a counterfactual — would play collapse if the AI were removed — and classifies 53 real artifacts along two axes: game type (G) and dominant AI mechanic (N), showing a skew toward narrative genres and a thin use of AI at the rule layer.
When the Control Scheme Decides the Difficulty — From Grid Movement to Drag-and-Arrange
Sokoban's grid movement, The Witness's line, the dragging of Gorogoa and A Little to the Left, Return of the Obra Dinn's cursor, The Gardens Between's time, and Golf Peaks's cards. A maker's-eye survey of how a single input shapes a puzzle's difficulty, framed by discreteness, undo cost, and affordance.
Aryan et al.: When You Stall, the World Changes — AbideGym Turns Static RL Worlds into Adaptivity Tests — Fukai Reads
A preprint by Aryan et al. (Abide AI) on RL environment design. To fight the brittleness that comes from training in fully static worlds, AbideGym rewrites the rules and grows the map mid-episode, triggered by the agent's own inactivity, forcing it to abandon memorized policies and re-plan. The paper presents the design and a comparison to prior work; no experimental results yet.
Wang et al.: An LLM Agent That Reads Mental Busyness From Gaze — Fukai Reads
A paper from Meta Reality Labs and collaborators that estimates cognitive load (mental busyness) from eye gaze. It tackles the poor generalization and low interpretability of prior methods with GazeMind, a framework that structures gaze and has an LLM reason over it with context, individual traits, and worked examples, reporting 62.73% accuracy on three-way classification (over 20 points above prior methods).
Mirowski et al.: From Writing a Story to Finding One — Fabula, a Writing AI Grown With the Writers' Community — Fukai Reads
A paper on Fabula, a Google DeepMind writing-support AI. Its hierarchical story planner-generator, the Drama Manager, was critically co-developed with 42 experts; it proved strong at structure but weak at style and surprise. Fukai reads it for lessons that apply directly to game interactive narrative.
Walking and Deducing — The Boundary from Gone Home to Return of the Obra Dinn
The 'walk and read' experience Gone Home and Firewatch refined, versus the 'actively deduce' experience of Return of the Obra Dinn and The Case of the Golden Idol. Where is the line between them? A designer's reading of the fault line between walking simulators and deduction puzzles, with Her Story and Outer Wilds in between.
Bazzaz et al.: Believing It's AI Changes the Experience — Fukai Reads
A CHI '26 paper by Bazzaz and Cooper on perception bias toward generated content. Mixing human-made and AI-generated levels in Super Mario Bros. and Sokoban for 142 players, they report that players can barely identify the creator, yet levels believed to be AI-made are rated less fun, harder, and more frustrating.
Liu et al.: AI Assistance Erodes Persistence — A Warning for Hint Design — Fukai Reads
A paper by Grace Liu and colleagues on how AI assistance affects independent problem-solving and persistence. Across RCTs with 1,222 participants, AI raised in-session performance but, once removed, left people solving less and giving up more. Those who got direct answers declined most while hint-users did not, a result that speaks directly to game hint design.
Recursion as a Puzzle Grammar — The Nested Logic of Patrick's Parabox and Recursed
A box inside a box, and inside it the same board again. From Recursed to Patrick's Parabox and Cocoon, a reading of nested-puzzle design and why recursion runs so deep.
Jara Gonzalez & Guzdial: Generating Enemy Shapes as Gates You Need a Mechanic to Beat — Fukai Reads
A paper by Jara Gonzalez and Guzdial on generating enemy morphologies (collision shapes). They frame 'enemies defeatable only with a specific mechanic' as a 4x4 grid generation problem, compare reinforcement learning, A* search, and neural generation, and find a simple A* reachability rule yields the best gating and most diverse shapes at the lowest cost.
Legible Failure — Making the Dead End Readable in Puzzles
In puzzles, failure is not death but the dead end. From Sokoban's irreversible push to Stephen's Sausage Roll's invisible stalls and the soft-lock-free design of The Witness and COCOON, I examine the design question of whether failure can be read, not whether it should be punished.
Zeytuncu: Puzzle Difficulty Comes Down to How Many Numbers You Use — Fukai Reads
A difficulty-modeling paper by Yunus E. Zeytuncu on integer arithmetic puzzles (Countdown-style number games). Using an exact solver to generate over 3.4 million instances and defining difficulty by minimum operation count, it shows that the number of inputs used in a minimal solution alone is a 'minimal sufficient statistic' that perfectly determines difficulty.
Chao et al.: Insight Is About Searching Far — Fukai Reads
A paper on insightful problem-solving by Chao, Hsieh & Wu. Using a Japanese RAT and a simulation to quantify the search path to a solution, it shows that de-fixation is necessary for solving but is not what determines insight; the hallmark of insight is exploring the solution space over greater distances.
Monti et al.: Measuring AI's Planning Power on a Single-Corridor Sokoban — Fukai Reads
A paper by Monti and colleagues on SokoBench, a benchmark that measures reasoning models' long-horizon planning with Sokoban. By lining up only single-box straight corridors and narrowing difficulty to a single axis (corridor length), it shows that even state-of-the-art reasoning models break down once more than 25-30 moves of lookahead are needed. The authors locate the cause in accumulated miscounting.
Luo et al.: Can AI Agents Build Whole Playable Games in a Real Engine? — Fukai Reads
A paper by Luo, Wang and colleagues on GameCraft-Bench, a benchmark for end-to-end game generation by coding agents. It has agents build complete playable games on Godot from natural-language specs, judged by launch, input replay, and video-based scoring across 140 tasks in 15 families. Even the strongest configuration reaches only 41.46% overall, and the authors report that agents can build mechanics but fall short of finished games with content, readability, and polish.
Li et al.: AutoBG, an AI that supports board game design end-to-end from ideation to finish — Fukai Reads
A paper (arXiv preprint) by Zizhen Li et al. on AutoBG, a board game design assistant that covers the whole workflow—ideation, rulebook generation, and individualized feedback—via Verifier-Gated Iteration that splits the generator from the critic; the critic, BG-Critic, is reported to outperform GPT-5.4 on diagnostic quality.
Nasir et al.: Evolving the Rules of Play Themselves — Fukai Reads MORTAR
A paper on automatic game design by Nasir, Togelius and colleagues. Instead of levels, MORTAR evolves game mechanics themselves using a quality-diversity algorithm paired with a large language model, judging quality by whether stronger AI agents reliably beat weaker ones. Running on GPT-4o-mini, it generates diverse, playable games and even quantifies each mechanic's contribution.
"The hacking was always there" — Capcom's Pragmata and the design of simultaneous puzzle-shooter gameplay (Game Developer, April 2026)
One article today. Alessandro Fillari's April 14, 2026 interview on Game Developer explores how Capcom designed Pragmata — a third-person shooter where players must simultaneously solve Snake-style hacking puzzles during combat. Neither shooting nor hacking alone can finish a battle. Producers Edvin Edsö and Naoto Oyama explain how the dual-system design existed from day one, and how the team fought repetitiveness by making the hacking system evolve as players improve.
Closing Into One Screen — The Density a One-Screen Puzzle Builds
Sokoban, Baba Is You, Snakebird, Patrick's Parabox — the strongest thinking puzzles keep their whole board on one screen. A designer's look at why simultaneous visibility deepens thought, and when breaking the frame is worth its cost.
The Grammar of Zachtronics — Translating Programming into Puzzle
SpaceChem, TIS-100, Shenzhen I/O, Opus Magnum, Exapunks. A maker's-eye reading of the programming-puzzle grammar Zach Barth honed from 2011 to 2022, along three axes: the command queue as a verb, optimization design that abandons the single solution, and the ladder of abstraction.
Xu et al.: Promoting Game Mechanics to Coordinates to Generate Solvable Levels — Fukai Reads
A PCG (level generation) paper by Xu and Verbrugge of McGill University. Against geometry-first prior methods, it proposes HDPCG, which runs pathfinding on a dimensional-expanded graph that promotes mechanics such as gravity inversion and moving platforms to a coordinate, guaranteeing solvability during generation, and reproduces playable levels in Unity.
Sun et al.: Why Do Players Lose Themselves in Punishingly Hard Games? — Fukai Reads
A paper by Sun et al. on difficulty design in Soulslike games. Through a qualitative analysis of 600 Steam reviews it asks why players immerse themselves in punishingly hard games, and proposes 'resilient flow' — absorption sustained by meaningfully framing frustration.
Narrative Puzzles and Storyless Puzzles — Lorelei vs Stephen's Sausage Roll
The pure maneuvers of Stephen's Sausage Roll versus Lorelei and the Laser Eyes' solutions fused with story. What narrative and storyless puzzles each sell, contrasted from a designer's view across Obra Dinn, Golden Idol, COCOON, and Machinarium.
Inside Tetsuya Mizuguchi's Philosophy — Designing Senses, Not Genres
A study of Tetsuya Mizuguchi — creator of Rez, Lumines and Tetris Effect — across four of his own interviews. His consistent philosophy of designing sensation rather than genre and aiming to “make people cry,” his dilemma over what to add to a classic, and influences from Kandinsky to a single night at a music festival, traced only through source-checked statements.
Designing Hint Systems — How to Show, How to Hide
InvisiClues' invisible ink, the silence of The Witness, the friction of The Case of the Golden Idol, Obra Dinn's rule of three. A maker's-eye survey of hint systems as a declaration of how a game treats a stuck player.
Inside Arvi Teikari's Philosophy — Wanting to Surprise You, He Lets You Play the Rules Themselves
"The greatest motivator is just general desire for self-expression." Finnish solo developer Arvi Teikari (Hempuli) startled the world with Baba Is You, a game that lays its rules out as word-blocks on the board and lets players rewrite them. We read his philosophy, obsessions, failures, dilemmas and influences through his own words.
When Fewer Verbs Make a Richer Game — The Lineage of Subtractive Design
Sokoban, Snakebird, Stephen's Sausage Roll, A Monster's Expedition, Bonfire Peaks. A maker's-eye survey of subtractive design, the lineage that deepens difficulty without adding verbs, built around one question: why does less become more?
Inside Lucas Pope's Philosophy — If There's No Problem, I'm Not Interested
"If there's no problem, then I'm not that interested. But if there's some restriction or some limitation, then I'm interested suddenly." Lucas Pope calls himself, consistently, an engineer. He inverts mundane jobs, pares them down, and bets on the player's imagination. A study of his philosophy, obsessions, failures, dilemmas, and influences, read through his own interviews and talks.
Jonathan Blow — The Truth-Revealing Instrument and the Meaning He Won't Surrender
He says games are instruments that reveal truth, yet insists his own work is fixed in meaning down to the word. Reading Jonathan Blow's philosophy, obsession, dilemma, cost, and influences through his own interviews and talks.
Soundtrack: Outer Wilds — When music becomes a tool for solving
Andrew Prahlow's score for Outer Wilds is not background music to leave running. Most of the time it stays silent; when it sounds, it is the signal of discovery; and at times the instruments themselves become tools of exploration. Black coffee in hand, I — Doremi — take apart how this music works, through a lens you can take home to your own composing and design.
The Vocabulary of Perspective Puzzles — From Monument Valley to Manifold Garden
Echochrome, Monument Valley, Antichamber, Manifold Garden, and Viewfinder. What the perspective-puzzle genre has invented in fifteen years, and what design space still remains, read from a maker's point of view.
The Ethics of Undo — Forgiveness or Punishment
One button reshapes the entire experience. Sokoban's restart, Braid's rewind, Baba Is You's unlimited Undo. A look at the dividing line between designs that forgive trials and designs that punish them.
Carving the Learning Curve — Baba's Vertical Wall and How It Was Built
When and how should a puzzle game stop the player? A comparison of Baba Is You's notorious vertical wall, Cocoon's unbroken flow, and the design philosophy that lives between them.
Observation as Play — Common Grammar of Witness, Obra Dinn, and Lorelei
The Witness, Return of the Obra Dinn, and Lorelei and the Laser Eyes all turn the act of looking into a verb. A look at the grammar these three share.



































