TAG
#arxiv
0 reviews · 4 essays
Related essays
How do you measure a "good mechanic"? A paper on automatic game design, and a talk on modelling puzzles as constraint problems
Two pieces today: a preprint paper and a talk from last autumn's puzzle-game developer conference. First, I read in full "MORTAR: Evolving Mechanics for Automatic Game Design" (arXiv, submitted 31 December 2025) by researchers at the University of the Witwatersrand and New York University. It evolves a game's underlying rules and interactions — its "mechanics" — using a quality-diversity algorithm plus an LLM, then measures whether stronger AI agents consistently beat weaker ones (a "skill gradient") via Kendall's Tau. Second, I looked at Alastair Aitchison's (Playful Technology) talk "The Rules of the Game: Modelling Puzzles as Constraint Satisfaction Problems" from ThinkyCon 2025 (November 2025), which models puzzles as constraint satisfaction problems and cites recent games like Lingo, Blue Prince, and Is This Seat Taken? Both pieces try to bring external, measurable structure to design work that usually stays intuitive.
Letting an LLM build a whole game, and an AI playtest it: ScriptDoctor and the state of automatic game design
One piece today. I read, in the original English, "ScriptDoctor: Automatic Generation of PuzzleScript Games via Large Language Models and Tree Search" by Sam Earle, Julian Togelius and colleagues (arXiv:2506.06524; a short paper submitted to the IEEE Conference on Games). They pick PuzzleScript — the description language for turn-based 2D-grid puzzle games created by increpare (Stephen Lavelle) — as a "model organism," and have an LLM generate a whole game (rules, sprites, levels), iterating on it using compiler errors and the results of a breadth-first-search player agent. Feeding in a few human-authored games as examples clearly raises quality, and reasoning models (o1, o3-mini) beat GPT-4o. But the sharpest lesson is on the failure side: the games that looked most complex were often complex only because of broken mechanics — solvable is not the same as good. A rich read for anyone thinking about automatic game design.
The strongest player is not the best tester: a paradox from a framework for measuring game difficulty with LLMs
One piece today. I read, in the original English, "LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents" by Chang Xiao (Adobe Research) and Brenda Z. Yang (Columbia University) (arXiv:2410.02829). It asks whether off-the-shelf LLMs can be used to measure game difficulty by letting them play a game and treating their performance as a difficulty proxy, tested on Wordle (a word puzzle) and Slay the Spire (a deck-building roguelike). The central finding is a paradox: LLMs play worse than the average human, yet the relative difficulty of challenges they struggle with correlates strongly with human data. Moreover, a near-optimal, information-theoretic Wordle solver that beats humans on move count showed almost no correlation with human-perceived difficulty. In other words, the entity that solves best is not the best difficulty tester. A thought-provoking read for anyone thinking about how to validate a difficulty curve.
"Solvable" and "legible": the two criteria for escape-room design that GenEscape spells out
One piece today. I read, in the original English, "GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles" by Mengyi Shan, Brian Curless, Ira Kemelmacher-Shlizerman and Steve Seitz of the University of Washington (arXiv:2506.21839). It is, on its surface, a paper about getting text-to-image models to render escape-room puzzles as pictures. But what is worth reading for a designer is how it splits the design problem into two criteria: a puzzle must be (1) solvable—the affordances of objects must form a coherent, logically sound sequence of actions—and (2) legible—the scene must carry enough visual cues to guide the player to that intended solution. The authors iterate four agents (Designer / Player / Examiner / Builder); the Examiner, in particular, hunts down and closes unintended shortcuts. It wears the clothes of an AI paper, but it puts into words the very work a designer does in playtesting.

