PAPER-DIGEST · 2026-09-30

Surendran & Cooper: In fill-in-the-blank microgames, a detailed design doc could not beat a short prompt — Fukai Reads

Generative AI × PCG — Mad Libs-style microgame generation

TL;DR

Place, number, adjective. Fill in blanks like these, and a generative AI writes a tiny game on the spot. This study brings the childhood word game Mad Libs into game making.

The authors compared two kinds of prompts: a detailed, design-document-like "Full" template and a "Simple" one that only passes the genre and the words. Across four genres and 96 games, the two were nearly tied on whether games ran and whether they could be completed.

The short Simple prompts, however, did better at getting the chosen words into the game and at producing games that looked and played differently from one another. Still, only about half of all games could actually be finished, and the evaluation was a small, preliminary one done by the two authors themselves.

Who wrote it, and where

The authors are Akash Surendran and Seth Cooper, both at Northeastern University in the US. Cooper has long worked on procedural puzzle generation and game-based research.

The paper appeared at the PCG Workshop held with FDG '26 (Foundations of Digital Games, August 2026, Copenhagen). PCG stands for Procedural Content Generation, making game content automatically with programs. It carries an ACM DOI. As a workshop paper it is shorter than a main-track paper and reads more like a progress report, as I understand it.

I picked it today because its subject is the microgame. The WarioWare lineage of games that throw one few-second challenge after another is close to home for a puzzle site. Most research on generating games with AI aims for bigger and more polished. This paper deliberately faces the other way.Gameplay screen of Scribblenauts UnlimitedScribblenauts Unlimited (5th Cell Media, 2012). Words you write appear as objects. An elder of the idea that words can decide a game's content (Steam screenshot)

Why games that are allowed to be unfinished

Making games is still hard. Tools and languages that lower the barrier have been studied for a long time. Recently there has been more work on having large language models (LLMs, AI trained on huge amounts of text that can write prose and code) write game code.

The paper's related-work section covers that thread briefly: METAGAME and Ludi, which generated board games; evolutionary generation using VGDL, a language for describing games; and LLM-based Unity prototypes and WarioWare-style mechanic generation.

Generating a large game whole, though, makes gaps in coherence and rough edges stand out. Here the authors flip the idea. In a microgame that lasts a few seconds, roughness is easier to overlook, and might even be funny. The charm of Mad Libs likewise lies in words chosen without context landing awkwardly in place.

So the goal is not to extend what LLMs can do. The authors write that they aim to use LLMs as they are, to make unpolished, small, personalized games.

How the games were made and checked

The mechanism is simple. A prompt template contains blanks such as {{PLACE}} or {{VERB_1}}. The player never sees the template, only a list of categories like "place" and "verb". The filled-in words go into the blanks, and the finished prompt is handed to an LLM, which writes the game's code.

For example, the line "It should give the feeling of {{VERB_1}} in a {{PLACE}}" becomes "It should give the feeling of running in a house." The model was Anthropic's Claude Sonnet 4.5. Games were written in Python with Pygame, drawn entirely in code with no image assets.

Two templates were compared. Full is close to a design document: gameplay, objectives, concrete values such as health, and even how to avoid common bugs. Simple contains only the genre, the list of words, and an instruction to include the words in the game.

The four genres were tower defense, platformer (jumping across platforms), reaction timing, and jigsaw puzzle. Each author made eight word lists, and from each list three games were generated with Full and three with Simple. All 96 games were played and rated by the author who had not made them, with genre and template type hidden.

Evaluation was a ladder of four questions: does it run, is it playable, is it completable, and is it fun. The authors also counted how many words appeared in the title, in on-screen text, or in other forms, and counted how many of the three same-prompt games were clearly distinct in gameplay or visuals.

What they found

First, do they work? According to Table 1, Full games ran 100% of the time and Simple 98%. Playable: Full 98%, Simple 92%. Both are high so far.

The wall comes after that. Completable: Full 50%, Simple 48%. Judged fun: Full 31%, Simple 35%. The authors write that nearly half of all games made reaching a completed state impossible or unreasonably difficult, usually from a lack of tuning and balance.

Genres differed. In tower defense, all 12 Full games were completable against 6 Simple ones, yet 0 Full games were judged fun against 5 Simple. Jigsaw: 3 Full and 6 Simple completable. Reaction: 4 Full and 8 Simple (out of 12 per genre per template).

Simple did better at bringing the words in. In the overall figures, unique keywords used were 41% for Full versus 64% for Simple; appearing as on-screen text, 28% versus 57%. The share of games clearly distinct from their siblings was 60% for Full and 75% for Simple.

In jigsaw, though, the picture flips. Full games drew the objects the words named, while Simple games tried to cram in as many words as possible, so the puzzle image was always a panel with the words written on it. A case where Simple wins on the numbers but not as play.

The authors' summary is cautious: results from both approaches were broadly comparable, neither is clearly superior, and simple prompts may work well enough for microgames.

Where the games broke

What I enjoyed most was the catalogue of how games broke. It works as a good checklist for makers who want to avoid the same traps in their own work.

Numbers first. The screen shows "5", implying five coins, but only two exist. The authors note that numbers enter blanks cut off from their context, so keeping them within sensible ranges is hard. In jigsaw, some puzzles had several blank white pieces, leaving the player to guess blindly where they go.

Platformers produced platforms too high to jump onto. Reaction games spawned targets faster than the authors could react. Each is a textbook case of playable but not completable.

Words showed up in recognizable patterns: shape pictures that assemble a word from geometric shapes and colors (glasses, tomatoes, stars); plain colored rectangles (blue water, yellow lard, green mucus); and games that seemed to capture a word's emotion, such as anger shown as fast-moving red objects. The authors write that seeing the word as on-screen text reinforced the sense that it had been incorporated.

How puzzle makers can use this

First: if you make a bonus mode for a daily puzzle, this fill-in-the-blank approach is ready to try. Collect three "words of the day" from players and ship a one-off microgame built on them. What matters is less polish than the feeling that "my words got in". As the paper observed, simply showing the word as on-screen text may help.

Second: never ship the output as is; always add a check. In the paper, nearly every game ran, but only half could be finished. For a Sokoban-like puzzle, confirm with a solver (a program that checks automatically whether the level can be solved) before release. The studies we covered on letting AI playtest and fix games and using a verifier as a curriculum work toward filling exactly this gap.

Third: do not hand number blanks to players. Values that decide success, such as coin counts and time limits, should be fixed within a range by the maker. Let players fill only words that set look and mood, such as nouns and adjectives. The paper's number failures show why that line is needed.Gameplay screen of Baba Is YouBaba Is You (Hempuli Oy, 2019). Rearranging word blocks changes the rules. A case of words acting on mechanics rather than looks, useful for thinking beyond fill-in-the-blank generation (Steam screenshot)

Fourth: vary prompt detail by genre. In tower defense, the detailed Full template was easier to complete, but Simple was more fun. In jigsaw, Simple turned the image into a board of words. This reads as saying that what needs constraining differs by genre, a clue for deciding what to fix and what to leave to generation in your own puzzle.

Fifth: make three from the same prompt and let players pick. In the paper, 60% to 75% of games were clearly distinct from their same-prompt siblings. Lining up a few and letting players choose might turn drawing a dud into the fun of a lottery.

How far to trust it

The authors acknowledge many limits. They used a single model, Claude Sonnet 4.5, which does not represent generative AI in general. They did not test whether much more detailed templates would improve results. Evaluation was informal, by the two authors, with each game seen by only one of them. Both are experienced game makers and part of the experiment itself; they state that no evaluation with unrelated people was run.

They also write that "is it fun" was the most subjective question and was mostly left to the rater's personal taste. The figures of 31% for Full and 35% for Simple are best read as rough tendencies.

What Fukai points out here is the scale of the comparison. Each genre-template cell has 12 games, built from only four word lists per genre. A difference of a few games between Full and Simple could flip with a different choice of words. The paper runs no statistical tests; the authors' conclusion of "no clear winner" is sound, but it does not support "Simple is better."

Also, the study did not measure players filling in their own words and then playing. The laugh in Mad Libs should come precisely from words you chose yourself. The authors list player studies as future work. I think that is where the real value of this approach will be decided.

Fukai's reading

From here on, this is my opinion. I want to place this study in a shift of interest from the quality of generation to its feel. PCG research has mostly measured whether content is solvable and whether its difficulty fits. This paper counted where the words appeared and how different the three games were. In the vocabulary of design criticism, it is closer to measuring the player's sense of involvement than the quality of the work. Even a game that half the time cannot be finished might make people laugh if their own word comes flying at them as an angry red object. I am waiting for the next experiment that tests that hypothesis.

What to read next

As future directions, the authors list comparisons across several LLMs, generation with small models that run locally, care about the possibility of offensive content, and studies with players. They also mention making microgames with world models (AI that generates video directly to create a playable world) instead of having code written.

Readers curious about that direction can start with Genie (Bruce et al., 2024) and Valevski et al. (2025) on diffusion models as game engines, both cited in the paper. Our digest of a paper rereading video-generation AI as a game engine is another entry point.

To feel words driving a game firsthand, try Scribblenauts Unlimited and Baba Is You. In the first, words become objects; in the second, words become rules. The fill-in-the-blank microgames in this paper sit just before both.

Sources

Papers and related material referenced in this article:

・Fill-in-the-Blank Microgames (Akash Surendran, Seth Cooper, 2026, PCG Workshop at FDG '26, full-text PDF)

・DOI: 10.1145/3815598.3815713 (ACM Digital Library)

・Templates and evaluation data released by the authors (OSF)

・Related work: Genie: Generative Interactive Environments (Bruce et al., 2024)

・Related work: Diffusion Models Are Real-Time Game Engines (Valevski et al., 2024 arXiv / 2025)

Reactions (no login)

Anonymous • one of each per visitor per day

関連シリーズ

Paper Digest第97回 / 全98回

Read next