PAPER-DIGEST · 2026-09-29
Endrovski et al.: The people using level generation least are the ones who build levels — Fukai Reads
PCG and tool design — what 120 developers actually want from generation tools
What does this paper show?
The short version: when 120 game developers were asked, the people using procedural level generation the least were the very people whose job is to build levels. Artists scored an average of 0.58 on a usage scale; designers scored 0.33, a gap the authors report as statistically significant (p = 0.001).
The second finding is that developers do not want tools that do everything for them. Asked what matters in AI-assisted generation tools, 73.3% chose creative control over the final output. Only 6.7% wanted full automation. The authors sum it up this way: developers will let AI ride shotgun, but they insist on staying in the driver's seat.
For puzzle makers, the takeaway is simple. Generation tools are chosen for the steering wheel they offer, not for how clever they are. If you build a generator, make it something a designer can stop, correct, and understand first. This article walks through the numbers behind that claim.
Townscaper (Oskar Stålberg, 2021). The player chooses where to build; the shapes of the buildings resolve themselves. One of the works the paper cites as practitioner-driven PCG. Image: Steam store page
Who wrote it, and where was it published?
The paper is titled "What game developers actually want from procedural level generation tools." Its authors are Bojan Endrovski (Delft University of Technology and Breda University of Applied Sciences), Joris Dormans (Ludomotion, a studio in Amsterdam), and Rafael Bidarra (Delft University of Technology): two researchers and one working game developer.
It was presented at the 17th PCG Workshop, held in Copenhagen on August 10, 2026. PCG stands for Procedural Content Generation, meaning levels, terrain and other game content produced by a program. The workshop ran alongside FDG '26 (Foundations of Digital Games), and the paper appears in the ACM proceedings with a DOI. Note up front that this is a short workshop paper, not a long main-track paper.
I picked it today because this series has covered many generation papers: a diffusion model that draws Sokoban levels, a system that evolves the programs that build levels, and more. Very few papers ask the practitioners directly whether they actually want these tools. Having read the map of the technology, I wanted to lay the map of its users on top.
Why hasn't procedural generation spread further in practice?
In research, the number of level-generation techniques grows every year. Examples include WFC (Wave Function Collapse, which fills a grid one cell at a time while respecting rules about which tiles may sit next to each other) and graph grammars (which first build the connections between rooms from rules). On the commercial side, Houdini and Unreal Engine's PCG framework are widely used.
Even so, few commercial games generate their whole levels. The authors looked at the top 20 best-selling roguelikes on Steam as of February 2026. None of them generated full level geometry procedurally, and nearly half used no level generation at all (Appendix Table 4).
One barrier the authors point to is that many prototype tools demand an "algorithmic mindset" from their users: you cannot steer the output without understanding the internals. Yet some tools, like Townscaper or the text generator Tracery, became widely loved. The paper notes that these came from practitioners who combine technical skill with design sensibility.
So what do designers actually value in a tool? According to the authors, this question matters to researchers and tool makers alike, yet has received little attention. The paper tries to fill that gap with a large survey.
Slay the Spire (Mega Crit, 2019). In the paper's Appendix Table 4 it is classified under layout generation among the best-selling roguelikes. Image: Steam store page
How did the authors collect developers' views?
The method is an online survey of 21 questions, mixing single choice, multiple choice, rating scales, rankings and open text. Every question was optional, so that nobody padded answers just to move on. Responses were anonymous, and no personally identifying information was collected.
The respondents were 120 game development professionals, recruited through the PCG mailing list, Discord communities, and LinkedIn via institutional networks. According to the public analysis tool, recruitment ran for 60 days. Completion time averaged about 45 minutes, with a median of 10 minutes. My guess, and it is only a guess, is that the high average comes from a few people leaving the survey open.
The analysis used three tools. The first was numerical comparison; the artist-versus-designer gap was tested with a Mann-Whitney U test (a way to check whether the values of two groups are shifted relative to each other). The second was thematic coding of open answers: the authors built 11 candidate themes and refined them into 9.
The third is the interesting one. The authors took 59 level-generation papers from past PCG Workshops and the survey's open-text answers, and counted which topics each side talked about, using TF-IDF (a way of surfacing words that are unusually frequent in a given document). In effect, they measured with word counts the gap between what researchers write about and what practitioners worry about.
What did the numbers show?
First, who answered. The largest groups were programmers and technical designers (38.3%), technical artists (21.7%), game designers (18.3%) and level designers (10.0%). Experience was spread fairly evenly from 0–2 years to 10+ years, and answers varied little by experience. For engines, 65.8% used Unreal Engine and 41.7% used Unity (multiple answers allowed).
What do they use generation for (multiple answers, Figure 5)? World building leads at 65.0%, then level layout and structure at 55.8% and enemy or NPC placement at 30.8%. Puzzle generation was 12.5%. For level generation specifically, 10.0% said always, 20.8% often, 23.3% sometimes, 35.8% rarely and 9.2% never (Figure 6).
Here the roles split. Scoring "never" as 0 and "always" as 1, artists (n = 30) averaged 0.58 and designers (n = 34) averaged 0.33 (Table 2): a difference of 0.25, or about 76% higher for artists. The paper states that no designer considered PCG essential to their process, and that around 60% of level layout generation in the sample was done by non-designers.
The survey also asked about concerns (Figure 7). The top answers were lack of artistic control (48.3%), time investment versus benefit (44.2%), difficulty debugging (41.7%) and unpredictable results (38.3%). In the "other" answers, many respondents named the "procedural oatmeal" problem: generated content that all feels the same.
On AI, the answers pointed one way. The most popular role was a tool that enhances specific components (43.3%); 35.0% preferred rule-based generation without AI; full automation drew 6.7% (Figure 14). The top AI concerns were unpredictable results (48.3%), black-box behavior (43.3%) and loss of designer agency (42.5%) (Figure 16). Regrouping related items, the authors found that 93% of respondents cited control-related concerns (Table 3).
Finally, the research–practice gap (Figure 22). Measured in occurrences per 1,000 words, research papers were dominated by "algorithms and models" (41.0), with "control and flexibility" at only 2.9. Practitioners' open answers were the mirror image: "control and flexibility" at 30.1 and "algorithms and models" at 2.7. The topic researchers discuss most is the one users barely mention.
How can puzzle makers use these results?
One. If you are writing a tool that generates levels for a Sokoban-style pushing puzzle, design it to be steerable mid-process rather than to output a finished level in one shot. In the paper's rankings, control over generation constraints, mixing handcrafted and procedural content, and visual previews came out on top (Figure 10). Let the designer lock the walls and boxes they placed, and have the generator fill only the rest; that satisfies all three at once.
Two. If you generate a daily puzzle, keep a record of why each board was chosen. Difficulty debugging (41.7%) and unpredictable results (38.3%) ranked high among concerns. Store each board's random seed (the number that reproduces it) along with the scores that made it pass your filters. When a strange puzzle ships one day, you can trace the cause right away. This helps even when the tool's author and user are the same person.
Three. If in your small team a programmer writes the generator and a designer judges the levels, name the controls in the designer's vocabulary. In the paper, generation was carried by artists and technical staff, while designers largely stayed away. A knob called "moves that need an insight" rather than "search depth" makes it easier for designers to put their own judgment into the tool.
Four. If you are torn between generating whole levels and hand-building everything, consider the middle path. Appendix Table 4 lists many best-selling roguelikes under "layout generation" or "chunk assembly", where handcrafted sections are put together by the game. Hades II falls in the latter. For a puzzle game, that means polished hand-made rooms arranged by a generator.
Five. If you bring an LLM (large language model, an AI that writes text after learning from huge amounts of it) into your pipeline, give it a narrow role. What respondents wanted was a tool that enhances specific components, not full automation. For example, let it draft hint text only, while a rule-based generator and a human own the board and the solution. That split avoids much of the "unpredictable results" worry that topped the list.
Hades II (Supergiant Games, 2025). The paper's Appendix Table 4 classifies it under chunk assembly: handcrafted sections combined by the game. Image: Steam store page
How far can these results be trusted?
The authors acknowledge four weaknesses. First, their own interpretation shaped both the question wording and the coding of answers; to counter this, they published the data and analysis tool. Second, they merged programmers and technical designers into one role, which in hindsight obscured meaningful distinctions.
Third, the research–practice comparison (Figure 22) is narrow: one open-ended question on the industry side, one venue, the PCG Workshop, on the research side. Fourth, the AI questions used "AI" as an umbrella term for generative AI and LLMs, setting aside the fact that many established generation techniques are also AI.
What I, Fukai, would add starts with how respondents were recruited. The main channels were the PCG mailing list and Discord, which naturally attract people already interested in generation. That may be why only 9.2% said they never use level generation. These figures should not be read as industry-wide usage rates.
Next, the groups being compared are small. The headline artist-versus-designer gap compares 30 people with 34. Level designers were 10.0% of the sample, about 12 people. There were no attention checks, and the median completion time was a brisk 10 minutes. None of this refutes the findings, but I think it is fair to read this as a survey that shows direction, not magnitude.
Finally, puzzle generation is a quiet voice here. Only 12.5% use generation for puzzles; most answers likely concern 3D worlds and action levels. Puzzles must be rigorously checked for solvability and unique solutions, which changes things. The use cases above are what I took home while keeping that difference in mind.
Fukai's reading
What follows is my own opinion. I want to place this paper in the lineage of mixed-initiative level design tools, where human and generator take turns editing, represented by tools such as Tanagra (2010). For over a decade that line of work has tried to build generators that people can steer. This paper confirms, with numbers, that such tools are still not standard practice. In the vocabulary of design criticism, what designers refuse to hand over is not labor but judgment. So however good the levels a generator produces, it will struggle to be chosen as a tool unless it shows the reasons behind its choices. The authors themselves speculate at the end that the barrier for designers is not technical. I would put it this way: a generator should first be a mirror of its maker's judgment.
What should you read next?
If you want to go deeper, start with the analysis tool the authors published. You can filter all 21 questions by role and experience and browse the answers yourself. I plan to find time to see what changes when you filter to level designers only.
For the technology side of the map, earlier installments of this series help: evolving the programs that build levels rather than the levels themselves, and using gravity and time as generation coordinates, both push how clever generators can be. Having an LLM probe design pillars reads well alongside today's paper as an attempt to put AI in the passenger seat.
The same PCG Workshop 2026 also produced a paper on crosswords that get easier when you are stuck, an example of a generator adapting to the player rather than the designer. Generation steered by the maker, and generation that adapts to the player: put the two side by side and the place of procedural generation becomes much clearer.
Sources
Papers and related material referenced in this article:
・DOI: 10.1145/3815598.3815682 (ACM FDG '26 proceedings)
・Procedural Level Generation Survey — the authors' public data and interactive analysis tool
・PCG Workshop paper database (including the 2026 accepted papers)
・Related work: Tanagra: a mixed-initiative level design tool (Gillian Smith, Jim Whitehead, Michael Mateas, 2010, FDG '10)
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 96 of 98
