PAPER-DIGEST · 2026-08-14

Bazzaz and Cooper: Comparing Generative AI to PCG Across 500,000 Steam Reviews — Fukai Reads

Generative AI / PCG / player reception / disclosure regime

TL;DR

"If we drop generative AI into our game, how are players actually going to take it?" A new paper by Mahsa Bazzaz and Seth Cooper at Northeastern tries to answer that question with 508,192 Steam reviews. They split Steam games into two groups: 5,186 titles that use Procedural Content Generation (PCG, the classic technique that builds terrain and items from rules and random numbers), and 5,970 titles whose developers openly say they use Generative AI (models like large language models or image diffusion).

The numbers are almost blunt. PCG titles are recommended 86.3% of the time; generative-AI titles, 68.4%. That is a 17.9-point gap in favour of PCG. On sentiment, PCG reviews split 69.0% positive to 31.0% negative, while generative-AI reviews split 53.0% to 47.0%. So the first finding is: PCG has been quietly accepted as a generative technique for years, but the moment a game admits to using generative AI, its ratings drop.

A thematic analysis of 600 reviews then shows why. Players read the use of generative AI as a signal of low developer investment; the more visible the errors, the stronger the signal. Alongside that sit ethical rejection on principle, mistrust when disclosure and evidence do not match, and an expectation about what a good use would look like. The paper is arXiv:2608.11539 (submitted 12 August 2026) and already carries an ACM DOI 10.1145/3831347, which means it has been accepted at one of the ACM venues.

Introduction

The authors are Mahsa Bazzaz and Seth Cooper, both HCI and games researchers at Northeastern University. Cooper is well known as the creator of the citizen-science game Foldit, and he has already collaborated with Bazzaz on a CHI '26 paper about how simply believing a level is AI-made changes the experience (I covered a related piece earlier under paper-bazzaz-perceived-creator). Today's paper is an arXiv preprint (submitted 12 August 2026); an ACM DOI, 10.1145/3831347, has been assigned, which means it has been accepted at some ACM venue, though the preprint itself does not name the conference. So I will introduce it as an author-accepted manuscript for an ACM venue.

Why I chose this paper today is that its question is unavoidable for anyone building puzzle games. Whether or not to add generative AI has become a daily discussion. But few people have tried to measure "how players will actually see it" from anything as flat and large as 500,000 Steam reviews, rather than from guesswork or the mood of Twitter. This paper does exactly that. It looks like a study about feelings, but it really is one about observational grounds for a design decision.

Background

A little history first. Generation in games long predates generative AI. Ever since Rogue in 1980, PCG (procedural content generation, the practice of assembling content with rules and random numbers) has run through Minecraft, Dwarf Fortress, Spelunky, No Man's Sky, Diablo and Hades. Generation itself is not new to players; it is bound up with a great deal of value.

Generative AI built on large models such as LLMs and diffusion image models has only appeared in commercial games in the last two or three years. Steam changed its policy in 2024 to require disclosure of AI-generated content, and developers now have a field where they must describe things like "assets made with AI" or "AI-driven NPCs". That disclosure regime is what makes this study possible: because disclosure exists, the authors can pick out only those games where the developer says generative AI is used, and build a real comparison group.

The purpose of the study was to put numbers on this contrast. The authors ask three questions. (1) On the Steam marketplace, how have PCG and generative AI spread differently? (2) How does player reception differ, and what context widens the gap? (3) What reasons do reviewers actually give? This three-step frame is what lifts the paper above the flat statement that "players hate AI".

Approach

The method is a two-stage combination of patient observation and a magnifying glass. In the first stage, Steam games are split into two groups: 5,186 titles that self-report using PCG and 5,970 titles that disclose generative AI. The split uses Steam's descriptions and the AI-disclosure field. From these roughly 11,000 games the authors collected 508,192 reviews: 341,447 for PCG and 166,745 for generative AI. The generative-AI side has more games but fewer reviews, reflecting how young the category still is.

The analysis is first quantitative: the recommend-rate (Steam's thumbs-up) and the sentiment split (positive versus negative) are compared across the two groups. Then a qualitative pass extracts 600 reviews from both groups for thematic analysis — the qualitative technique where humans read carefully and cluster recurring bundles of meaning under labels. This is not an automated summary; it is the human-read heart of the paper.

As a designer I want to flag one methodological point: the choice of comparison group. PCG is placed here as the elder statesman of generative techniques — Minecraft-like terrain, Diablo-like dungeons, Spelunky-like levels all sit in this pile. Contrasting generative AI against PCG lets the authors show that players are not against generation per se. Without such a comparison the finding would flatten into a plain "AI is disliked", and much of what is interesting would be lost.

Findings

The quantitative results are the part of the paper you can read most coolly. PCG titles have a recommend rate of 86.3%; generative-AI titles, 68.4%. The gap is 17.9 points. On sentiment, PCG splits 69.0% positive to 31.0% negative; generative AI splits 53.0% to 47.0%. Nearly half of generative-AI reviews being negative is the strongest quantitative finding.

The thematic analysis surfaces five bundles. (1) Perceived Quality Deficit: players read generative AI use as evidence of thin investment, and treat visible errors — malformed hands, repeated lines, broken textures — as "proof that it was not really made properly". (2) Ideological and Ethical Resistance: a rejection unrelated to output quality, framing the use of generative AI as taking work from illustrators, translators and writers, and expressing fear of the industry normalising it.

(3) Acceptance: a real, if smaller, current of voices that does not reject generative AI wholesale. The acceptance is conditional: "in proportion", "where it truly adds value", "as new gameplay rather than as cost cutting". (4) Transparency and Trust: even when a developer says the AI was used lightly, mismatches between that claim and the actual credits or the grain of the art collapse trust in the developer. Disclosure is being evaluated against evidence, not just as an isolated statement.

(5) The Bar Keeps Rising: a small but interesting theme where a minority of reviewers criticise games for not using AI where they think it should be — say, translation or accessibility. What this shows, as I read it, is that player expectation is not a binary of "AI in or out", but a fine network about where to use it and where explicitly not to.

Usecases — How Makers Can Apply This

Read this paper from the perspective of a team making a daily puzzle site like Puzzlebyrinth and at least three concrete uses come out. First: do not put PCG and generative AI on the same shelf. Players now look at the two through different lenses. If your puzzle boards come from deterministic rules plus a seed, that is classical PCG, not generative AI. There is no need to advertise "AI generates the boards" — and if you do, ratings drop. Framing it as "a near-infinite space of boards from classical algorithms" puts you on the Diablo / Spelunky side of generation, where players actually value it.

Second: disclosure is not where the fight ends; it is where the alignment with evidence begins. If you do use an LLM for translation or hint text, disclose it, but also line the disclosure up with things a player can check — named human editors per language, a review history for hints, a small marker on AI-generated hints. The failure mode the Transparency & Trust theme flags is the developer who says "lightly used AI" while most of the art is clearly generated. That gap is the most damaging pattern in the paper.

Third: match the Acceptance theme by using generative AI for new play, not for asset reduction. What players there accept is generative AI used where it enables things AI uniquely enables — reshaping a hint's voice in response to a player's past solve path, or, for a word puzzle in Japanese, rewriting hints to a speaker's age while preserving meaning. If AI is doing a move only AI can, reviews shift. Simply "illustrations were prepared with AI" is, in this paper's frame, the worst-positioned use.

A fourth companion use: do not leave the disclosure field empty. Steam's disclosure regime has been in force since 2024, and refusing to say anything is itself starting to read as suspicious. On a web puzzle site the same logic applies: even one sentence at the end of each article — "Written by Fukai, no AI assistance was used" — hardens the foundation of trust. The paper studies Steam reviews, but the disclosure-and-evidence loop transfers directly to article writing.

Limitations

The authors themselves list three limitations. First, the analysis is restricted to English-language reviews. How players in Japanese, Chinese or Spanish-speaking communities perceive this is not visible here. Attitudes to AI and labour ethics vary across cultures, so this is a reason to soften the paper's conclusions by a step. Second, Steam release dates can be edited later by developers, so any time-series analysis carries an approximation. Third, the thematic analysis was restricted, via keyword filters, to reviews that explicitly mention AI, missing implicit reactions that never surface as the word "AI".

What I would add as Fukai is a reservation about causation. The study observes a relationship between "the developer says generative AI is used" and "the rating drops". Whether that is because AI itself is disliked, or because the population of AI-disclosing games happens to include many small indie titles with thinner reviews, cannot be fully disentangled from this analysis alone. A regression that statistically controls for third variables — studio size, price bracket, genre distribution — would give the claim a stronger footing.

One more thing I noticed as Fukai: games that do use AI but do not disclose it are automatically excluded. Steam's AI-disclosure field only started operating in 2024, and works from before that period, or ones that avoid the field, do not fall into the "generative-AI games" bucket. So the conclusion is really "how players react to disclosed generative AI", not "how they react to games that used AI". For a practitioner that distinction is not small.

Fukai's Reading

I want to place this paper inside a design-critical claim: disclosure is not the outcome, it is the entrance. Quantitatively the paper shows the coarse fact that a single declaration in Steam's disclosure field is worth 17.9 recommend-rate points. But the qualitative material — the Transparency and Trust theme, the Bar Keeps Rising theme — really says something else: that players go on stubbornly reconciling that declaration with the artifact for as long as they play. In the vocabulary of design critique — in the vein of what Michael Cook and Georgios Yannakakis have long said about games continuously producing objects to evaluate — this is the generative-AI edition of that argument. Disclosure is not a breakwater; it is the entrance to a long conversation.

Closing

For readers who want to widen the map, the same Bazzaz and Cooper CHI '26 paper, Human Bias in Perceiving AI-Generated Content in Games (an experiment showing players cannot really tell AI from human levels apart, but rate the ones they believe are AI more harshly), is a good companion. Today's paper handles the macro side with 500,000 reviews; the earlier one handles the micro side with 142 participants. Reading only one is single-eyed viewing.

For readers who want to trace the theoretical thread, the Player-AI Interaction line of HCI work from Guzdial and others, plus industry reports around Steam's 2024 disclosure policy change, together give the atmosphere for why by 2026 disclosure has become the axis players rate on. What I want to see next, personally, is a follow-up at CHI or CHI PLAY that reruns this comparison with price, genre and studio size statistically controlled, and asks how much of the 17.9-point gap really survives.

Sources

Papers and related material referenced in this article:

・Player Perceptions of Generative AI in Games: A Steam Review Analysis (Mahsa Bazzaz, Seth Cooper, 2026, arXiv preprint / accepted at an ACM venue)

・DOI: 10.1145/3831347 (assigned by ACM)

・Related (earlier work): Human Bias in Perceiving AI-Generated Content in Games (Bazzaz and Cooper, CHI '26)

・Related: Procedural Content Generation in Games: A Survey with Insights on Emerging LLM Integration (Zohaib et al., 2024)

・Related: For Steam's 2024 generative-AI disclosure policy see Valve's official announcement "AI Content on Steam"

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 57 of 91

Read next