PAPER-DIGEST · 2026-08-07
Geheeb et al.: Let an LLM Poke at Your Game Design Pillars — Fukai Reads
mixed-initiative design support / formalising design pillars / LLMs in early game design
TL;DR
Game studios lean on design pillars — a handful of natural-language statements that declare what a game is and then act as the yardstick for every later decision. Pillars are authored per project: Combat in God of War, Realism in Duskers. Yet Julian Geheeb and colleagues at the Technical University of Munich report that despite heavy industry use, the academic literature is nearly silent; they found only two high-level sources. This paper lays a first foothold with three contributions: a formal definition of the game design pillar, a dataset of 55+ documented pillars from real games, and a prototype tool called SPINE.
SPINE asks a large language model (LLM — a model trained on large text corpora to generate and interpret natural language) to check whether a pillar is well formed, to rewrite it, to test a pillar set for contradictions and coverage, and to rate on a 1–5 scale how well a proposed feature fits the pillars. Evaluation comes in three parts: a pre-study comparing gemini-2.0-flash against GPT-4o mini, a case study at a 42-hour game jam, and expert interviews with four developers. The conclusions are modest. The tool helps at the moment of putting a pillar into words; the decision-support side looks promising but is thinly evidenced; output-quality variance and transparency remain unresolved. Peer-reviewed at FDG '26.
Introduction
Today's paper is 'LLMs are the Ideal Candidate for Mixed-Initiative Game Design Pillar Workflows' by Julian Geheeb, Marvin Julian Schwarz, Daniel Dyrda and Georg Groh, all at the Technical University of Munich, Germany. It is on arXiv as 2605.09767 and appears in the proceedings of Foundations of Digital Games (FDG '26, 10–13 August 2026, Copenhagen), DOI 10.1145/3815598.3815653, dated 1 February 2026 on the paper itself. So this is a peer-reviewed conference paper rather than preprint-only material — but the conference has not happened yet, citations are effectively zero, and the work has not been exposed to community discussion. Read it with that discount applied.
My reason for picking it is simple: it addresses the stage nearly everyone building a puzzle or a game passes through — deciding what to decide first. And while the title asserts that LLMs are 'ideal', the authors take the trouble in section 1 to note that 'ideal denotes a promising fit worthy of evaluation rather than a proven outcome'. That single sentence changes how the whole paper should be read. Note too that this is not a study deciding whether the tool is good. Following Types 1–3 of Ledo et al.'s evaluation strategies for HCI toolkits, the authors deliberately arrange three angles: performance probing, a demonstration in the wild, and qualitative expert assessment.
Background — why pillars went unstudied
Software engineering has design pillars like Security or Scalability that are broadly agreed across domains. The authors start by contrasting that. Game design pillars must be the opposite of universal: if two games share the same pillars, players simply call one a clone of the other. So pillars are authored fresh each time, in words specific to that project. That per-project nature is exactly what makes them hard to study — a pillar is a fragment of natural language whose meaning shifts with the project and which usually lives buried in internal documents. They appear constantly in industry articles and talks, yet the authors report locating only two high-level academic sources: Zagal (2023) and Luo et al. (2021).
Meanwhile, working with pillars in practice has known pain points. Citing concurrent work by colleagues (Dyrda et al., 2026), the authors list vagueness, conflicting interpretations, difficulty applying pillars to concrete decisions, and insufficient documentation structure. Pillars, in other words, are a tool that is good to have but awkward to operate. That is the opening for an LLM. Creating pillars is natural-language generation; using pillars is natural-language judgement. Both have the shape of things LLMs are good at — and that shape-match is the authors' starting point.
Approach — define the pillar, then put it in SPINE
The authors first synthesise academic and industry sources into a formal definition. A game design pillar, they write, is a normative design construct functioning as a high-level principle for directing and constraining decision-making in game development, composed of (1) a succinct title naming the principle and (2) an expository statement specifying the intended experiential or structural property the game should embody. A structural property means a characteristic of the game's formal system — mechanics, dynamics, progression architecture, interaction loops, rule-based organisation. An experiential property means the intended emotional, cognitive or aesthetic experience of the player: player experience, in short.
Four quality criteria accompany the definition. Clarity (the meaning is unambiguous to stakeholders), Unicity (one pillar articulates one coherent principle), Conciseness (minimal, economical language), Actionability (applicable to concrete design decisions). Three set-level constraints follow: mutual non-contradiction, completeness (the set collectively covers the core experiential goals), and bounded size (deliberately kept small). A typical set is three to five pillars depending on scope, with finer considerations pushed into subordinate sets. This definition-plus-checklist is, to my mind, the most portable part of the paper.
On top sits SPINE — System for Pillar-based INteractive Experience design — a Django backend with a Nuxt4 frontend and a swappable API-based LLM. The interface lets you write three kinds of content: a core design idea, a set of pillars (title plus description), and a feature idea. Four LLM-powered functions: structural pillar analysis (does the title match the description; is the description continuous text; is the intent clear; is it focused on one aspect — each flagged with a 1–5 severity), structural pillar repair (the user chooses whether to keep the original or the LLM version), pillar set validation (three separate prompts for coverage, contradictions and additions), and feature validation (rate a proposed feature 1–5 against the pillars with an explanation). Conciseness and Actionability were left out of the checks for now, since both usually require domain-grounded interpretation.
Findings — what the three evaluations showed
The pre-study (section 4.1) compares gemini-2.0-flash with GPT-4o mini, both chosen for speed and cost since SPINE is meant to be used iteratively. The material is two sets of three pillars: one from a student project called Ordinary, one reverse-engineered from Sea of Thieves, a game the authors know well. Each pillar is queried three times to see whether the severity ratings hold steady. Results appear in Table 1 (Gemini) and Table 2 (GPT). GPT returned essentially the same ratings regardless of the pillar and regardless of whether it had already repaired it — title 3-3-3, Clarity 4-4-4, and so on across the board. Gemini stayed largely consistent, varied between pillars, and cleared nearly all warnings when scoring its own revisions. The authors' reading is cautious: this pattern 'suggests potential limitations in GPT's ability to fully interpret the given task'.
Their prose habits differed too. GPT ran long, expanding descriptions into explanatory text; Gemini reformulated more succinctly. The shared failure is the interesting part. On Sea of Thieves, both models pulled the first and last pillars toward variations of player agency, reducing the set's distinctiveness and obscuring the game's specific design focus. Neither flagged that those two pillars had become too similar. The authors diagnose this as a prompt-design limitation rather than a model limitation: the models were asked to find explicit contradictions, not duplication or redundancy. It is an honest diagnosis and one that transfers directly to practice.
The game jam case study (4.2) ran 42 hours, with one researcher operating SPINE as the primary design tool while two master's students who also work at indie studios contributed discussion and ideation (P1, 26, female; P2, 23, male; both with prior pillar experience). The theme was the 'This Is Fine' meme. The team refined a first pillar into embrace sarcastic resilience; entering a stress-meter idea prompted SPINE to suggest unleash controlled rage, then to flag a contradiction between the two. The team initially disagreed but kept iterating; at four pillars SPINE flagged contradictions again, and renaming 'rage' to 'composure' reduced the conflict while revealing that the set as a whole was too broad for the project's scale. They discarded everything and rebuilt with manage your composure and comical exaggeration. To me this sequence is the paper's most persuasive evidence — not because the output was right, but because it prompted the team to discover their own scale.
The feature side is equally concrete. A workplace setting with a hostile boss and constant calls and emails, entered in the feature field, came back 5/5 with a detailed justification; alternative settings and different levels of description granularity scored lower, matching expectations. The next day the team debated whether the space should be vertically expansive or tightly constrained; SPINE favoured the constrained layout, citing stronger support for composure management. Participants questioned parts of the explanation, which led to the suggestion that AI feedback should explicitly reference the text it evaluated. SPINE was not used after that — the design vision had stabilised and no further high-level decisions remained before the deadline.
The expert interviews (4.3) involved four developers from two studios, in person, recorded, 50–60 minutes each. P3 (24, male, one year professional) and P4 (32, male, five years) work at the same studio, building its second commercial title. P5 (23, female, two years freelancing in game art) and P6 (27, male, co-founder, roughly 1.5 years) are collaborating on their first commercial title. All knew pillars and had used them at least once, though several were unsure how to articulate them precisely. Reception was mixed overall: one very positive, two emphasising potential, one more critical. That critic, P4, opened by saying he was 'kind of biased against LLMs', citing job replacement, distrust of the companies behind them, and environmental impact — yet still called the tool 'pretty decent' for certain use cases.
The individual remarks form a clear pattern. After several repair iterations P4 said, 'Now it's super watered down and abstract.' P5 said, 'Fixing it doesn't really change much. Every time I fix it, it's like adding more water.' Same observation: iterate and the distinctiveness drains out. Against that, P3 said of one revision, 'I kept this pillar vague on purpose with vibe. And I can see the generated version is more specific how this would be achieved, which is more helpful so I take this one,' and elsewhere, 'I would prefer it if it would poke me in the direction of "hey this is missing" instead of completely suggesting a new pillar. I would prefer that for my workflow, and it would be more respectful towards me.' P6 called the feature feedback the best thing about the tool, precisely because it offered 'a more objective look on things'. The authors record it as a promising direction that the least developed feature drew the most praise.
Where to use it
First, the definition and checklist work standalone, with no LLM involved. A pillar is a short title paired with a statement naming an experiential or structural property; it should satisfy Clarity, Unicity, Conciseness and Actionability, and the set should be non-contradictory, complete and small (three to five). If I were building a Sokoban-like, I might write 'every move costs' (each action carries weight that is hard to undo) and 'the board is fully visible' (never surprise the player with hidden information), then re-read them against Unicity. If I had written 'every move costs and yet it feels exhilarating', that is the signal to split it in two. No model required.
Second, contradiction detection is more useful as a ruler than as an answer. The game jam team was flagged repeatedly, kept rewording to reduce conflict, and what they finally gained was not correct pillars but the realisation that the set was too broad for the project's scale. If you are doing hyper-casual PCG (Procedural Content Generation — automatic generation of game content), write out every property you want the generator to hold as a pillar, and read the number of contradictions as a scale indicator rather than something to eliminate. Many contradictions usually mean you are being greedy.
Third, feature validation is most realistic as a first-pass filter for newcomers. P2 in the case study noted that in larger teams, members not deeply involved in design decisions could use it to quickly validate ideas before broader discussion. If you run a daily puzzle with several collaborators, pin the pillars in the README and add two fields to the feature-proposal template — which pillar does this support, which pillar does it strain — and you get the same effect without an LLM. If you do add one, confining it to P6's 'more objective look' role is the safer bet.
Fourth, the paper also tells you where not to use it. Iterating the rewrite function converges toward thinner output; P4's 'watered down' and P5's 'adding more water' are the same phenomenon. Handing the model anything that needs an edge — a pillar title, a game's hook, a deliberate juxtaposition — is a poor bet. Conversely, asking it to point at places where intent has been left vague fits well: being poked is welcome, being written for is not. And Appendix A's dataset of 55+ documented pillars is worth reading on its own. Lay other people's pillars side by side and two schools appear — those who write experiential properties and those who write structural commitments.
Limitations
The authors acknowledge a great deal. The game jam study is limited by its small scale and number of participants, and they explicitly name the bias arising from the researcher's dual role as tool designer and primary user. The interview study has a moderate sample and focuses on exploring functionality rather than measuring performance quantitatively. They also concede that the decision-making side is thinly evidenced because it was used only rarely during the jam. And the pre-study is exploratory by design, since no dataset, baseline or benchmark exists; building a proper dataset is listed as future work.
On the system side they name three limitations. A recurring one across both studies was perceived variability in output quality, consistent with their own analysis in section 4.1. Next, concerns about trust and transparency — one developer suggested the system should be more transparent. Proposed remedies include better prompting, fine-tuning with annotated datasets, and grounding outputs in context through retrieval-augmented generation (RAG — retrieving external material to anchor what the model generates) or agentic workflows.
The third is the most design-interesting: several participants noted that contradictions between pillars are not always undesirable and may be an intentional design choice creating productive tension. The current system's mutual-exclusivity assumptions produced feedback that did not reliably recognise such intentional juxtaposition. The authors' redesign implication is to soften contradiction feedback — flag it as a risk while inviting the designer to confirm whether it is deliberate and to articulate its purpose.
What follows is what I noticed as a reader. First, every pillar in this study is handled at the moment of creation; nothing addresses the moment when existing pillars drift away from what the project has actually become. P6 says from his own experience that pillars' relevance decreased over the course of a project, yet SPINE does not touch that phenomenon. As far as I can read it, the real trouble with pillars lies not in writing them but in the six months afterwards.
Second, every observation window is short — a 42-hour jam and a 30-minute hands-on session — even though the authors themselves write that formulating pillars is unlikely to be completed in a short timeframe. They offer as one explanation for the team's early alignment that defining pillars early, supported by the tool, contributed to a shared vision; but they name small team size and clear roles in the same breath, and the paper contains nothing separating the two. I would not read that as causal.
Third, model selection happens once, in the pre-study, after which every evaluation is fixed on gemini-2.0-flash. GPT's habit of returning identical ratings could plausibly be a sensitivity-to-prompt-structure issue rather than a capability issue, and the paper does not separate those. Since the authors correctly diagnose the missed redundancy as a prompt-design limitation, I think the same suspicion should be pointed at the uniformity of the ratings.
How Fukai reads it
This section is my own interpretation. I would place this work not in the lineage of automating generation but in that of automating forced restatement. Line up the moments where SPINE actually worked and almost none of them involve the model writing something good. What worked was being poked about the clash between rage and composure until the team recognised their own scale; getting a 5/5 back and feeling settled about a setting; and P3 articulating, in his own words, that he had kept a pillar vague on purpose. In each case an outside question arrives and the designer rebuilds their own phrasing. In the vocabulary of design criticism, this is rubber duck debugging (noticing your own error in the act of explaining it to someone) applied to design documents — not a co-author. So when P3 and P6 both asked to be poked rather than written for, I read that not as a complaint about missing features but as an accurate identification of what shape this device should take. And if that reading holds, the paper's largest limitation — variance in output quality — is avoidable not by reducing the variance but by never presenting the output as an answer.
Closing
The nearest companion piece is the same group's SPARC (Geheeb et al., 2025, arXiv:2509.24730) — an attempt to have locally runnable medium-sized LLMs refine game concepts against ten predefined aspects. The authors themselves note that concept creation and pillar creation often go hand in hand, placing the two tools at effectively the same step, and that SPINE's focus on pillars complements SPARC. Reading them back to back is the natural move.
The other is Despain (2013), which the authors cite as prior guidance that pillar development is a process of structured questioning rather than prescription. The 'nudging' feedback participants asked for here is, in part, a return to that older idea. One caution to close: FDG '26 has not yet taken place as of this article. The paper is peer-reviewed but has not been exposed to community discussion, and citations are effectively zero. What to take home is the framework — the definition and the quality criteria. On SPINE's effectiveness, nobody can yet say more than 'four experts and two students found it broadly agreeable after a short encounter'.
References
Papers and materials referenced in this article:
・DOI: 10.1145/3815598.3815653 (FDG '26 proceedings, peer-reviewed conference paper)
・Related work by the same group: Diamonds in the rough: Transforming SPARCs of imagination into a game concept by leveraging medium sized LLMs (Geheeb, Ivan, Dyrda, Anschütz, Groh, 2025, AI4HGI '25 at ECAI '25)
・In-paper citations mentioned in this article: Zagal (2023) and Luo et al. (2021), the two academic sources on pillars the authors located; Despain (2013); Ledo et al. (2018) on evaluation strategies for HCI toolkits; Dyrda et al. (2026), concurrent work on challenges in pillar workflows; Eigner and Händler (2024) on determinants of LLM-assisted decision-making
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 52 of 90
