PAPER-DIGEST · 2026-10-02
Xu & Verbrugge: Small Local AIs Translate a 5×5 Grid into a Different Skill Every Time — Fukai Reads
Skill generation — resolving the coherence–diversity tension with local language models
TL;DR
Roguelikes give us a different run every time, but rarely a different skill every time — skills and units are still hand-authored. Kaijie Xu and Clark Verbrugge of McGill University tried plugging small, locally-run language models into exactly that gap.
Their testbed is a browser game, Neural Spore Genesis. Players arrange Fire, Water, Earth, Wood and Metal cells on a 5×5 grid, and a model 'translates' the layout into a combat skill. The models almost always produced a skill matching the dominant element, while still producing different skills from the same grid each time — a combination none of the four conventional baselines reached. But the models barely read 'vertical line' shapes, and no player study has been run yet.
Introduction — a paper about generating skills, from FDG 2026
Good morning. Strong hot drip coffee, one printed paper. Today's is Generate Diverse Skills with Large Language Models (listed on the workshop page as “Generate Infinite Skills with Large Language Models”), by Kaijie Xu and Clark Verbrugge of McGill University. We have covered an earlier paper by the same pair, on using gravity direction and time as coordinates for level generation.
It appeared at the PCG Workshop held with FDG '26 (Foundations of Digital Games, August 10–13, 2026, Copenhagen), and has a DOI. PCG means Procedural Content Generation — generating game content automatically. Workshop papers are shorter than main-track papers and often report work in progress. It is also very new, so it has not yet been widely discussed.
I picked it because the subject is skills, not levels. For people making puzzles or roguelikes, the system that gives meaning to what the player built is often the core of the game. Can that translation be handed to a machine? Few studies have checked with numbers.
Background — randomness reaches the run, not the mechanics
According to the authors, the replayability of roguelikes and auto-battlers (games where you set up a team and the fights resolve on their own) comes from randomized rewards and encounters. That randomness rarely reaches the substance of skills and units. Unconstrained randomness tends to produce skills that are mechanically noisy, strategically unreadable, and unrelated to what the player was building.
Existing generators are not naive. Noita's spell combinations are cited in the paper as a rich spell engine. But such systems run inside a fixed vocabulary of effects and fixed ways of combining them. LLM-based game generation has grown too, but mostly uses text as the control channel. The authors find almost no work that turns a player-built spatial layout into executable combat mechanics.
Noita (Steam store screenshot), cited in the paper as a commercial example of a rich spell engine
Two demands collide. One is coherence: a Fire-dominant cross should yield a fire-based area effect, not a water projectile. The other is diversity: building the same shape in another run should yield a different skill. The authors call this the coherence–diversity tension.
Approach — summarize the grid, let the model write the skill blueprint
The game loops through three phases. Buy: purchase from four randomly drawn elemental cells (you can pay to re-roll). Construct: distribute cells across five skill slots, each with its own 5×5 grid. Battle: the compiled skills fire automatically at procedurally scaled enemy waves, and drops feed the next buy phase.
The grid is not handed to the model raw. It is first summarized into features: dominant element, shape category (horizontal, vertical or diagonal line, dense block, L-shape, T-shape, corner cluster — seven in all), density, a left–right symmetry score, and element mix (mono, binary, ternary, diverse). A simple heuristic also adds an intent signal: aggressive, defensive, utility or balanced.
From that summary the model writes a skill blueprint in JSON (a standard text format for data). The blueprint follows a schema — a fixed set of fields. There are 11 action types (projectile, zone, heal, shield, buff, dash, summon, beam, nova, trap, echo) and 11 modifiers (lifesteal, stun, slow, knockback, damage-over-time, pierce, chain, splash, homing, crit, bleed). Damage, range and cooldown have bounds. The model also writes a name and description.
Four baselines were compared: a rule-based generator that maps shape and intent to a fixed skill; a random generator that picks anything schema-valid; a template generator using 20 hand-written skills with per-grid power scaling; and a calibrated sampler drawing from hand-set probabilities per element and shape. The two models, Llama 3.1 8B and Qwen2.5 14B, ran locally on an RTX 4080. The corpus was 7 shapes × 5 elements × 3 mix variants = 105 grids, sampled 5 times each — 525 skills per condition — scored with 11 automatic metrics.
The yardsticks — measuring 'fits' and 'varies' separately
The 11 metrics fall into three groups. First, is it broken? Schema validity and whether numbers are in range. Second, does it fit? The centerpiece is element coherence: 1 if the skill's main element matches the grid's dominant element, 0.5 if only a secondary element matches, 0 if absent. Another metric checks whether the chosen action type suits the shape.
Third, does it vary? The key metric is intra-config diversity: how different the five skills from the same grid are, measured as the mean distance between their embeddings (an embedding turns text or data into numbers where closeness reflects similarity). Zero means the same skill every time. Others measure how evenly action types are used and how many of the 11 modifiers appear at least once.
There is no player study; every score is automatic. The authors state this as a limitation themselves, which I return to below.
Findings — small models were both on-target and different every time
Key numbers from Table 2: element coherence was 0.994 for both Llama 3.1 8B and Qwen2.5 14B. The rule-based generator scored 1.000 but had intra-config diversity of 0.000 — the identical skill every time — as did the template generator. Random and template generators reached only 0.301 and 0.371 coherence. Diversity was 0.429 for Llama, 0.433 for Qwen, 0.371 for the calibrated sampler.
In the authors' words, the two models occupy a high-coherence, high-diversity region that no baseline reaches (Figure 6). Kruskal-Wallis tests (a check of whether differences among three or more groups could be chance) gave p < 10⁻⁵ for the coherence metrics. Modifier utilization was 1.000 for both models (all 11 used) versus 0.455 for the calibrated sampler.
The weakness is just as clear. Shape–action coherence was 1.000 for rule-based, but 0.422 for Llama and 0.554 for Qwen. By shape (Figure 4), dense blocks scored 0.97 (Llama) and 1.00 (Qwen), but vertical lines only 0.07 and 0.09. The authors suggest the models' spatial prior does not reliably distinguish line orientation. On how well the description matches the actual mechanics, hand-written templates scored highest at 0.503, models 0.42–0.43.
A temperature ablation (temperature is the knob for how random the model's output is; Table 3, Llama only, 50 grids × 2) raised diversity from 0.297 at 0.3 to 0.472 at 1.2, while element coherence stayed at 1.000 up to 0.9 and was 0.995 at 1.2. The authors conclude temperature is the primary control for creative diversity. The larger Qwen wins on shape coherence and validity but runs at about 8 tokens/s versus 15 — roughly 1.9× slower.
Use cases — for people making puzzles and roguelikes
One. If you are making a game where arrangement determines power, like Backpack Battles, the paper's pipeline is a ready template: summarize the layout into a few features, and let the model write only a fixed-schema blueprint. Don't ask the model to understand the raw board. That summarize-first step plausibly underpins the 0.994 element coherence.
Backpack Battles (Steam store screenshot), an example of a game where arrangement itself is power (not covered in the paper)
Two. Keep the rule-based generator alongside the model. Rule-based scored perfectly on shape fit; the models scored near zero on vertical lines. So let rules fix the skeleton for shapes players clearly mean something by (like line orientation), and let the model handle names and modifier combinations. A player who builds a vertical line and gets nothing vertical will read it as unfair.
Three. Use temperature as a rarity knob. Higher temperature raised diversity while keeping element coherence almost intact. Low temperature for ordinary rewards, high for chests or special events, could stage rarity. But the paper does not measure balance, so you must check yourself that high-temperature skills are not overpowered.
Four. Apply it to puzzles where a drawn shape becomes a rule. Summarize the player's figure into features, and have a model translate it into a short rule blueprint (push, pull, freeze…). Schema bounds act as a fence against board-breaking rules. Schema validity was 1.000 for Qwen and 0.984 for Llama — still a few percent broke, so validating the blueprint on the receiving side is a must.
Limitations — what the authors admit, and what I noticed
(a) The authors acknowledge four weaknesses. No human evaluation: embedding similarity and rule conformance cannot measure perceived coherence, surprise, agency or satisfaction. A closed schema of 7 shapes, 5 elements and 3 mixes, so no claim of generalization to messier constructions. Only 8B–14B local models, so it is unknown whether scale fixes the shape deficit. And a single prompt template, with no prompt-style ablations or few-shot examples.
(b) What I, Fukai, would point out first is that balance is not measured. As far as I read, the paper does not evaluate skill power or win rates. In an auto-battler, making every different skill worth using is harder than making skills different. There is still distance between 'diverse' and 'playably diverse'.
I would also point out the nature of the yardsticks. Shape–action coherence is scored against a set of actions 'expected' for each shape. The rule-based generator's perfect score can be read as a consequence of it running on nearly the same expectations. Conversely, whether a model's different choice for a vertical line is truly 'wrong' can only be settled by players. And because diversity is measured by embedding distance, different wording in names or descriptions alone could raise it; that should be kept apart from how different the effects actually are.
Finally, speed. The paper reports tokens per second, but not how many seconds one skill takes, nor whether skills are generated during play or prepared in advance. Anyone shipping this must design waiting times and pre-generation themselves.
Fukai's reading — the player's layout becomes source code
This part is my own opinion. The paper describes the grid as being 'compiled' into a skill (compiling means converting something a person wrote into a form a machine can run). I want to place that word choice in a larger design current. A spell engine like Noita executes the player's combination literally: rules decide what a build means, and surprise comes from the combinatorial explosion of rules. This work goes the other way, handing the job of interpreting what a build means to a language model. In design-criticism terms, it is close to automating the interpretation of player expression. Read that way, the vertical-line weakness is more than a precision gap. When interpretation drifts, do players feel it as a delightful discovery or as a betrayal? Measuring that with people is, as I read it, the heart of this research's next step.
Closing — to widen the map
If you want to go deeper, read Smith and Whitehead's 2010 paper, the origin of measuring a generator's expressive range; this paper's diversity metrics sit in that lineage. For practical LLM-based PCG, Nasir and Togelius (2023) is a good entry point. Withington's (2025) analysis of the possibility space of a billion spells, which the paper cites, is a counterpart map from the rule-built side.
The authors name a player study, larger models and few-shot prompting as next steps. For auto-battlers they propose passing race, class and synergy variables as structured input, while leaving balance and validation to game-specific systems. When the follow-up appears, I will print it out and reach for the colored pens again.
References
Papers and materials referenced in this article:
・Related: Practical PCG Through Large Language Models (Muhammad U. Nasir, Julian Togelius, 2023, arXiv)
Reactions (no login)
Anonymous • one of each per visitor per day
関連シリーズ
Paper Digest第99回 / 全99回
Read next
Related reviews
Desktop Dungeons: Rewind
Un puzzle-roguelike au tour par tour, où l'on explore un donjon tenant sur un seul écran et où l'on abat, un par un, des monstres plus forts que soi, en s'appuyant sur une règle : révéler une case inexplorée restaure la santé. QCF Design a refait en 3D son Desktop Dungeons de 2013, avec un système de rewind pour annuler des coups et la construction de royaume conservée.
Isle of Arrows
Un tower defense à tendance puzzle : on place une à une des tuiles distribuées au hasard sur une île flottante, on prolonge les chemins, on aligne des tours contre les vagues d'envahisseurs. Une tuile indésirable peut être passée, une à la fois, contre des pièces. Le croisement de Gridpop entre jeu de plateau et tower defense.
Duskers
On envoie des drones dans des vaisseaux abandonnés, on récupère carburant et pièces, puis on passe à l'épave suivante. Tous les ordres se tapent au clavier, et l'on ne voit qu'une caméra et un point de capteur qui signale une présence sans dire laquelle. Un jeu solo de Misfits Attic où l'on organise ses coups dans le noir.


