PAPER-DIGEST · 2026-10-03

Li: 'Make It Tall' Worked, 'Make It Hollow' Didn't — Mapping Which Requests a Neural Generator Honors — Fukai Reads

Neural PCG — which structural constraints can a generator learn?

TL;DR

"Make the towers tall and the buildings symmetric." Which of those requests will a learned generator actually honor? Alex Chengyu Li, an independent researcher in Tokyo, turned that question into a map for a neural generator of Minecraft buildings.

The 14 measured properties split into three regimes: 9 Controllable (they move a lot when asked), 4 Approachable (they move a little), and 1 Unresponsive (the share of air inside the building). And how much an un-requested property moves could be partly predicted from the training data (rank correlation 0.879). But only 10 properties went into that test, and the author says plainly that the sample is too small to generalize.

Introduction — knowing which requests will stick, before you pay for them

Good morning. Strong hot drip coffee, one printed paper. Today: Which Structural Constraints Are Learnable? A Regime Map for a Minecraft Voxel Generator.

The author is Alex Chengyu Li, listed as an independent researcher in Tokyo, with no university or lab affiliation. It appeared at the PCG Workshop (Procedural Content Generation) held with FDG '26 (Foundations of Digital Games, August 10–13, 2026, Copenhagen). It is a peer-reviewed workshop paper with a DOI.

I picked it because the question is so practical. There is plenty of work on putting conditions on generative models. There is much less on estimating, before training, whether a condition will work. Retraining costs time and money. Getting a rough answer up front is valuable on its own.

The paper is about two months old and has not yet been widely discussed. Please read it with that in mind.

Background — inside one model, some requests work and some don't

The paper opens with a level designer's problem: "I want tall towers and symmetric buildings — which will the generator honor?" Learned generators (ones that pick up how to build from data) can be steered by conditions. But some conditions work well and others barely work, and the author notes that this happens within the same model.

With hand-written rules, things are simpler: whatever the rules say is always kept. In Townscaper, buildings assemble themselves around what you place. When you write the rules, you know what they guarantee. With learned generators, what is kept depends on the mix of data and model.Screenshot of TownscaperStore screenshot of Townscaper (Oskar Stålberg), where buildings assemble themselves around what you place. Image: Steam

The author's target is this "you only find out by trying" situation. There is no rule of thumb to lean on before running expensive experiments. So the paper asks two things: which constraints are learnable, and can that be predicted from the training data alone?

Approach — six condition tokens in, fourteen properties measured out

The data: 10,310 Minecraft builds from three sources including 3D-Craft, normalized to 32×32×32 voxel grids (shapes made of cube blocks), with block types remapped to 513 tokens.

The generator has two stages. A VQ-VAE (which compresses a shape into a short sequence of "part numbers" and can decode it back) turns each building into an 8×8×8 grid of codes. Then a model predicts those codes one at a time — the same idea as a language model continuing a sentence word by word. Each stage has about 40 million parameters (the model's adjustable knobs).Screenshot of TeardownTeardown (Tuxedo Labs), another game built from voxels. Store screenshot. Image: Steam

Conditioning uses CFG (classifier-free guidance, which exaggerates the difference between conditioned and unconditioned outputs to push toward the condition). There are six condition tokens: height, size, footprint, symmetry, enclosure, and complexity, each bucketed into categories.

For evaluation, 40 samples were generated per condition (5 seeds × 8), and 14 properties measured. Four are the conditions themselves (height, block count, symmetry, enclosure). The other ten were never requested; the paper calls them emergent properties — floor count, elongation, surface ratio, layer consistency, the share of air inside the box, and so on. Responsiveness is the largest relative shift from the unconditioned mean, with bootstrap 95% intervals (resampling the data many times to gauge spread).

The cut lines are simple: over 100% is Controllable, 20–100% Approachable, under 20% Unresponsive. Applying these three bins to all 14 properties gives the "regime map" in the title. Once you have the map, you can compare where a new request is likely to land.

Findings — nine moved a lot; only 'air inside' stayed put

Key numbers from Table 1. Nine properties shifted by more than 100% (Controllable): height 266%, block count 632%, symmetry 237%, floor count 182%, vertical aspect 333%, and others. For seven of them, the whole 95% interval stays inside Controllable. Two sit on the edge: connected components (128%) has an interval that dips below, and enclosed ratio (564%) only looks large because its baseline is near zero — a mean of 0.00003 and an absolute shift of about 0.006, per the paper.

Four were Approachable (20–100%): elongation 64%, surface ratio 32%, layer consistency 44%, footprint convexity 51%. Only one was Unresponsive (under 20%): hollowness, the share of air inside the bounding box, at 14%.

Then the prediction (Figure 1). For the ten emergent properties, the author multiplied two quantities: how strongly the property is tied to some conditioned property, and how much it varies in the training data (coefficient of variation). That composite's rank correlation with actual responsiveness was 0.879 (p=0.002, n=10). Each factor alone gave 0.624 and 0.636, neither significant. One reading: properties that are both linked to a condition and varied in the data move most easily.

Finally, the CFG scale was varied over 0, 2, and 4 (Table 2). Floor count rose 88% → 167% → 197%, crossing from Approachable to Controllable; the author treats this as a frequency floor (too few examples, but it moves if you push). Hollowness crept up 6% → 12% → 19% and stayed below 20%. Whether stronger guidance would push it over is, in the author's words, untested.

Use cases — three habits a designer can take home

First: if you are training a generator for voxel or tile-based buildings or dungeons, list the properties you want to request before training, and measure two things in the data — is each property tied to something you condition on, and does it vary? Where both are low (here, hollowness), plan from the start to fix it with rules afterward rather than rely on learning. The author also suggests layering explicit constraint methods for unresponsive properties.

Second: if you generate Sokoban-like puzzle levels with a learned model, you can sort requests such as "solvable" or "medium move count" the same way. Sweep the guidance strength over two or three settings and watch how responsiveness grows. If it grows, more data, stronger guidance, or fine-tuning may get you there. If it barely moves, invest in changing the architecture or in filtering outputs with a solver (a program that checks solvability).

Third: label each condition in your generator's spec or tool UI as strong, loose, or ineffective. Level designers learn which knobs to trust and are disappointed less often. For items that look huge in percent but tiny in absolute terms (enclosed ratio here), show the baseline value next to the percentage to avoid misreading.

Fourth: if you run a daily puzzle that generates one level a day, check first whether the property you want to request, such as difficulty, actually varies enough in your data. A dataset of mostly easy levels will struggle to give you a hard one on request. This is my application of the paper's frequency-floor idea, and hand-authoring extra levels for the thin buckets may be the cheapest fix.

Limitations — one pipeline, small data, ten points

The author admits many weaknesses. The data (10,310 builds) is small for neural 3D generation and heterogeneous. There is a single pipeline. The predictive correlation rests on 10 properties, with a wide 95% interval of 0.55–0.97, and the composite is described as a hypothesis, not a validated predictor. All four directly conditioned properties responded well, so there are no negative cases there to generalize from.

Regime calls near boundaries depend on which conditions were tested. The 40 samples per condition are clustered within 5 seeds, which may understate variance. And only structural response was measured; visual quality needs a separate evaluation. The conclusion itself says the result is pipeline-specific and should be tested on larger datasets, other architectures, and other content domains.

What I, Fukai, would add is twofold. First, there is no human evaluation. "Controllable" means a number moved, not that designers got the shapes they wanted. Second, why hollowness did not move is not yet pinned down. Its low variation in the data (CV 0.23) is reported, but whether stronger guidance would push it over is untested. I would not take "unresponsive properties need architectural change" as settled on the strength of this single case.

Fukai's reading — writing the generator's user manual

This part is my opinion. I would place this work in the line of expressive range analysis that Smith and Whitehead began in 2010 (mapping the range of outputs a generator can produce). That work charted what comes out; this one charts which requests make things move. In design-critique terms, it labels each knob of a generator with how well it works. It is less about making the model smarter and more about writing a user manual for people who use the model as a tool. An independent researcher releasing code and trained checkpoints fits that stance.

Closing — where to read next

If you want to go deeper, start with Smith and Whitehead on expressive range — the origin of the idea of measuring generators. Then read Karth and Smith on WaveFunctionCollapse to see the rule-based side, which keeps shapes by construction. When this paper says to layer explicit constraints over unresponsive properties, that is the kind of method it means.

For the conditioning mechanism itself, Ho and Salimans' classifier-free guidance paper is the source. Read the three together and you get a map of what comes out, what is kept, and how to steer. The author has archived code, results, and checkpoints on Zenodo, so you can try the same diagnostic on your own data.

References

Papers and materials referenced in this article:

・Which Structural Constraints Are Learnable? A Regime Map for a Minecraft Voxel Generator (Alex Chengyu Li, 2026, PCG Workshop at FDG '26)

・DOI: 10.1145/3815598.3815669

・Code, results and checkpoints (Zenodo): 10.5281/zenodo.20821894

・Related: Analyzing the expressive range of a level generator (Gillian Smith, Jim Whitehead, 2010, PCG Workshop 2010)

・Related: WaveFunctionCollapse is constraint solving in the wild (Isaac Karth, Adam M. Smith, 2017, FDG '17)

・Related: Classifier-Free Diffusion Guidance (Jonathan Ho, Tim Salimans, 2022, arXiv)

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 100 of 100

Read next