PAPER-DIGEST · 2026-10-01

Lessard et al.: Can a Machine Dig the Good Stories Out of a Simulation? — Fukai Reads

Emergent narrative — an experiment in story sifting, and its unexpected ending

TL;DR

This paper honestly reports an attempt to have a machine pick only the 'interesting stories' out of a pile of events a simulation produced on its own. The conclusion was a surprise. The story 24 readers liked best was the 'control' one, chosen with no clever selection at all.

The authors tried two yardsticks for picking stories out of 228K events: 'improbability' and 'dramatic situations.' The six stories they picked were each turned into a short story twice, once by a human writer and once by GPT-4. Readers rated those 12 pieces on a 5-point scale.

The yardsticks made no clear difference. What emerged instead were two forces: how good the raw material is (a 'story effect') and how well it is told (a 'narrative effect'). The authors end by asking whether such stories can only be enjoyed by people who already know the world they came from.

Who wrote it, and where

The authors are Jonathan Lessard, Stephen Friedrich, Samuel Paré-Chouinard and Leila Kosseim, all at Concordia University in Montreal, Canada. The simulation used in the experiment was built by LabLabLab, a research lab there.

It was presented at the PCG Workshop held alongside FDG '26 (Foundations of Digital Games), August 10–13, 2026, in Copenhagen. PCG (Procedural Content Generation) is a small venue, and the paper is short, as workshop papers tend to be. Its DOI is 10.1145/3815598.3815687, and it is published under CC BY 4.0.

I picked it today for a simple reason. Papers that say 'this did not work' this candidly are rare. And by digging into why it failed, the authors surface a question that matters to anyone who makes games. The paper is also very new, so it has not yet been widely discussed.

Sifting sand for gold

Games like Dwarf Fortress and RimWorld have no fixed plot. Inhabitants act by their own rules, and 'stories' fall out of the result. This is called emergent narrative (stories no author wrote, arising naturally from how a system behaves).Dwarf Fortress gameplay screenDwarf Fortress (Steam store screenshot). The paper cites its player retellings as a precedent

The catch is that most events are dull. Someone cut a tree; someone ate a meal. Buried among them, now and then, is a chain of events worth telling. In his 2018 PhD thesis, researcher James Ryan framed this as 'overgenerate and curate.' Finding stories after the fact is called story sifting (running a sieve over the pile of events to pick out chains that could become stories).

The paper likens this to 'the gold prospector sieving large quantities of sand.' That is where the title's 'gold digging' comes from.

There is prior work. In 2022, Kreminski and colleagues proposed ranking stories by 'unexpectedness.' But most studies have checked whether the algorithm works or whether the tool is easy to use. Few have measured whether readers actually find the results interesting. That is the gap this paper aims at.

How they looked for interesting stories

The testbed is Chroniqueur, a social simulation LabLabLab built from 2020 to 2024. Groups of hunter-gatherers gradually fill up a procedurally generated island. As resources run short, farming, herding, social ranks and towns emerge. In effect, it runs the whole Neolithic transition.

Chroniqueur records every event, who was involved, and what caused what (what Ryan calls 'causal bookkeeping'). In the world used for the experiment, 'Preparipate,' the simulation ran for 61 in-game years, producing 228K unique events across 27K event chains.

The authors first tried to use narratologist Schmid's five criteria of 'eventfulness' as yardsticks: relevance, improbability, irreversibility, non-iterativity, and lasting effect. Most proved hard to implement. 'Lasting effect,' for example, kept ranking the same kinds of events at the top no matter how it was counted.

What survived was 'improbability.' They treated the whole flow of events as chains of probabilities and looked for the least likely sequences. (The paper only says it models the event space as Markov chains, a model that describes a flow by how likely each next step is; it gives no further detail.)

The other yardstick was 'dramatic situations.' Of the 36 dramatic situations Georges Polti listed in 1912, ten could be searched for in this world, such as 'supplication,' 'crime pursued by vengeance,' 'vengeance taken for kin upon kin,' and 'an enemy loved.' The authors also built a dedicated interface to search and sort stories by these criteria.

How readers judged the stories

Six stories were chosen: the most improbable 'crime pursued by vengeance' and a random one; the most improbable 'supplication' and a random one; and the most improbable story with no dramatic situation and a random one. That last one, random with no dramatic situation, serves as the control that uses no yardstick at all.

Here the authors made a careful move. If readers saw the simulation's template text directly, clumsy prose would sway the ratings. So each story was written up as a roughly 500-word short story twice: once by the same single human writer, once by GPT-4. GPT-4 was told to stick to the outline and write 'like mainstream pop fiction or pop history.' The aim was to separate the quality of the material from the quality of the telling.

There were 24 readers, faculty and students from Fine Arts, History and English at Concordia, split into two groups of 12. Each group read one version of each of the six stories, arranged so that three were human-written and three GPT-written. They rated six items on a 5-point scale: interesting, made sense, enjoyed reading, well-rounded characters, surprising, and easy to follow.

What they found

The headline result ran against the hypothesis. The paper writes: 'Ironically, the narrative that stood out as the clear favorite was the control one.' It is a story called 'Gossip,' neither particularly improbable nor featuring any dramatic situation. The stories picked with the yardsticks could not beat the one picked without them.

Comparing human and GPT versions, readers generally preferred the human ones. According to the paper, 'only two GPT4-written stories surpassed human ones in average enjoyed reading,' and only among the less-liked stories (Figure 4, average enjoyment per story).

The power of telling showed most clearly in the 'Meteor' story, highly improbable but with no dramatic situation. Its human version ranked 3rd of 12, while its GPT-4 version ranked 12th, dead last. The same chain of events can be judged very differently depending on how it is told. The authors call this the 'narrative effect.'

There was also a 'story effect.' The gossip and love stories ranked high in both human and GPT versions. In the paper's words, 'good story material might shine in spite of mediocre narrativization.' The theft and vengeance stories, by contrast, ranked low in both versions, which reads as weak material.

Note that the paper reports no specific mean values and no statistical tests (calculations that check whether a difference could be chance). The results are presented mainly as the rankings in Figure 4 and the authors' discussion. Some readers also commented that 'affairs' came up again and again.

How game makers can use this

First: if you are making a colony or village simulation, this bears directly on a 'this week's events' feature. The paper's point is that picking statistically rare events is not enough. It reads as better to prioritize events involving inhabitants the player knows by name and places they remember.RimWorld gameplay screenRimWorld (Steam store screenshot). A signature emergent-narrative game where stories arise from colonists' actions

Second: the lesson carries over to picking a 'puzzle of the day' from generated boards. The board with the rarest solution shape is not necessarily the most fun. The paper's 'improbability is not interestingness' is a warning that applies to puzzle selection as-is. Narrow candidates by rarity if you like, but let real players confirm the final pick.

Third: if you plan to have an LLM (large language model, an AI that learns from vast amounts of text and generates text) narrate events, the 'Meteor' result is sobering. Weak telling sank the same material to last place. One option is to prepare human-written templates for the key moments and let the LLM fill in details.

Fourth: if you use story patterns like dramatic situations as templates, watch out for repetition. The authors warn that without enough variation in detail and background, each pattern becomes 'cookie-cutter.' Readers noticing that 'affairs' kept recurring is an early sign. Giving the same pattern different contexts seems more effective than adding more patterns.

Fifth: there are hints for result sharing and replay sharing. The authors conclude that emergent stories shine brightest when told among people who know the same world. That suggests building paths that deliver a story to others who played the same game, rather than explaining it to outsiders.

How far to trust it

First, the weaknesses the authors themselves admit. Few kinds of stories were read, and 24 readers is a small sample. So, they write, 'it would be premature to discredit the approach altogether.' There was only one control story, and it was chosen at random. They also note that Chroniqueur's own quirks may have mattered, and that the yardsticks could have been modeled in many other ways.

The authors also flag problems with improbability itself. Being statistically rare is not the same as being meaningful as a story. Long runs produce extreme outliers that defy narrative explanation. And because they standardized the background information given for every story, the causal chains became quite abstract.

What Fukai points out here are three things. First, there are no statistical tests; even 'clear favorite' can only be read off a ranking chart. Second, there was only one human writer. The strength of the human versions may reflect that one person's skill, not 'humans' in general.

Third, the readers were humanities faculty and students, not players of the game. This actually connects to the authors' conclusion: once readers did not know the world, the appeal of 'improbability' may have been hard to convey. It is also worth remembering, when reading the telling comparison, that GPT-4 is a model a generation behind.

Fukai's reading

This part alone is my opinion. I want to read this paper as a rediscovery of something game designers know well: interest lives not inside the material but in the relationship between the material and the reader. In puzzles, too, only someone who has learned the rules can see the surprise in a move. As the authors say at the end, citing Steven Sych, Dwarf Fortress retellings are fun because only people who know that world's 'normal' can notice its 'weirdness.' So the next task for story sifting, as I frame it, is not only choosing stories. It is also a design problem: how to hand readers a sense of the world's 'normal' in advance.

What to read next

This paper's starting point is James Ryan's PhD thesis, 'Curating Simulated Storyworlds' (2018, UC Santa Cruz). It argues head-on for the 'overgenerate and curate' view of emergent narrative. It is long, but it sits at the center of this field's map.

For prior work using 'unexpectedness' as a yardstick, see Kreminski and colleagues' 'Select the Unexpected' (ICIDS 2022). Read side by side with this paper, which reports that the direction did not pan out, it should become clearer where the two designs differ.

Sources

Papers and materials referenced in this article:

・Narrative Gold Digging: Sifting and Narrating Stories from a Procedural Simulation (Jonathan Lessard, Stephen Friedrich, Samuel Paré-Chouinard, Leila Kosseim, 2026, PCG Workshop at FDG '26)

・DOI: 10.1145/3815598.3815687

・Related: Curating Simulated Storyworlds (James Ryan, 2018, PhD thesis, UC Santa Cruz)

・Related: Select the Unexpected: A Statistical Heuristic for Story Sifting (Max Kreminski et al., 2022, ICIDS 2022)

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 98 of 98

Read next