PAPER-DIGEST · 2026-08-19

Battleday et al.: Measuring AI Discovery With 70 Games That Never Explain Their Rules — Fukai Reads

Discovery benchmark / hidden rules / human comparison

TL;DR

Today I read a preprint about a set of 70 games built to measure one thing: when an AI is dropped into a small world whose rules and win conditions were never explained, can it act, experiment, and work out the regularities for itself? The benchmark is called DiG-bench (Discovery in Games). Each game is presented as a short string of text, has between 1 and 16 levels, and each level has a limited number of steps. The player is given only the list of available actions and the opening screen; both what the rules are and what counts as winning must be inferred through play.

The results lean further toward humans than I expected. All 70 games were beaten by at least one human on their first attempt. The strongest single model beat 50; pooling every model tested, 57 games fell in total, and in the top two tiers (20 games combined) only 9 were beaten. And when the ground-truth rules are simply handed over, the same model jumps from 18 games to 69. The bottleneck, then, reads as the finding of rules rather than the following of them.

Read from the puzzle-maker's side, the paper doubles as a level-design handbook. Confine the observation to a short string. Keep the action set under roughly ten options. Provide a separate sandbox whose steps do not consume the budget. Stratify difficulty by machine failure while separately guaranteeing a human floor. Every one of those is a design decision I can carry straight back to my own boards. Note that this is an arXiv preprint (arXiv:2608.12593, submitted 12 August 2026) and has not yet been through peer review.

Introduction: Who Wrote It, and Where

The paper is titled "DiG-bench: Discovery in Games". Its arXiv identifier is 2608.12593, submitted on 12 August 2026, with cs.AI (Artificial Intelligence) as the primary category and a cross-listing in cs.LG (Machine Learning). The author list runs to sixteen names — Ruairidh M. Battleday first, and including Kai Sandbrink, Tri Dao, Jürgen Schmidhuber, Joshua Tenenbaum, Thomas L. Griffiths and James C.R. Whittington; the full list is in the references. It is a roster drawn from both cognitive science and machine learning. The public site is at digbench.ai.

I chose it today because its central question overlaps almost exactly with the craft of making puzzles. What DiG-bench measures is the ability to act on a system whose rules are hidden and infer those rules from what happens. That is nearly a restatement of what a good puzzle game asks of its player. Baba Is You has no manual; neither does Chants of Sennaar. The player touches things, fails, and assembles the rules from the wreckage. What arrived today is a document that lines up seventy instances of that process and puts numbers on it.

A caveat up front: this is a preprint posted to arXiv, not a peer-reviewed paper (a preprint being a manuscript the authors release before any journal or conference review). It is only days old and has accumulated no citations yet, so I want to handle its numbers carefully. Every figure I quote below is quoted as the paper's text and tables state it.

Background: What Existing Benchmarks Were Missing

There is no shortage of tests for AI capability. Even so, the authors write that almost nothing in the current landscape directly probes the capacity to discover new knowledge through experimentation in an environment whose objective is unknown. Their definition of discovery is crisp: finding and making sense of previously unexplained regularities in a way that is useful for compressing existing observations and for reasoning about new states of a system. In short, being able to summarise what you have seen and predict what comes next.

The authors name the gaps specifically. ARC-AGI-3 is framed as a test of discovery and fluid reasoning, but confounds it with visual perception. ARC-AGI-1/2, Bongard-LOGO and IOLBench probe rule induction — inferring a rule from examples — but leave out active experimentation. DiscoveryWorld does test experimentation, yet entangles it with prior scientific knowledge or leans on a menu of possible discoveries. The claim, then, is that no controlled test isolated discovery itself from perception and prior knowledge.

Their instrument is games. Games are a traditional stage for AI research, but DiG-bench uses them differently. Each game here is a self-contained miniature world with its own laws, and both those laws and the objective are hidden from the player. Translated into a puzzle designer's terms, the brief is: design a board that remains solvable with the tutorial and the goal display both removed. All 70 games are handcrafted by human experts and new, existing nowhere on the internet. The worry that answers are already sitting in the training data — contamination — is cut off by building fresh and keeping most of it private.

Approach: How the 70 Games Are Built

The construction is strikingly plain. There are 70 games, assigned to seven tiers by machine difficulty, tier 1 easiest and tier 7 hardest. Three games from each tier — 21 in total, named P-1 through P-21 — are released publicly; the remaining 49 are held private to keep evaluation honest. For the public set there is an API and SDK, so anyone can run their own model against it.

A game's observation is a short Unicode string. Most games display using only letters, digits and simple punctuation. A game has between 1 and 16 levels, and each level pushes the player to infer more of the game's structure. Each level has a step limit. And the action set is small: the paper states that most games have fewer than ten possible actions in total. That strikes me as the important design point. The fewer the actions, the more surely a player who is stuck is stuck because the rule is unclear, not because the move is unfindable.

The device I found most interesting is creative mode. It is a sandbox with rules similar or identical to the main game, but where steps do not count towards the per-level limit; the player enters it with a "/" action. Crucially, the sandbox's initial state is made different from the main game's, so a solution found there cannot simply be copied across. Free experimentation, no free answers — one design move that buys both.

The evaluation setup is worth recording too. In the basic harness, the model is sent the task description and initial state at the start of a run, then returns a single legal action per turn. In very long games the history is truncated to fit the context window, dropping the oldest steps first; this happened in 6.1% of runs. Where a reasoning-effort setting was exposed it was set to high. As separate conditions, the games were run inside agentic harnesses — Claude Code, Codex, Kimi Code, Prime Agent, PRO-LONG, frameworks in which the model calls its own tools and manages its own context — and, as a control, with the ground-truth rules given before play. Humans received identical information to the agents, plus guidance to take their time and encouragement to use pen and paper.

Findings: Where the Numbers Landed

The overall shape first. The weakest model tested (Qwen 3.6 27B) beat 1 game out of 70. The strongest, Opus 5, beat 50. Pooling across every model — counting any game that some model beat — gives 57. The models run in the basic harness were Claude Opus 5, Gemini 3.1 Pro, GPT-5.5, Kimi K3, GLM-5.2, DeepSeek V4 Pro Preview, DeepSeek V4 Flash 0731 and Qwen 3.6 27B. The lowest tier is mostly beaten even by a one-generation-old model like Gemini 3.1 Pro, while the higher tiers remain hard for the state of the art: across the top two tiers, 20 games combined, the models took 9.

Humans, by contrast, beat all 70 — each by at least one person, on a first attempt. The authors do not claim this was easy: they state explicitly that many of the games are substantially effortful for humans, and that discovery often takes work. There is a step-count comparison too. Restricted to levels both humans and Gemini 3.1 Pro beat, Gemini took 46 ± 63 steps per level and humans 49 ± 78, with no statistical difference (Wilcoxon p=0.15, Figure 5A). On the levels they both solve, the effort is about the same. The gap is not in speed but in whether the thing gets solved at all.

The most eloquent number in the paper, I think, is the rules-given control. Without the rules, Gemini 3.1 Pro beats 18 of 70. Given the ground-truth rules, it beats 69 (Figure 5B). Same model, same games, same step limits, and the only added ingredient is a statement of the rules. Given the rules, that model also almost stops using creative mode (Figure 5C). The experimentation, the numbers suggest, was there because the rules were not. The other surprise was that agentic harnesses did not help: Fable 5 (Claude Code), GPT-5.6 Sol (Codex), Opus 5 (Prime Agent) and the rest all failed to improve over basic-harness Opus 5, which the authors read as suggestive that existing harnesses offer limited gains in discovery ability.

The authors also check that the public and private halves have not drifted apart. Gemini's win rates are similar between public and private games, and human step efficiency likewise (Figure 5D) — so an impression formed on the 21 public games should roughly carry to the 49 hidden ones. On run conditions: most runs used a $200 cost cap, Kimi K3 used a 12-hour wall-clock cap, and context sizes span 262k to 1M tokens. Runs stopped by hitting a cap number between 0 and 6 per configuration (Table 3).

Use Cases: What a Puzzle Maker Can Take Home

First: make the practice sandbox one you cannot copy from. Creative mode is directly transferable. Suppose you run a daily symbolic puzzle and want newcomers to try the controls without burning their real step budget. Build it naively and players will simply replay the practice solution on the live board. DiG-bench's answer is to make the sandbox's initial state different from the main game's. Freedom to experiment stays; transcription is closed off. That single move turns a practice mode into a learning device.

A board from Baba Is You, where pushing the word blocks that state the rules rewrites the rules of the game itselfBaba Is You (from official Steam screenshots)

Second: stratify difficulty by machine failure, and guarantee the human floor separately. DiG-bench assigns its seven tiers by machine difficulty, then separately confirms that all 70 games were beaten by at least one human on a first attempt. That double structure is worth copying. For your own puzzles, let your solver or a general model attack the boards and put the ones it cannot crack in the upper tiers — while making "at least one human tester solved it cold" a shipping condition. The first check alone lets unfair boards drift upward; the second alone leaves the tiers blurry.

Third: price your hints using this paper's numbers. If handing over the rules moves a model from 18 games to 69, then in a puzzle a statement of the rules is very nearly the answer itself. When designing tiered hints we tend to treat "just tell them the rule, briefly" as a gentle nudge; in this frame it is the strongest hint available. Keep the rule itself for the last tier and put attention-steering earlier — where to look, which action to try again — so the experience of discovery survives.

Fourth: confine the observation to a short string, and keep the action set small. DiG-bench excluded visual perception for reasons of AI evaluation, but the move pays off in puzzle-making too. A board expressible in letters and punctuation survives a small phone screen, works with a screen reader, and produces light share images. Above all it keeps "I cannot see it" from mixing with "I cannot understand it." Holding the action set under ten runs in the same direction: fewer controls mean the player spends less attention on what to press next and more on what is happening on this board. Cutting actions is not cutting expressiveness; it reads as consolidating the causes of being stuck into one.

Limitations: What the Authors Concede, and What I Noticed

The weaknesses the authors concede in the text first. The largest is the number of runs: most model-by-game pairs had only a single run, so the results are effectively single-seed, and variability across seeds is explicitly not reported. Whether a gap like 50 versus 57 sits inside the noise cannot be settled from this document. On top of that, very long games had their histories truncated, which happened in 6.1% of runs, and some runs were stopped by cost or time caps — between 0 and 6 per configuration. That the caps themselves vary from $100 to $200 by configuration also blurs direct comparison between models a little.

From here on, this is what Fukai points out. First, the human 70-of-70 and the model 50 are not quantities you can line up as they stand. The human 70 aggregates "someone, on a first attempt, beat each game"; the model 50 is what a single model beat. The like-for-like comparison is against the pooled model figure of 57, and 57 versus 70 is closer to the real distance. The paper's own conclusions keep these apart carefully, but lifting the numbers alone makes the human advantage look larger than it is.

Second, the number and characteristics of the human players do not appear in the text. The recruitment channels do: personal and professional networks, direct outreach to academics, professors nominating students, and puzzle-solving communities. And the authors themselves write that the players tended to be motivated by solving puzzles. This human baseline, then, belongs not to people in general but to people who like puzzles, were told to take their time, and were encouraged to use pen and paper. For puzzle-design purposes that is arguably the more useful population, but it does not support a reading of "humans in general versus AI." The data are presented in aggregate, anonymised form only.

Third, this preprint has no dedicated limitations section; the acknowledgements of weakness are scattered through the discussion. That is common enough in a pre-review manuscript, but as a reader it is inconvenient not to be able to check in one place how far the authors' own sense of their weaknesses extends. Fourth, 49 of the 70 games are private. That is the right call for preventing contamination, but it carries the cost that no outside party can audit whether the difficulty tiers are sound or the games are fair — an unavoidable trade, which I would still like written down as a trade. Fifth, the models used for evaluation move fast in both name and version; reading that table again in six months, separating which numbers were limits of a model from which were limits of a particular release will be hard.

Fukai's Reading

This section is my own reading. I want to place this work in a line of movement that takes the hiddenness of rules — until now left to a designer's instinct — and shifts it toward the side of measurable resources. In the vocabulary of puzzle design criticism, what DiG-bench did is less the automation of difficulty than the calibration of unknownness. If the same board swings between 18 and 69 solved depending on whether the rules are concealed or stated, then what you conceal is not a component of difficulty but its principal component. Perhaps we have talked about difficulty in terms of board complexity and step counts only because we had no instrument for measuring the amount of withheld information. I take these 70 games as a first edition of that instrument.

Closing

For readers who want to go deeper, I would start by skimming the prior benchmarks the paper names — the ARC-AGI family, Bongard-LOGO, IOLBench, DiscoveryWorld. Seeing what DiG-bench subtracted in order to exist sharpens the outline of these 70 games considerably. Among articles on this site, the piece on Ahn and colleagues' CogARC covers AI inferring rules; the piece on Chao and colleagues covers the relation between insight and exploration; and the piece on Waugh and colleagues covers evaluating machines on pencil puzzles.

My own homework is to play the 21 public games by hand. Start at tier 1, work upward, and find out where my own instincts stall. Pen and paper are encouraged, and I intend to comply.

References

Papers and materials referenced in this article:

・DiG-bench: Discovery in Games (Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, Jürgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas L. Griffiths, Zeb Kurth-Nelson, James C.R. Whittington, 2026, arXiv preprint arXiv:2608.12593)

・Full HTML text of the paper (including Tables 1-3 and Figures 1, 4 and 5)

・Official DiG-bench site (play the 21 public games, leaderboard, API/SDK)

・discos-research/dig-bench (public implementation repository)

・Prior benchmarks the paper positions itself against: ARC-AGI-3 / ARC-AGI-1 and 2 / Bongard-LOGO / IOLBench / DiscoveryWorld

・Related articles on this site: on Ahn et al.'s CogARC / on Chao et al. on insight and exploration / on Waugh et al. on pencil-puzzle evaluation

・Review status: an arXiv preprint (submitted 12 August 2026; primary cs.AI, cross-listed cs.LG) with no DOI assigned. Posted only days ago, it has accumulated no citations and has not yet been widely discussed

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 63 of 101

Read next

Related reviews