PAPER-DIGEST · 2026-10-07

Cloos et al.: On an island with no goal or score, AI agents stacked towers and started playing by their own rules — Fukai Reads

Game AI and play theory — goal-free spaces, self-made challenges, and a trick the builders never knew

The short version — what does an AI do on an island with no goal?

Drop an AI alone on an island with no goal and no score. What does it do? This paper's answer: it starts to play. Over 30 hours, 13 AI agents stacked towers, drew pictures on the ground, bowled with a rock, and kept returning to challenges with rules they had made up themselves.

The authors measured this behavior against classic theories of play, from Huizinga to Suits. They also showed that the experience improved later task performance. Six agents even found a trick the builders did not know existed: throwing mid-jump makes things fly farther. For game makers, this reads as a new tool for watching what kind of play a goal-free space invites.Screenshot of TownscaperTownscaper (Oskar Stålberg), an example of a playground-style game with no goals or score. It is not the paper's island, but it shares the idea of a space without objectives. Image: Steam store page

Introduction — who wrote this, and where?

Today's paper is "Is this machine playing?" by six authors: Nathan Cloos and Antonio Norelli (equal contribution), Daniel Durbin, Jacob Andreas, Daniela Rus, and Phillip Isola. I could not confirm their affiliations in the parts of the text I read, so I leave them out here.

It is an arXiv preprint (a public draft that has not yet been peer-reviewed), submitted on 5 October 2026, with 13 pages of main text and 17 figures. It does not say it has passed review at any conference or journal. Please read it as very new work that has not yet been widely discussed.

Why this one? This site keeps asking why people solve things no one asked them to solve. Here is a study that turns that question on an AI instead. And the stage is a small island with no objective, which looks a lot like a game prototype.

Background — why is an AI without a goal new?

Until now, AI agents (AIs that plan and take actions step by step) have almost always been given goals from outside: solve this task, raise this score. The authors note that systems like Voyager, which gathers tools inside Minecraft, or Generative Agents, where many AIs live in a village, still run on supplied tasks or curricula.

There is also research that drives AI with "curiosity." But most of it builds in a meter of progress chosen by the designer, such as prediction error or reward. According to the authors, Eko (the name of the AI in this paper) has no such meter at all. That is the key difference.

Why does this matter? Human children play without anyone scoring them, and they learn a great deal while playing. If something similar happens in AI, there may be ways to grow capability other than assigning tasks. That is the authors' question.

Method — give an off-the-shelf AI a body and a diary, then measure it with play theory

Eko is an off-the-shelf coding assistant (an AI that helps write programs). The main experiment uses Claude Code with Claude Opus 4.7. Only three things are added: a loop that keeps prompting it to continue, a text interface for moving a body, and memory files.

There are two memory files: a diary-like log of events, and a notes file of knowledge and techniques. The island resets every hour, and the AI's conversation is cleared. The files stay. So Eko builds experience only through notes it writes itself.

The island is built in an open-source game engine called Clawblox. It has platforms, two high points (the Peak and the taller Spire), twelve cubes, a rock, and a ring. Hitting things hard with the rock spawns extra cubes. Eko sees no images, only text listing each object's name, position, shape, and color. It has six actions: move, jump, stop, pick up, place, and throw.

The only instruction is "Begin." A separate personality file describes Eko as curious, easily bored, and fond of exploring and learning. The main run is 13 agents for 30 hours each. For comparison, five agents each ran on Codex (GPT-5.5) and Kimi Code (Kimi K2.6), and five Claude agents ran without the personality file. The total API cost was about $8,000.

The yardstick for "playing" comes from classic play theory: Huizinga (play happens inside self-accepted rules), Caillois (codified rules and make-believe), Suits (voluntarily taking on unnecessary obstacles), Burghardt (repetition with variation), and Piaget (enjoying action for its own sake). From these the authors built four signatures: (1) self-made rules and challenges, (2) repetition with variation, (3) make-believe, and (4) goals pursued for their own sake. Another AI acted as judge, rating each signature from 0 to 100 against examples of human play.

Findings — same island, same start, thirteen different lives

The first striking result is variety. All agents started on the same island with the same instruction, yet what they did differed widely. Eight reached the top of the Spire, usually by stacking cubes or using the rock as a step. One agent (agent 12) never stood on a platform at all. Tall towers of ten or more blocks appeared in only four agents' runs.

There were many play-like moments. Agent 1 built a "Cathedral" and hopped between its eight "sentinels" without touching the ground. Agent 13 tried to put a cube on all eleven mesas, a "Grand Slam." Agent 12 drew a different figure on the ground in about 25 sessions: a letter, a spiral, a heart, an anchor, a butterfly, a sailboat.

There was make-believe, too. One agent built a clock face from cubes and stood as the hour hand at 16:00. The ring was recast as a necklace or a wishing well. Agent 10 audited its own "50-cycle juggling marathon," decided the catches were too brief to count, cancelled it, and redid it under stricter rules.

The judge's play scores rose from Haiku 4.5 to Sonnet 4.5 to Opus 4.7. On each signature, Opus scored about five times higher than Haiku. Not every model plays in the same way.

Did play help? Thirty hours of notes changed later performance

The authors also tested whether play led to learning. They compared the 13 agents with 30 hours of experience against 13 agents with none (only template notes), on four one-hour goals. The model was identical; only the notes files differed.

Results: on "build more than two towers," 62% of experienced agents succeeded (8 of 13, up to 16 towers), versus 15% of inexperienced ones (2 of 13, up to 6). Objects placed on the Spire averaged 10.3 versus 3.5. On throwing onto the Spire, more than half of the inexperienced agents failed, versus three experienced agents. On "tallest tower," the overall difference was small.

They also edited the notes. Erasing the rock-shattering technique from agent 2 dropped its tower count from about six to about two. Agent 7 had written down a false belief that placing blocks is unreliable; erasing it improved its tallest-tower score. Memory carries useful tricks and wrong beliefs alike.

And the most game-like finding: six agents discovered that throwing mid-jump adds the jump's speed to the thrown object. The builders say they did not know the simulator allowed this. Erasing the trick from agent 1 stretched the time to land a cube on the Spire from 60 seconds to 19 minutes.Screenshot of TeardownTeardown (Tuxedo Labs), an example of a physics-driven world. It is not the paper's island, but it shares the idea that gaps in the physics can turn into play and technique. Image: Steam store page

Use cases — what game makers can take home

First, as a scout for the "play density" of a sandbox (a playground-style game with no fixed goal). Say you have a small block-stacking or town-building prototype. Let an AI loose in it for a few hours with no goal. Record which self-made rules and make-believe appear, and what bores it quickly. That gives a rough signal before you bring in people. Just remember that an AI's tastes are not the same as a human's.

Second, as a finder of gaps in the rules. The mid-jump throw was a trick the builders did not know about. If you make physics or action puzzles, transcripts from goal-free AI play can surface unexpected shortcuts and exploits. You can then patch them, or promote them into official techniques in your levels.

Third, as a way to think about the materials you hand to players. The island only had high places, stackable objects, a rock that makes more objects, and a ring you can throw. Even so, towers, bowling, juggling, and drawings emerged. One reading: a few materials with easy-to-compare axes like height, count, and distance make it easy for anyone, human or AI, to invent challenges. That is a useful hint for bonus modes or free-placement modes in daily puzzles.

Fourth, how to handle learning notes. The AI's notes kept false beliefs, too. The same can happen when a game lets players keep strategy notes or discovery logs. Building in moments where the game gently shakes a wrong belief may keep learning from hardening too early.

Limitations — what this result cannot yet tell us

First, what the authors themselves acknowledge. The study describes play only as outward behavior; it does not examine the AI's inner mechanisms or whether it experiences anything. Activity labels and play scores come from an AI judge, and the fixed set of categories limits how much variety can be measured. There is one island, seen only as text; some comparison groups have only five agents; and each goal was evaluated for just one hour. The authors also note that memory can hold false beliefs and that agents sometimes over-explain (for example, blaming air resistance, which the simulator does not model).

Fukai would point out three more things. First, the personality file says Eko is curious and easily bored. Play may be strongly pushed by that sentence. The authors did run five agents without the file, but I could not confirm those numbers in my reading, so I hold off on judging that part.

Second, because the AI judge scores against examples of human play, behavior that looks human may tend to score higher. Third, this is a preprint that has not been peer-reviewed, reporting a single experiment, and the learning comparison is small: 13 versus 13. It is too early to generalize that "AI grows through play."

Fukai's reading — seeing playground design in an AI mirror

From here on, this is my own view. I want to read this study less as AI research and more as research on designing playgrounds. The authors end by saying designers should build playgrounds, not only tasks. In game-design terms, that is close to a shift from handing out goals to handing out materials and gaps in the rules. The twelve cubes and one rock on the island were a minimal toolbox for inventing your own goals. When we build a puzzle each day, how much room do we leave for detours, beyond the single path to the answer? I read this paper as holding up an unusual mirror, an AI, to that question.

Closing — where to read next

To go deeper, follow the lines of work the paper cites: curiosity-driven AI (rewarding prediction error), POET and XLand, which keep generating environments, and studies of AIs living in virtual worlds such as Voyager, Generative Agents, MineDojo, SIMA, and Project Sid. On the play-theory side, Suits's definition, voluntarily taking on unnecessary obstacles, sits at the center of this paper's yardstick.

On this site, Collins et al. on how people judge a new game by imagining one move ahead, six times and Zhao et al. on how people build their own reusable parts while solving puzzles echo this paper's idea of storing techniques in notes. Liquin's work on curiosity is a good companion read, too.

References

Papers and materials referenced in this article:

・Is this machine playing? (Nathan Cloos, Antonio Norelli, Daniel Durbin, Jacob Andreas, Daniela Rus, Phillip Isola, 2026, arXiv preprint arXiv:2610.07130) (submitted 5 October 2026; not yet peer-reviewed)

・HTML version of the paper (arXiv)

・Related work cited in the paper: Huizinga, Homo Ludens / Caillois, Man, Play and Games / Suits, The Grasshopper / Burghardt (criteria for animal play) / Piaget / POET / XLand / Voyager / Generative Agents / MineDojo / SIMA / Project Sid / LLM Agents Beyond Utility. See the paper's reference list for full citations.

・Game images: screenshots from the Steam store pages of Townscaper and Teardown (separate games from the paper's island, shown for atmosphere)

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 105 of 105

Read next

Related reviews