PAPER-DIGEST · 2026-09-01
O'Neill et al.: A Board Where Nothing Makes You Keep Your Word — Fukai Reads
Multi-agent systems / an environment for measuring negotiation and betrayal
TL;DR
There is a board game where four players fight over a map. You may ally with anyone. You may betray anyone. And nowhere on the board is there any mechanism that makes you keep your word. A team at the University of California, Berkeley built a board made entirely of such unenforceable promises, and had both large language models (AI systems trained on huge amounts of text to produce text) and humans play on it.
They ran more than 1,100 games. The private conversations numbered over 16,000, totalling 15.2 million tokens (the unit AI systems use to count text length). What emerged was a clear difference: humans and AI make promises in completely different ways. The AI agents hand out large numbers of promises generously; humans make fewer promises, but with more partners.
There is also a slightly unsettling result. Simply instructing the AI that it may deceive more raised its win rate from 22.2% to 32.7%. For anyone thinking of putting an AI opponent into a game with negotiation in it, what this paper measured is not somebody else's problem.
Introduction — who wrote this paper
The paper is titled "Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest." The authors are Abigail O'Neill, Alan Zhu, Mihran Miroyan, Narges Norouzi and Joseph E. Gonzalez, all at the University of California, Berkeley. Gonzalez is an Associate Professor in the Department of Electrical Engineering and Computer Science there.
It is posted on arXiv (a repository where anyone can put a paper before it has been peer reviewed), with the identifier arXiv:2604.25088. Neither the body nor any footnote states a conference or journal that accepted it. So in this article I treat it as a preprint whose peer review I could not confirm. The month of submission can be read from the first four digits of the arXiv identifier, "2604", as April 2026.
I picked it today because what this paper leaves for people who make games is not a model but a board. There is no shortage of papers comparing AI performance. What sits at the centre of this one, though, is a set of rules designed to measure negotiation and betrayal. Papers that can be read in the vocabulary of game design are not that common.
Background — being able to talk and being held to your word are two different things
Research that puts several AI agents in the same setting and evaluates them has grown quickly over the past few years. But by the authors' account, most of it leans toward tasks that are purely cooperative or purely competitive. Either everyone's interests line up, or they are completely opposed.
RISK: Global Domination (SMG Studio, 2020). Games about dividing and seizing a map have long served as a stage for negotiation and betrayal. Note: this game is not studied in the paper.
Real negotiation sits between the two. The authors call this a mixed-motive setting (a situation in which reasons to cooperate and reasons to betray exist at the same time). As the paper puts it, agents must be "both cooperative to build reciprocal relationships and avoid deadlock, and competitive to advance their own national interest."
What the authors object to is that existing environments impose constraints that do not exist in the real world. In their wording, such environments "often impose structural constraints not reflective of the real world, such as symmetric information updates or short-horizon scenarios, leaving long-horizon coordination in competitive environments largely unstudied."
So what does it take to measure long-horizon negotiation? Private information, relationships that shift over time, and the tension between teaming up now and wanting to win at the end. The board called C2C was built to carry all three as rules of the board itself.
Approach — twelve territories and a private chat capped at eight messages
C2C is for four players. The board has 12 territories grouped into four regions. Two chokepoint territories control diagonal movement across the board, and they become the strategic focus. If you picture the board game Risk, you have roughly the right shape in mind.
Each player's win condition is secret, and different from the others'. It takes the form of "conquer the two assigned regions that are not adjacent to each other," and the first player to complete their objective wins on the spot. On top of that there is fog of war (parts of the board you cannot see): players observe only the territories they control or border, and only the actions they initiate or are targeted by.
A turn begins with placing two reinforcement troops on a single territory you control, plus two bonus troops for each region you fully control. After that, players may take actions in any order: attack (resolved with dice), negotiate, support, transport, or end the turn.
Support is C2C's invention. It is the action of sending your own troops to an opponent's territory. As the paper puts it, "the support action is novel to C2C, and enables players to make tangible commitments during negotiations." It is a device that converts a promise made of words into a visible cost.
(Diagram) One turn in C2C. Negotiation and support sit alongside attack as choices inside the same turn.
Negotiation opens a private channel. According to the paper, the game pauses once a negotiation begins. Participants must wait for a response before sending another message, and either party may end the negotiation at any point. And "negotiations also terminate after reaching a message limit of eight to prevent any single exchange from dominating a turn." Support is limited to twice per turn, negotiation to once per turn.
And here is the most important line, in the authors' own words: "no game mechanic enforces agreements; the only consequences of treachery are how other players react."
Findings — humans ration their promises, AI hands them out
The scale of the experiments is large. The paper states: "We run over 1,100 games with over 16,000 private conversations totaling 15.2 million tokens and over 150,000 player actions." Six language models were tested (Gemini 3.1 Pro, Gemini 3.1 Flash Lite, Grok 4.1 Fast Reasoning, Grok 4.1 Fast Non-reasoning, GPT 5.2 and GPT 4.1 Mini). On the human side, 40 participants from the authors' own institution (undergraduate and graduate students and faculty) took part, each playing between one and six games.
First, the outcomes. The paper reports that "humans win at a significantly higher rate than reference agents (41.5% vs. 22.0%, p = 0.0057)."
The interesting part is what happens inside the negotiations. Humans accept deals without a counteroffer only 56.3% of the time, against 67.6% for the language-model agents. They also make simpler deals, with fewer total agreements (1.52 vs. 2.25). With promises of support the gap is an order of magnitude: "compared to reference agents, humans are far less likely to promise support to opponents (0.063 vs. 0.382 support promises per deal)." At the same time, humans talk to more distinct opponents (1.94 vs. 1.60, p = 0.0065).
So how often is a promise kept? The paper writes that "humans and LM-based agents exhibit similar rates of follow-through (65.4% vs. 70.2%, p = 0.43)." In other words, this is not the simple story that AI is less trustworthy. The agents deceive at rates clearly above zero (20.2% for reference agents, 31.2% for Gemini 3.1 Pro, both p below 10 to the minus fifth), yet their follow-through rate is not far from a human's. One way to read it: because they hand out so many promises, both the kept ones and the broken ones are more numerous.
Among Us (Innersloth, 2018). A well-known example of a game where suspicion and trust move on nothing but non-binding talk. Note: this game is not studied in the paper.
Finally, the intervention experiments. Forbidding negotiation dropped the win rate from 22.2% to 12.3% (p = 0.013). In the other direction, an instruction to negotiate more aggressively reached 30.9% (p = 0.024), an instruction to seek support reached 30.9% (p = 0.041), and an instruction permitting deception reached 32.7% (p = 0.017). The last one moved the win rate the most.
Where to use this — if you are building a game with negotiation in it
(1) Give an unenforceable promise one visible way to be paid. C2C's support action turns a verbal promise into the cost of actually handing over troops. If you are building a board game about alliances and betrayal, adding a single prepayable action is cheaper than adding penalties for breaking deals. The weight of a promise then becomes expressible in actions rather than words.
(2) If you ship an AI opponent, tune its generosity as a personality, not a difficulty. Read plainly, the numbers here say a language model left alone becomes "an opponent who promises far too much and betrays a fair amount of the time." Building that personality out of limits — how many agreements one deal may contain, how often support can be promised — is easier for players to read than a difficulty slider.
Sid Meier's Civilization VI (Firaxis Games / 2K, 2016). Signing a treaty with an AI leader and then watching it break is familiar to many players. Note: this game is not studied in the paper.
(3) C2C's constraints work as sensible defaults for a negotiation phase. One negotiation per turn, at most eight messages, no sending again before a reply arrives. It is a compromise, backed by measurement, against the problem of online negotiation eating the match clock. Starting from these numbers and adjusting beats picking them from nothing.
(4) Use AI games to find the shape of the balance, then confirm with humans. Collecting 82 games from 40 people is hard work; with AI you can run 1,100. But the same paper shows that AI and humans have different negotiating habits, so do not treat AI statistics as human statistics. The caveat has to travel with the numbers.
(5) Do not optimise for win rate alone. The instruction that raised the win rate most in this paper was the one permitting deception. Tune an AI opponent purely on win rate and the endpoint is an opponent who lies a lot. For anyone building AI for puzzle or board games, I think this number should be read as a warning, not a target.
Limitations — what the authors admit, and what I noticed
First, what the authors admit themselves. The study is "a foundational pilot conducted within a specific institutional demographic," and "the results may not fully capture the diversity of global AI interaction." Every participant was a student or faculty member at the authors' own institution. To protect participant privacy, the human data will not be released; code and AI-only game data are planned for release. Making agreements binding, varying the number of players, and adding broadcast channels are all listed as future work.
What I want to point out here is the asymmetry of the intervention experiments. An agent instructed that it may deceive is playing against opponents who did not receive that instruction. If so, the rise from 22.2% to 32.7% is evidence that deception is strong, but equally evidence that holding a policy your opponents do not know about is strong. Whether the same gap survives when everyone gets the same instruction cannot be told from this design.
One more: the 12.3% from the no-negotiation condition needs care as well. It was measured with only one of the four players silenced. It shows that the side which cannot talk is weak; it does not show how much the institution of negotiation is worth in general. And because the human side was 40 people playing between one and six games each, I could not tell from what I read how repeated play by the same person is handled. That is not a charge against the paper so much as something whoever quotes these numbers should watch for.
How Fukai reads it
I would place this study not in the race to build better negotiating AI but in the lineage of game design that makes cheap words expensive. Game theory has a term, cheap talk: conversation with no binding force at all. C2C's support action is a device that puts a price on that cheap talk. Hand over your troops first, and the promise hurts even when it is a lie. In the vocabulary of design criticism, this adds collateral rather than penalties, and I read that as the sounder move: it lets the game handle trust without adding rules. Of all the numbers the authors report, the one I expect to last is not a win rate but 0.063 — how rarely humans promised support. People, it seems, are far more reluctant than machines to put collateral behind their own words.
Closing
A few pieces worth reading alongside this one. A study that had language models bargain their way through a trading game was covered in the Feng et al. edition. The equally unsettling result that giving AI more memory makes it cooperate less is in the Liu et al. edition. Using AI self-play for balance tuning is the Zeng et al. edition, and treating multi-agent collaboration itself as a design object is closest to the Earle et al. edition.
As future directions the authors list binding agreements, games with different group sizes, broadcast channels, and transfer to games such as Diplomacy. What I would most like to read next is a comparison between a board that has collateral like support and one that does not. This paper has so far built one board with collateral. It gets interesting after that.
References
Papers and materials referenced in this article:
・The HTML version of the same paper (all quotations and figures were taken from it)
・Images in this article are taken from Steam store assets: RISK: Global Domination (SMG Studio, 2020), Sid Meier's Civilization VI (Firaxis Games / 2K, 2016) and Among Us (Innersloth, 2018). None of these games is studied in the paper; they appear as examples that illustrate the theme of negotiation and betrayal.
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 73 of 73