PAPER-DIGEST · 2026-08-21

Kelidari et al.: A Card-Game Agent Is Only as Strong as the Yardstick You Build First — Fukai Reads

Imperfect information / building a yardstick / over a hundred ablations

TL;DR

What I read today is a preprint that takes Gin Rummy — a two-player card game — and checks, one ingredient at a time across more than a hundred controlled runs, which training choices actually make a lightweight reinforcement learning agent (a system that learns by trial and error to take higher-reward actions; here small enough to train on a single GPU) stronger. The authors start from two problems any practitioner eventually hits: an agent never outgrows the opponents it practises against, and there is no cheap way to grade it.

Their answer is unglamorous and effective. First, hand-build a rule-based "fixed expert" and use it as an immovable yardstick. Measured as win-rate against that expert, plain PPO (one standard reinforcement learning method) reaches 15.0%, TRPO (a method that keeps each update inside a trusted region) reaches 22.5%, and stacking every ingredient that worked reaches 34.2±2.1%. Conversely, learned feature embeddings, imitation learning, fine-grained step rewards and using an LLM as the sparring partner all failed to help.

The finding most useful to designers is this. Swapping the network shape — fully connected, convolutional, permutation-invariant set, recurrent — leaves win-rate inside one narrow band with overlapping confidence intervals, whereas a search that can peek at the hidden cards and one that cannot split into 85% and 26%. The ceiling, then, reads as set by how much information is visible rather than by model size. Note that this is an arXiv preprint (arXiv:2607.06854), submitted to AIIDE 2026 and therefore not yet peer-reviewed.

Introduction

The paper is titled "A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong", by Nima Kelidari, Mohammadsaeed Haghi and Mahdi Salmani. A footnote in the body states: "Preprint. Submitted to the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE 2026)." So this is a manuscript submitted to AIIDE, a venue for AI in games and interactive entertainment — not a peer-reviewed publication. Its arXiv ID is 2607.06854, posted in July 2026.

I picked it today because its subject is small. This is not a story about a giant self-play system; it is a study of agents that fit on one GPU, taking each training trick out in turn to see whether it mattered. Removing components one at a time to measure their contribution is called an ablation study, and this paper stacks more than a hundred such runs into tables. For an individual or a small team putting AI into a game, that is exactly the grain size that helps.

There is a second reason. The core claim is not about a learning algorithm but about measurement. In the authors' words, an agent beats a random opponent over 99 percent of the time and only ties copies of itself, so neither tells you how strong it is. That absence of a yardstick is not unique to AI — it shows up in the same shape when you try to grade puzzle difficulty. So today I read a card-game paper as a level-design paper.

Background

The rules, kept to a minimum. Gin Rummy is a two-player game with a 52-card deck: you form runs and sets from your hand (melds) and try to reduce the points left over in unmatched cards (deadwood). Once deadwood is low enough you declare a "knock" and settle the hand. Settling with zero deadwood is "gin", worth the most points. You cannot see your opponent's hand, which makes this an imperfect-information game (a game in which some of the state is hidden from you).

Plenty is already known in this area. In fully observable games like Go and chess, the AlphaGo Zero and AlphaZero line showed that self-play (repeatedly playing copies of yourself to improve) can surpass humans. In imperfect-information games like poker, the CFR (counterfactual regret minimization — an iterative procedure that minimises regret) family is the established route toward equilibrium. Silver et al. and Zinkevich et al. both appear in the paper's references.

What was missing is the practice in between. You lack the compute for equilibrium solving, yet win-rate against random play tells you nothing. Records that compare, on one shared yardstick, which tricks are worth paying for when you build small were thin. The authors call this the "opponent bottleneck": as the paper puts it, "Training against a weak fixed opponent teaches a weak ceiling, and pure self-play can chase its own tail."

Approach

The first thing the authors build is an opponent that does not learn: a rule-written "fixed expert" with four steps. (1) Enumerate every way to split the hand into melds and take the split that minimises deadwood exactly. (2) Draw from the stock only when it strictly lowers deadwood. (3) Discard whichever card minimises the resulting deadwood. (4) Knock the moment it is legal, and go for gin only when gin is reachable without delay. Because the enumeration is exact rather than approximate, this opponent returns the same judgement every time — which is what makes it an immovable yardstick.

What I find striking is that the expert itself almost never gins. The paper states that "the expert gins in 0.7 to 1.7 percent of games and wins by knocking early with low deadwood." The highest-scoring move is the one the strongest player declines to use. That sentence pays off later, when we get to reward design.

The agent's input is a 4×52 grid of zeros and ones (own hand, discards, cards whose location is still unknown, laid out as a table), and its action space is 110 discrete choices with legal-action masking (a mechanism that prevents illegal moves from being selected). The headline algorithm comparison runs for 2 million steps; the strongest self-play configuration for 3 million. The compute is a single GPU.

The ingredients tested, briefly: update method (PPO versus TRPO); reward design (gin-to-knock payoff ratios, knock-first, gin-first, deadwood-reduction bonuses, and dense step rewards at several horizons); an opponent curriculum (three stages — random, then a growing pool of the agent's own past checkpoints, then self-play mixed in; the pool sampled both uniformly and with PFSP weighting that favours harder opponents); warm-starting from trained weights; keep-the-best checkpointing; input representation (raw 4×52 versus a learned embedding); network shape (fully connected, convolutional, set-based, recurrent, attention); imitation learning (DAgger); and using an LLM as the sparring partner. Headline numbers come from 2000 games with 95% confidence intervals and seats rotated; broader sweeps use 400–600 games per cell. Results across random seeds are aggregated with IQM (the interquartile mean) and stratified bootstrap confidence intervals.

Findings

Five ingredients helped. In the authors' own summary: "Trust region updates, a well-aimed reward, a curriculum of tougher opponents, warm starting, and keeping the best checkpoint all help." In numbers, win-rate against the expert is 15.0% for PPO and 22.5% for TRPO; the strongest self-play configuration reaches roughly 30%; stacking everything that worked reaches 34.2±2.1%. Warm-starting and keep-the-best are each reported as recovering 2–3 points. Turned around: the expert wins 70–99% of games against the trained agents. The yardstick still stands well above them.

The list of things that did not help taught me more. A learned embedding did worse than the raw 4×52 input. Imitation learning "drove training loss to near zero but produced an agent winning almost no games" — the authors file this under causal confusion, where the student reproduces the expert's moves without learning the reasoning behind them. Dense step rewards make the agent "farm instant points instead of trying to win". Using an LLM as the opponent produced competent play at 9 to 27 seconds per move, orders of magnitude too slow to train against.

On reward design the wording is blunter still: "Paying three times more for a gin than a knock leaves the gin rate under one percent." Raising the payout does not change behaviour if the move is structurally hard to reach. That sits neatly beside the earlier fact that the expert itself hardly ever gins.

And then the comparison at the paper's core. An ISMCTS search that samples possible hidden hands and searches the resulting trees (run at 60 rollouts), graded fairly, reaches only 26% against the expert; the same search granted oracle access to the hidden cards reaches 85%. The authors write that this gap "quantifies the value of hidden information". Meanwhile the network-shape sweep landed inside one narrow band with overlapping intervals: convolutional 31.1%, set-based 30.4%, fully connected 26.7%. As a check in a second game they also run Leduc Hold'em, a standard miniature poker, where tabular learning closes to a return of -0.085 against the CFR optimum (random sits around -0.78).

Where this is useful

The first takeaway is the plainest. Before you put AI into your game, write the opponent that does not learn. A solver will do; so will a greedy heuristic. The requirements are only that it return the same judgement every time and that it be decently strong. If I were building a competitive card or board game, I would spend the first sprint on this before touching machine learning. Once a yardstick exists, every later improvement can be stated in points moved. The paper's closing line says exactly this: "Build that reference first, and the rest measures itself."

A match in Inscryption: the player faces one fixed opponent across the table, trading cardsInscryption (from the official Steam screenshots)

The second is to reread reward design as player-incentive design. "Paying three times more for a gin than a knock leaves the gin rate under one percent" has the same shape as a symptom every designer has seen: nobody uses the special move no matter how far you raise its damage multiplier; the high-scoring combo's point value goes up and its frequency does not move. On this paper's logic, the thing to fix is not the payout but the reachability — the number of moves, the information required, the conditions that must coincide. If you are balancing scores in a hypercasual title, measure what fraction of positions actually admit the move before you touch its multiplier.

The third is the difficulty curriculum. What the paper found effective was not a pool of fixed strength but a rising schedule — random, then a pool of the agent's own past checkpoints, then self-play mixed in — combined with PFSP weighting that draws harder opponents more often. That is a training-side result, but it reads directly as onboarding design. If I were making a Sokoban-like, I would reorder the early levels so that instead of "many boards at one fixed difficulty" the player more often draws boards resembling the one they just failed. Keeping a pool of your own stumbles transplants cleanly.

The fourth is treating the amount of hidden information as a difficulty knob. What produced the gap between 26% and 85% was not smarter search but whether the hidden cards were visible. In puzzle terms, covering part of the board, letting the player flip one hint, or showing versus hiding the remaining move count may move difficulty more than anything algorithmic. For those of us whose reflex is to enlarge the board when we want a harder level, this result is worth remembering.

The fifth is operational detail. Standing an LLM up as a live, responding opponent does not work at this scale, judging by 9 to 27 seconds per move; the reasonable placement is asynchronous work — drafting levels, rephrasing text, summarising play logs. The evaluation protocol is also worth copying: headline claims on 2000 games with seats rotated and 95% confidence intervals, finer comparisons on 400–600 games per cell. When you are unsure how many matches you need before you can assert a balance change, those are usable practical floors.

Limitations

The weaknesses the author acknowledges are clear. First, the paper states that "specific numbers are about Gin Rummy". Second, the expert used as the yardstick is a strong heuristic, not a game-theoretic optimum, so 34.2% says how the agent fares against this particular expert, not against optimal play. Third, the agent's policy is a reactive feed-forward network with no machinery for estimating the opponent's hand. Fourth, the LLM was tested only as a live opponent, not as a generator of training data. The listed future directions are opponent-hand estimation, belief-state search conditioned on inferred cards, equilibrium-finding methods such as CFR and NFSP, stronger memory, and offline learning from LLM-generated games.

What I would flag here is, first, the number of random seeds: two for the main algorithm comparisons, eight for Leduc, per the text. The careful aggregation with IQM and bootstrap intervals is welcome, but generalising the gap between 15.0% and 22.5% into "TRPO is better" on two seeds is more than I would want to claim. To the authors' credit, the seed counts are stated as plainly as the 400–600 and 2000 game counts, so a reader can go and check.

Second, the interpretation of the ceiling. "Information rather than network size" is an attractive reading, but its central evidence is the gap between a fair and an omniscient ISMCTS. What is being compared there is a different kind of player — a search procedure — not a direct experiment in which the learned agent is handed more information. So I take it as strong indirect evidence that hidden information is valuable, not as a general law that scaling the model is pointless.

Two last points. One is that no human player appears anywhere in this study. What is measured is win-rate against a fixed opponent; enjoyment and felt challenge are out of scope, so a designer borrowing this for difficulty work has to build that bridge themselves. The other is review status: this is a preprint under submission to AIIDE 2026, posted recently enough that it has not been widely discussed. The numbers could still change through review.

How Fukai reads it

I want to place this study in the lineage of building the reference first, not the lineage of building a stronger AI. In the vocabulary of design criticism it is closer to fabricating your own measuring instrument than to improving a model. The moment one immovable yardstick is fixed in place, a hundred scattered runs turn into a table you can read across — and my reading is that what really helped was less TRPO or the curriculum themselves than the existence of that table. The place where puzzle difficulty repeatedly defeats us is the same place: solve-rate as a yardstick keeps moving. So the import worth taking from this paper is not the list of five ingredients but the ordering of the procedure — write the opponent that does not move, first.

Closing

If you want to go deeper, reading the two lineages this paper builds on will give you the map: the self-play line represented by Silver et al.'s AlphaGo Zero and AlphaZero, and the equilibrium-computation line for imperfect-information games that begins with Zinkevich et al.'s CFR. On method details, Schulman et al.'s TRPO, PPO and GAE sit underneath; on the search side, Cowling et al.'s ISMCTS; and on the pitfalls of imitation, de Haan et al.'s work on causal confusion. For those who want to touch the code, the authors have released the pipeline together with the expert implementation.

We have covered neighbouring ground on this site before. On the practicalities of putting reinforcement learning into game AI, there is our piece on Sestini et al.; on how curricula actually work, our pieces on Zhou et al.'s verifier curriculum and Ponnock et al.'s Mario curriculum both mesh directly with today's reading. If the "build the yardstick first" idea appeals to you, those are the next stops.

References

Papers and materials referenced in this article:

・A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong (Nima Kelidari, Mohammadsaeed Haghi, Mahdi Salmani, 2026, arXiv preprint arXiv:2607.06854)

・Full HTML text of the paper (including the win-rate tables, the ablation list, and the Leduc Hold'em check)

・Nikelroid/adversarial-coevolution (the authors' released implementation, including the fixed expert, the LLM serving stack and a web client)

・Prior work the paper builds on: Schulman et al. (2015, 2016, 2017) on TRPO, PPO and GAE; Silver et al. (2017, 2018) on AlphaGo Zero and AlphaZero; Zinkevich et al. (2007) on CFR; Ng et al. (1999) on potential-based reward shaping; de Haan et al. (2019) on causal confusion; Cowling et al. (2012) on ISMCTS

・Related articles on this site: on Sestini et al. on reinforcement learning and game AI / on Zhou et al.'s verifier curriculum / on Ponnock et al.'s Mario curriculum

・Review status: an arXiv preprint (arXiv:2607.06854, posted July 2026) whose body footnote states it is under submission to AIIDE 2026. It has not been peer-reviewed and has no DOI; posted recently, it has accumulated no citations and has not yet been widely discussed

Reactions (no login)

Anonymous • one of each per visitor per day

Part of these series

Paper DigestEpisode 64 of 102

Read next