PAPER-DIGEST · 2026-08-23
Pfau & Vrettis: What Happens When Players Generate Their Own Pokémon Cards — Fukai Reads
PCG / trading card games / generative AI and the sense of authorship
TL;DR
"The Pokemon card you just made up in your head becomes a real-looking card in twenty seconds." Today I read a paper that actually did this. A player types a name and a short bit of flavour, and the AI produces the artwork, the attacks and the numbers as one finished card. Forty-nine students made 196 of them.
Magic: The Gathering Arena (Wizards of the Coast, 2023), gameplay image from its Steam store page
What interests me is the shape of their satisfaction. Visual satisfaction averaged 4.25 out of 5. And when asked whose idea the final card was, 93.5% answered that it was their own. The AI did the making, yet the card still felt like theirs. The authors call this "procedural relatedness".
One caution up front. None of these cards has ever been played. Neither the overpowered ones nor the weak ones have been checked in an actual match. I will come back to that properly later on.
About this paper
The paper is titled "From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study". The authors are Johannes Pfau and Panagiotis Vrettis, both at Utrecht University in the Netherlands. It appeared on arXiv on 30 April 2026 as a preprint, that is, a manuscript posted before peer review. It is listed as arXiv:2604.27972v1 under cs.AI, with a CC BY-NC-ND 4.0 licence.
I picked it for two reasons. First, the papers I have covered here recently leaned heavily towards "give an AI a game and score it". Second, this study unusually puts the maker's feelings at the centre. Not many generative-AI papers measure whether the result felt like your own rather than how good it was.
Only about four months have passed since posting, so it is safe to assume the citation count is still near zero. This has not been widely discussed yet. I also found no statement in the text that it has passed peer review. I read it with that in mind.
When only the strong cards survive
Trading card games, where you collect cards, build your own deck and play, are a big industry. The paper starts from Magic: The Gathering in 1993 and notes the genre is now worth billions of dollars. Magic alone, it says, engages over 50 million players and passes one billion dollars in revenue.
The genre has a recurring problem. Every time a new set arrives, the strongest combinations harden within a few months. The paper describes the resulting metagame, meaning the overall picture of what is considered strong right now, as "very limiting, stale, and repetitive". Thousands of cards exist on paper; only a handful get played.
The second problem is power creep. To make new cards attractive, their power drifts upward release after release. The paper says this happens in "as good as every live game".
What the authors regret most is the loss of attachment to individual cards. People used to treasure one card because they loved the character on it. In an environment sorted purely by strength, that relationship does not survive. Recovering it is where this research starts.
Slay the Spire (Mega Crit, 2019), gameplay image from its Steam store page
Automatic card generation itself is not new. The paper lines up earlier work: RoboRosewater (2015), which had a neural network write Magic cards; Summerville and Mateas (2016), who completed card text with a sequence model; and Chen and Guy (2020), who generated Hearthstone cards from a grammar and predicted their balance.
What this paper claims as new is two additions. One is generating the artwork alongside the attacks and numbers. The other is putting the player at the entrance of the pipeline. The authors pose three questions: how to unite generative AI across multiple facets of design, visual and mechanical; how far player-centric generation can preserve both fidelity to the person's own idea and visual quality; and what iterative strategies best close the gap between an idea and what gets generated.
Five steps to build one card
The system has five steps: collect reference cards, retrieve similar ones, fill in the numbers and attacks with a text model, draw the artwork with an image model, and finally lay it all out as a card.
The five-step flow described in the paper (diagram made by me; not a reproduction of the paper's figures)
Step one. They pull 15,411 cards from the official Pokémon TCG Developer Portal API and narrow that to 993, keeping only basic Pokémon up to generation 9. Each card is handled as structured data in JSON, a format of field names paired with values: name, flavour text, types, HP, attacks, weaknesses and so on.
Step two is retrieval. The system looks for existing cards close to what the player wrote, using the name, flavour text and types. It uses an embedding model called nomic-embed-text-v1.5, a tool that turns text into a row of numbers so closeness in meaning can be measured. Handing those close examples over as references is what the paper calls RAG, or retrieval-augmented generation.
Step three has a text model fill in the blanks. The model is Qwen3-14B. The instruction is to "complete the JSON that I started" with the fields hp, abilities, attacks, resistances, weaknesses and retreatCost. It is not writing from a blank page; it is filling in fixed slots.
Step four is the picture. An image model, FLUX.1-dev-Q8, runs locally through ComfyUI. To keep the art style consistent they blend two LoRAs, a technique that nudges a model's style with a small amount of extra training: a Niji style and a Pokémon (Ken Sugimori) style. Step five assembles the card with a web tool called Pokécardmaker. One card takes about 20 seconds, computed mainly on an NVIDIA RTX 5090.
What the 196 cards said
Forty-nine people took part: master's students in Game and Media Technology and in Artificial Intelligence, 75.5% male and 24.5% female. Each made four cards, giving 196 in total. Participants could install the whole pipeline locally or use a server the authors provided, and were free to change the models, the parameters and the system prompt.
Ratings used five-point scales. Visual aesthetics averaged 4.25 (SD 0.9). "How closely the outcome image resembled what you had in mind" came to 3.95 (1.07). "How well the attacks and values suited the concept" came to 4.04 (0.75). All three sit on the good side.
The breakdown is more informative. For the artwork, 35.5% were convinced at the first try and 38.4% after slight polishing. For the mechanics, first-try conviction was higher at 51.0%. Never convinced: 11.6% for images, 10.2% for mechanics. In total 88.4% of visual outcomes and 89.8% of mechanical ones ended up convincing.
The ways people fixed things are also broken out. Rewriting the prompt was the most common at 51.8%. Regenerating with the same prompt came to 18.8%, changing the original idea 13.4%, tuning parameters 6.3%, and manual touch-up 0.9%. And 8.9% gave up.
Finally, 93.5% attributed the finished design to their own idea, and only 1.4% to the AI. Fourteen participants explicitly said the result exceeded their expectations, and seven said they preferred the generated design to what they had first pictured.
The open-ended answers were organised by thematic analysis, with two of the authors independently applying open codes. That work also shows that ten participants read their failures as limits of the model rather than flaws in the approach: the tool was immature, but the method was sound.
The paper's figures show named examples. Successes include Glister, Ferrafox and Möbiusect. Failures include Zagarin (under-specified mechanics), Aurellune (imbalanced) and Oricrane (repeated text).
What builders can take away
First: design for filling in fixed slots rather than free creation. This pipeline uses the structure of existing cards as its template. If I were building a deck-builder, I would hand the AI a form with holes in it, not a blank page. The more you narrow the freedom, the higher the share of usable results should be.
Second: make "say it differently" the primary repair path, not "make it again". More than half of participants (51.8%) solved their problem by rewriting the prompt. A screen with a big regenerate button and little else does not match how people actually fix things. An easily editable input field does more.
Balatro (LocalThunk / Playstack, 2024), gameplay image from its Steam store page
Third: instrument the 8.9% who gave up. You cannot judge a generation tool by looking only at the results people kept. Log which attempt they dropped out on and which step stalled, and the next fix becomes visible.
Fourth: ask about looks and mechanics separately. This study measured artwork satisfaction and mechanical satisfaction apart, and the numbers differed. Often only one side needs fixing. For generated puzzles too, I think "does it look right" and "did it feel good to solve" deserve separate questions.
Fifth, a note on scale. This whole thing runs on one desktop machine: a 14B-class text model, a quantised FLUX for images, roughly 20 seconds per card. No external API is required, which puts it inside the reach of a small team. That participants could install and run it themselves backs this up.
Sixth, a caution. Do not drop generated cards straight into a place where winning and losing matter. The reason is in the next section. Start with solo play, collection and showing cards off, where nobody is hurt by a loss.
What we still do not know
The authors are clear about their own weaknesses. The largest is that none of the cards has been played. The paper states plainly that they were not deployed into printed or digital play. Balance therefore remains unverified.
They also acknowledge that balance matters. There is a line to the effect that generated content should be viable without strictly dominating, or being dominated by, the alternatives. They mention their own combat simulator but defer the work, saying it would exceed the scope of this paper.
They further list the narrow sample of 49 master's students in game and AI programmes, the fact that only the Pokémon TCG was tested, and concerns about copyright and intellectual property. On the model side they name hallucinations such as inventing mechanics that do not exist, limited memory and reasoning, poor handling of numbers, and non-determinism. They also admit each card is generated in isolation, with no view of the set as a whole.
What I would add is a point about who did the scoring. The person judging whether a card matched what they had in mind was the person who made it. We tend to rate our own effort generously. How a third party would score the same 196 cards is not something this study can tell us.
My second point is that there is nothing to compare against. A figure of 93.5% for "my own idea" is close to the ceiling. But with no control group, for instance people simply handed cards the AI made on its own, there is no ruler to say whether 93.5% is high. The authors' "procedural relatedness" also reads to me as claimed rather than measured with a scale for attachment itself.
How Fukai reads it
I want to read this study as a question of placement: where do you put the generative AI? Here it does not do the imagining. The player supplies the idea, and the AI is relegated to filling in fixed fields. That, I think, is why the sense of authorship survived. In the vocabulary of design criticism, this is closer to automating the fair copy than automating the creation. Had the order been reversed, with the AI proposing the concept, I doubt the figure would have been 93.5%. That is my reading, not something the paper demonstrated with a controlled comparison.
Closing
For anyone who wants to widen the map, here is some of the work this paper builds on. RoboRosewater (2015) was an early attempt at having a neural network write Magic cards. Summerville and Mateas (2016) completed card text with a sequence model. Chen and Guy (2020) generated Hearthstone cards from a grammar and went as far as predicting their balance. I name these only as citations from the paper; I have not gone back to the originals.
We covered AutoBG here before, a system that supports board game design from concept to finished product. That one sits on the "let the AI design" side. Reading it alongside today's paper sharpens the outline of how much the seating position of the generative AI changes the result.
References
Papers and materials referenced in this article:
・Google Scholar profile of Johannes Pfau (first author, Assistant Professor at Utrecht University)
・Prior work the paper builds on: Milewicz / RoboRosewater (2015) on Magic card generation, Summerville & Mateas (2016) on sequence-model card completion, Chen & Guy (2020) on grammar-based Hearthstone generation with balance prediction, Summerville et al. (2018) on PCG, and Maleki & Zhao (2024) on LLMs in PCG. I name these as citations from the paper; I have not re-read the originals.
・Sources of the game images: the Steam store page for Magic: The Gathering Arena (Wizards of the Coast, 2023), for Slay the Spire (Mega Crit, 2019), and for Balatro (LocalThunk / Playstack, 2024). None are connected to the paper; they are real card games cited to match the topic.
・The five-step diagram is my own work and is not a reproduction of any figure from the paper.
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 66 of 104
