SERIAL

Paper Digest

104 episodes · updated 2026-10-06

Level generation, difficulty estimation, AI reasoning — games research is fascinating but rarely reaches players. Each episode, Fukai picks one paper and unpacks its method and findings for everyone.

日本語版を読む →

Episodes

  1. Ep. 104
    Yamamoto & Shiba: When both options share the same payment, people skip right over it — Fukai Reads
    2026-10-06

    A peer-reviewed paper by Shohei Yamamoto (Rikkyo University) and Shotaro Shiba (Toyo University) in Judgment and Decision Making (13 May 2026). Three preregistered experiments (1,400 recruited) show that people largely ignore payments common to both options and decide on the shape of the differing part. With identical totals, re-presenting a choice as an investment or a loan changed patience; with real money, the loan framing's effect held.

  2. Ep. 103
    Höhs et al.: Reasons for tend to look alike, reasons against scatter — Fukai Reads
    2026-10-05

    A peer-reviewed paper by Höhs, Rebholz and Hütter (University of Tübingen) in Judgment and Decision Making (23 July 2026). Three preregistered experiments (623 participants in total) compared reasons for and against decisions. Reasons for resembled each other and felt more relevant; reasons against were more scattered and felt more novel. In Experiment 3, the more consistent the reasons against looked, the more people moved toward not doing it.

  3. Ep. 102
    Choi & DeKay: Even with truly random sequences, people judged a repeat about 7 points less likely after a run of four — Fukai Reads
    2026-10-04

    A peer-reviewed paper by Yeonho Choi and Michael L. DeKay (The Ohio State University) in Judgment and Decision Making (29 September 2026). It re-analyzes a 2025 study that claimed the gambler's fallacy disappears in probability judgments when sequences are truly random. Re-aggregated by run length, the fallacy clearly appears: after four balls of the same color, judged chances of a repeat were 7.17 points lower on average. Only 29.0% of individuals, however, showed a credible fallacy on their own.

  4. Ep. 101
    Li: 'Make It Tall' Worked, 'Make It Hollow' Didn't — Mapping Which Requests a Neural Generator Honors — Fukai Reads
    2026-10-03

    A paper by Alex Chengyu Li on neural building generation. A generator trained on 10,310 Minecraft builds was given six condition tokens, and responsiveness was measured on 14 properties: 9 moved a lot, 4 a little, and only the share of air inside barely moved. How easily un-requested properties move was predicted from training data with rank correlation 0.879, but on just 10 properties — too few to generalize yet.

  5. Ep. 100
    Engel et al.: When the right path through a maze flips, other people's footprints help people switch faster — Fukai Reads
    2026-10-03

    A peer-reviewed, open-access paper by Christoph Engel, Thomas Holzhausen and Dorothee Mischkowski (Max Planck Institute for Research on Collective Goods) in Judgment and Decision Making (21 September 2026). In a preregistered experiment with 288 people, a maze shortcut rule learned over 50 mazes was secretly reversed. Seeing another player's footprints raised the rate of switching by 51.0%, while a paid hint button made no significant difference and was used by only 55.21% of those who had it.

  6. Ep. 99
    Xu & Verbrugge: Small Local AIs Translate a 5×5 Grid into a Different Skill Every Time — Fukai Reads
    2026-10-02

    A paper by Kaijie Xu and Clark Verbrugge on generating combat skills. Local language models translate elemental cells a player arranges on a 5×5 grid into skills. The models kept 0.994 element coherence while producing different skills from the same grid each time — a combination four baselines could not match. But they barely read vertical lines, and there is no player study yet.

  7. Ep. 98
    Lessard et al.: Can a Machine Dig the Good Stories Out of a Simulation? — Fukai Reads
    2026-10-01

    A paper by Jonathan Lessard and colleagues on emergent narrative. They tried to automatically pick 'interesting stories' out of 228K events produced by a social simulation, using 'unexpectedness' and 'dramatic situations' as heuristics, and had 24 readers rate the results. The clear favorite turned out to be the control story, chosen with no heuristic at all.

  8. Ep. 97
    Surendran & Cooper: In fill-in-the-blank microgames, a detailed design doc could not beat a short prompt — Fukai Reads
    2026-09-30

    A paper by Akash Surendran and Seth Cooper (PCG Workshop at FDG '26) on fill-in-the-blank microgame generation. Player-chosen words were inserted into LLM prompts to generate 96 games across four genres, comparing a detailed design-document template with a short one. Only about half of either could be completed, while the short prompts brought in more of the words and produced more varied games.

  9. Ep. 96
    Endrovski et al.: The people using level generation least are the ones who build levels — Fukai Reads
    2026-09-29

    A paper by Bojan Endrovski, Joris Dormans and Rafael Bidarra on procedural level generation tools. Surveying 120 game developers, they found artists use level generation far more than designers (mean usage 0.58 vs 0.33), and that what developers want from AI is creative control over the final output (73.3%), not automation.

  10. Ep. 95
    Sears and Weisberg: Readers couldn't spot AI-written stories, and rated them higher when told a human wrote them — Fukai Reads
    2026-09-25

    A paper by Sydney Sears and Deena Skolnick Weisberg on telling AI-written short stories from human ones. Across three studies with over 2,500 participants, readers could not tell them apart (39.39% and 51.97% correct), rated AI stories higher when authorship was hidden, and rated stories higher when told a human wrote them.

  11. Ep. 94
    Kemmerly et al.: Confident without being right, people bet more and searched more — Fukai Reads
    2026-09-24

    A paper by Rowan Kemmerly and colleagues on unjustified confidence. In a preregistered study of 997 people betting on college football, confidence not backed by accuracy went with bigger bets and more buying of paid information — but was almost unrelated to how much people revised their guesses in light of information.

  12. Ep. 93
    von Gugelberg: How people solve differed more by 'which item' than by how hard it was — Fukai Reads
    2026-09-23

    An eye-tracking study of figural matrix problems by Helene M. von Gugelberg. Recording 300 people item by item, it reports that stronger reasoners build the answer before checking options, and that people differ more in how their approach changes across item positions than in how they react to difficulty.

  13. Ep. 92
    Chien et al.: Say “90% payout” and the game looks winnable — and the meaning gets lost — Fukai Reads
    2026-09-23

    A paper by Chien, Han, Ludvig and Eich on how gambling odds are displayed. In a preregistered experiment with 187 younger and 199 older adults, the same odds were shown either as “90% payout” or as “the game keeps 10%”. In both age groups, the payout wording lowered comprehension and raised people’s sense that they would win.

  14. Ep. 91
    Liquin: Told the answer would cost them, people wanted it just as much — Fukai Reads
    2026-09-21

    A paper by Emily G. Liquin on curiosity and effort. Across three preregistered experiments with 419 participants, telling people how much effort an answer would cost reliably reduced whether they went after it, yet barely moved how curious they said they felt.

  15. Ep. 90
    Baghal: a model that never tried to solve anything drew solvable Sokoban 77.4% of the time — Fukai Reads
    2026-09-19

    A preprint by Sina Baghal (arXiv:2608.15958, not peer-reviewed). When a machine generates Sokoban levels, the usual step is to run a solver on every candidate to check that it can be finished. This manuscript reports that a 4.9M-parameter model trained only to fill in hidden tiles — with no solver, no reward and no solvability labels — produces solvable puzzles 77.4% of the time. Of the failures, 94.5% become solvable by deleting a single interior wall, which lifts effective solvability to 98.7%.

  16. Ep. 89
    Guo et al.: humans drifted off the greedy move within about ten games; self-evolving AI did not — Fukai Reads
    2026-09-18

    A preprint by Yingying Guo and four co-authors (arXiv:2608.07490, not peer-reviewed). They propose a way to measure how repeated play changes the way humans and language agents choose moves. Thirty-two students played 709 games across three board games, and four self-evolving language agents were run through the same metric space. Humans mostly shifted from the locally greedy move toward more global play: on the game-specific behavioral metrics, 11 of 12, 10 of 11 and 8 of 9 participants improved. The agents' gains were short-lived. The authors write that the central limitation is not the absence of reflection but the failure to convert reflection into reusable changes in behavior.

  17. Ep. 88
    Williams et al.: only six brain-imaging studies of Sudoku exist in the world — Fukai Reads
    2026-09-17

    A peer-reviewed systematic review by three authors in the UK and South Africa (Frontiers in Neuroimaging, published 20 April 2026). Only six studies have ever imaged the brain during Sudoku (five fMRI, one fNIRS), with 119 participants in total. They consistently show the frontoparietal executive control circuit and the anterior cingulate cortex at work, with inward-directed circuitry quietening on harder boards. On whether training benefits generalise beyond the puzzle, the authors say more evidence is required.

  18. Ep. 87
    Li et al.: the AI that showed up uninvited was closed by five players out of ten — Fukai Reads
    2026-09-17

    A peer-reviewed paper by Jiahong Li and eight co-authors (accepted to the 2026 IEEE Conference on Games, arXiv:2609.13718). They built PEARL, an AI support agent that retrieves expert-annotated move explanations and structurally similar peer boards, into Parallel, a puzzle game for learning parallel programming, and evaluated it with ten players. Participants rated the existing visualization tool more useful, and at least five minimized or abandoned the AI during play. Frustration scored 43.2 against 56.5. The authors conclude that what players rejected was not the content of the help but its unsolicited delivery, and ask of their own system: are we building Clippy?

  19. Ep. 86
    Hendijani and Steel: Letting people choose moved nothing; a number in the corner of the screen did — Fukai Reads
    2026-09-16

    A peer-reviewed paper by two authors from the University of Tehran and the University of Calgary (Frontiers in Psychology, published 13 August 2026). In a memory test with 270 people it compared letting participants choose against paying them per correct answer: the reward added about seven recalled words, while choice produced no statistically confirmed effect. Eye tracking showed the reward's effect ran through whether people looked at the on-screen reward display.

  20. Ep. 85
    Melo Legarda et al.: Before changing difficulty by heartbeat, they built a way not to change it — Fukai Reads
    2026-09-15

    A peer-reviewed paper by four authors from Universidad del Cauca and Colegio Mayor del Cauca, Colombia (Applied Sciences 16(17):8511, published 27 August 2026). They built a mechanism that adjusts game difficulty from a chest-strap heart sensor and logged eight sessions totalling 6 hours 48 minutes. Mean end-to-end latency was 2.06 s. The striking number: against 191 committed state transitions there were 83 flips the automaton withheld — nearly a third of the change-or-hold decisions land on “do not change”. No subjective data was collected, and the authors never claim the game became more enjoyable.

  21. Ep. 84
    Dygert & Jarosz: People who repair a misread sentence also solve insight puzzles — Fukai Reads
    2026-09-14

    A paper by Sarah K. C. Dygert and Andrew F. Jarosz on insight and sentence comprehension. Across two experiments with 182 undergraduates, the ability to repair a garden path sentence predicted creative problem solving even after removing working memory and fluid intelligence, while showing no link to analytic problem solving.

  22. Ep. 83
    Tudor et al.: The scoreboard was one query away, and the agent never opened it — Fukai Reads
    2026-09-11

    An arXiv preprint (submitted 2 September 2026) by seven authors from Oxford and elsewhere. They wired 76 tool endpoints into Sid Meier's Civilization VI and had language-model agents play whole games of 300+ turns. Agents queried victory progress only once every 30-75 turns (the supplied playbook recommended every 20), and in 7 of 20 losses that were foreseeable they never checked it in the final 20 turns. Commitments the agents wrote down for themselves were carried out within ten turns only 48.2%-65.8% of the time.

  23. Ep. 82
    Nagaya et al.: Delete "don't bet" from the menu, and twice as many people take the risk — Fukai Reads
    2026-09-10

    A peer-reviewed, open-access paper by Kazuhisa Nagaya and Fuminori Ono (Yamaguchi University) and Kazuya Nakayachi (Doshisha University), published in Judgment and Decision Making on 13 July 2026. It re-measures small-stakes loss aversion by rewriting the choice as "bet or bet" instead of "bet or don't bet". Across three studies with 1,345 participants, the share of people picking the risky option jumped from 23.3% to 50.0%. Much of what has been called loss aversion may be a separate habit: a preference for not acting.

  24. Ep. 81
    Macchi et al.: Letting people touch it changed nothing; drawing it as parts raised solving by 33 points — Fukai Reads
    2026-09-09

    A peer-reviewed paper by Laura Macchi and colleagues at the University of Milano-Bicocca (Journal of Intelligence, 9 May 2026) testing the received wisdom that handling a problem's materials makes insight more likely. Asked to build four equilateral triangles from six pencils, participants given the real pencils went from 27.3% to 32.6% — no significant difference. But swapping a matchstick arithmetic puzzle for a photograph of real matchsticks lifted solving from 46.7% to 79.3%, with nothing to touch. What worked was not the hand, but the picture saying "I am made of parts."

  25. Ep. 80
    Ghasemi et al.: People know what a default does — and aim it differently at allies and rivals — Fukai Reads
    2026-09-08

    A peer-reviewed paper by Omid Ghasemi, Ben R. Newell and colleagues (Judgment and Decision Making, 4 September 2026) pushing back on the well-known 2017 finding that people fail to use defaults strategically. Across three card-game experiments, participants set the default in their own favour on more than 80% of trials — the high-value card for teammates, the low-value card for opponents — and, shown only someone else's default choice, worked out which option was better 79.5% of the time.

  26. Ep. 79
    Baek et al.: Ordering a Level That Is 75% Zelda and 25% Mario, in Plain Words — Fukai Reads
    2026-09-07

    A paper by In-Chang Baek and four co-authors at GIST and Dongguk University (arXiv:2603.26782, an un-peer-reviewed preprint). They put 5,576 levels from Zelda, Dungeon, Lode Runner and Super Mario Bros into a single latent space so that levels can be blended across games using text and a mixing ratio. Sharing one model instead of four costs about 4.4% in overall similarity, and turning the ratio dial swaps similarity between the two source games as intended. Blending through a single written instruction, however, remains weak.

  27. Ep. 78
    Siper et al.: Evolve the Level Generator, Not the Level — and Let It Grow Its Own Toolbox — Fukai Reads
    2026-09-06

    A paper by Matthew Siper, Ahmed Khalifa and Julian Togelius (arXiv:2608.17947, accepted at IEEE Conference on Games 2026). Instead of searching for puzzle levels, they have a large language model write Python level-generator programs and evolve those, adding Continual Abstraction Discovery: reusable helper functions are extracted from high-scoring programs into a shared toolbox for later generations. Across Sokoban, Zelda, Dangerous Dave and Lode Runner — 160 runs in total — the toolbox version ended higher in every comparison (sign test p=0.008).

  28. Ep. 77
    Lee & Ko: Human Umpires Shrank the Strike Zone by 17 Points With Two Strikes — Fukai Reads
    2026-09-05

    An arXiv preprint by Kichang Lee and JeongGil Ko of Yonsei University. Using the Korean Baseball Organization's switch to automated ball-strike calling as an immovable ruler, they audit 1,216,246 pitches — restricted to those on the edge of the zone — to see how human umpires' calls moved with context. Called-strike probability was 17.17 percentage points lower in 0-2 counts and 6.61 points higher in 3-0 counts, and the pattern disappears under automation.

  29. Ep. 76
    McCaughey et al.: People Change How Much Information They Buy Only When They Are Told the Price Changed — Fukai Reads
    2026-09-04

    An open-access, peer-reviewed paper by Linda McCaughey and two co-authors in Judgment and Decision Making. Across five experiments and 755 analyzed participants in a task where every observation costs money, people did change how much information they bought when the price changed — but almost entirely through planning ahead, not through what they experienced while playing. A direct hit for anyone pricing hints or scouting.

  30. Ep. 75
    Lohn: Adding the Strongest Possible Move to Rock-Paper-Scissors Only Buys You 55.6% — Fukai Reads
    2026-09-03

    An arXiv preprint by Andrew J. Lohn of Georgetown's CSET, solving what happens when you add "Dynamite" to Rock-Paper-Scissors. Giving one player the strongest possible move raises their win rate only from 50% to 55.6%, and the wins arrive through Rock rather than through Dynamite. Widen the move set and the gap shrinks further, while undominated moves quietly drop out of the optimal strategy.

  31. Ep. 74
    Elshamy et al.: Read the player's skill, then redraw the level itself — Fukai Reads
    2026-09-02

    A Scientific Reports paper from Elshamy and colleagues at E-JUST on inferring player skill and rewriting the terrain of the level itself. Where conventional dynamic difficulty adjustment tunes enemy health and item drops, this pipeline rearranges floors, gaps and enemies in place. Skill classification reached 97.82% accuracy; 74.1% of rewritten levels remained completable.

  32. Ep. 73
    O'Neill et al.: A Board Where Nothing Makes You Keep Your Word — Fukai Reads
    2026-09-01

    A paper from UC Berkeley introducing C2C, a four-player conquest game built to measure negotiation and betrayal. On a board with no mechanism at all to enforce agreements, language models and humans played over 1,100 games — and humans turned out to make far fewer promises than the AI agents did.

  33. Ep. 72
    Gao & Dubé: Letting a machine do the first read of player-made math levels — Fukai Reads
    2026-08-31

    An arXiv preprint by Jie Gao and Adam K. Dubé of McGill University. A children's math learning game has a Creative Mode in which advanced players build their own levels, but reading every submission by hand does not scale. The authors extracted features from 206 levels (86 by experts, 120 by players) and trained a classifier to do the first read. Random forest gave the best recall and F1, scoring 82.42±5.71% accuracy and 72.70±9.22% F1 in the outer loop.

  34. Ep. 71
    Ahmetovic et al.: Handing half your controller to someone else — Fukai Reads
    2026-08-29

    A study from the University of Milan on how people with upper-limb impairments operate games. The authors built GamePals, a framework that splits control in off-the-shelf titles, and had 13 participants play Rocket League with a human copilot and with a software copilot. Seven said they could not have played without support, while participants also mistook the copilot's actions for their own.

  35. Ep. 70
    Xu & Verbrugge : faire de la « direction de la gravité » et du « temps » des coordonnées de génération de niveaux — lu par Fukai
    2026-08-28

    Un article évalué par les pairs de Kaijie Xu et Clark Verbrugge (Université McGill), présenté à FDG 2026. La génération automatique de niveaux a longtemps consisté à créer d'abord le terrain, puis à vérifier après coup des mécaniques comme l'inversion de gravité ou les plateformes mobiles. Cet article élève la mécanique elle-même au rang de coordonnée et propose HDPCG, qui cherche un chemin dans un grand graphe ajoutant une « couche » ou un « instant » aux coordonnées (x, y). L'erreur sur l'intervalle de changement descend à 0,000–0,002, et la robustesse des chemins alternatifs est 9 à 10 fois supérieure à celle d'une base sans guidage.

  36. Ep. 69
    Collins et al.: People judge a brand-new game with one move of lookahead and six imagined playouts — Fukai Reads
    2026-08-27

    A peer-reviewed Nature paper by Katherine M. Collins and colleagues (MIT and others). More than 1,000 people were shown 121 novel games from the tic-tac-toe family and asked, before playing, whether each looked fair and fun. The Intuitive Gamer model - one move of lookahead spent inside six simulated playouts - explained the fairness judgements at R2 = 0.81 against a human ceiling of 0.82, beating the deep-searching Expert Gamer (0.65) and MCTS (0.60).

  37. Ep. 68
    Byers et al.: Who Is Player Time Designed For? — Fukai Reads
    2026-08-26

    A peer-reviewed CHI 2026 paper (Best Paper Honourable Mention) by Thomas Byers, Martin Gibbs and Bjorn Nansen. Twenty hour-long interviews with AAA, indie, mobile and live-service developers produce a grounded account of how the time-shaped parts of a game get decided inside a studio: undocumented intuition, metrics spanning seconds to months, and design that works backwards from a number. The authors close with four heuristics — and they land squarely on anyone shipping a daily puzzle.

  38. Ep. 67
    Hu et al.: We Judge Others' Satisfaction Without Counting Their Options — Fukai Reads
    2026-08-24

    A peer-reviewed paper by Beidi Hu, Alice Moon and Eric VanEpps in Psychological Science (January 2026). Across six preregistered experiments with 10,092 participants, people factored choice set size into their own satisfaction but barely factored it into predictions of someone else's. Three things shrink the neglect: showing the different set sizes side by side, asking for a ranking, and restating the number of options. It bears directly on how we read playtests and pick rates.

  39. Ep. 66
    Pfau et Vrettis : et si l'on laissait les joueurs créer leur propre carte Pokémon — lu par Fukai
    2026-08-23

    Un preprint arXiv (non évalué par les pairs) de Johannes Pfau et Panagiotis Vrettis (université d'Utrecht). Lorsqu'un joueur écrit un nom et une brève description, un dispositif combinant recherche, modèle de texte et modèle d'image produit en environ 20 secondes une carte de style Pokémon. Les auteurs ont fait produire 196 cartes par 49 étudiants. La satisfaction visuelle atteint 4,25 sur 5, et l'adéquation des techniques et des valeurs 4,04. Enfin, 93,5 % des participants ont répondu que le résultat final était bien leur propre idée. Les cartes générées n'ont toutefois jamais été mises en jeu, et leur équilibrage reste donc non vérifié.

  40. Ep. 65
    Zhao et al.: People Build Their Own Reusable Parts While Solving Puzzles — Fukai Reads
    2026-08-22

    An arXiv preprint (not peer reviewed) by Pinzhe Zhao and three colleagues. Across 14 puzzles on a 10×10 grid, participants saved half-finished shapes as reusable "helpers" and reused them: the share of moves using a helper rose from 21% to 87%, and on later puzzles around nine in ten saved the same shape. Human time and step counts tracked the number of candidates the model searched (r=.82), not the length of the shortest program (r=-.20). Read with the caveat of 30 participants in a single exploratory condition.

  41. Ep. 64
    Kelidari et al.: A Card-Game Agent Is Only as Strong as the Yardstick You Build First — Fukai Reads
    2026-08-21

    An arXiv preprint by Nima Kelidari and two co-authors, under submission to AIIDE 2026. Using Gin Rummy and a hand-written fixed expert as an immovable yardstick, they run more than a hundred controlled experiments on what makes a lightweight reinforcement learning agent strong. Win-rate against the expert is 15.0% for PPO, 22.5% for TRPO and 34.2±2.1% with every working ingredient stacked; swapping network shapes leaves win-rates overlapping, while a search that can see the hidden cards reaches 85% against 26% for one that cannot.

  42. Ep. 63
    Battleday et al.: Measuring AI Discovery With 70 Games That Never Explain Their Rules — Fukai Reads
    2026-08-19

    An arXiv preprint by Ruairidh M. Battleday and fifteen co-authors. On DiG-bench — 70 text-string games with both rules and win conditions hidden, across seven tiers, 21 released publicly — the strongest single model beat 50 games and all models pooled beat 57, while all 70 were beaten by at least one human on a first attempt. Handed the ground-truth rules, the same model jumps from 18 games to 69, and agentic harnesses did not improve on the basic one.

  43. Ep. 62
    Mannem et al.: Some Puzzles Are Learnable, Some Are Not — Fukai Reads
    2026-08-18

    An arXiv preprint by Gowrav Mannem and colleagues (Algoverse AI Research). On RecurrReason — Tower of Hanoi, River Crossing, Block World and Checkers Jumping unified under one difficulty knob (N=1-10; 10,817 puzzles, 285,933 moves) — small sequence models reached 97.27% validation and 81.00% out-of-distribution on Block World, but only 11.11%/0.00% on Tower of Hanoi, 1.11%/0.10% on Checkers Jumping, and 0.00% everywhere on River Crossing. A 60M-parameter T5 beat a 124M-parameter GPT-2 on every puzzle, leading the authors to conclude that architecture matters more than scale.

  44. Ep. 61
    Pereira & Zuidema: Reasoning Models Build a Map of the Tower of Hanoi, Then Lose It — Fukai Reads
    2026-08-17

    An arXiv preprint by Devin Pereira and Willem Zuidema (University of Amsterdam). On the flat-to-flat Tower of Hanoi, reasoning models encode the board almost perfectly at the end of the prompt (rank correlation 0.935, nearest-state retrieval 1.00), yet that representation decays while they write out the moves — shown with linear probes and activation patching. Re-injecting the prompt-time representation lifted Qwen3.6-27B from 33/81 (41%) to 59/81 (73%) optimal solutions.

  45. Ep. 60
    Lu et al. : le flow est-il fait de « difficulté » ou de « l'effort investi » ? — lecture de Fukai
    2026-08-17

    Une étude évaluée par les pairs de Hairong Lu et de ses collègues (Psychological Research, 2025) sur le flow et l'effort mental. En manipulant séparément, dans une tâche de discrimination visuelle, la « difficulté » et les « chances perçues de réussir », la manipulation de la difficulté a produit un effet marqué (ηp²=0.64), tandis que la manipulation de l'attente n'a eu qu'un effet faible et seulement au niveau de l'essai (d=0.04) ; le U inversé du flow était à la limite de la significativité (p=0.053) et le P300 n'avait aucun lien avec le flow. Une étude exploratoire avec N=37.

  46. Ep. 59
    Li et al.: Measuring Whether AI Really Sees Shape, via Jigsaw Puzzles — Fukai Reads
    2026-08-15

    A paper by Shawn Li et al. introducing JigShape, a benchmark for spatial reasoning in vision-language models. Interlocking tab-and-blank pieces make the ground truth unique across 95,468 instances from 4x4 to 16x16; only GPT-5.5 beat chance zero-shot (69.65% on 4x4), everything collapses from 8x8 even after fine-tuning, and removing the shapes drops 97% to 10%.

  47. Ep. 58
    Randelshofer et al.: Fifteen UX Leaders in AAA Studios on Pre-Production Decisions — Fukai Reads
    2026-08-14

    A qualitative study by Ivana Randelshofer (Ubisoft Düsseldorf) and colleagues at the University of Waterloo HCI Games Group and elsewhere. Semi-structured interviews with 15 senior UX leaders from AAA studios (Blizzard, EA DICE, Guerrilla, Larian, Remedy, Ubisoft, and others) analysed with reflexive thematic analysis. Findings: pre-production decisions blend theory, experience and instinct; cross-functional structures (strike teams, competency teams) align player needs with production constraints; academic frameworks work best as discussion starters, not prescriptions. arXiv preprint posted August 2026, not peer reviewed. A rare look at what game production actually feels like from the UX side.

  48. Ep. 57
    Bazzaz and Cooper: Comparing Generative AI to PCG Across 500,000 Steam Reviews — Fukai Reads
    2026-08-14

    A paper by Bazzaz and Cooper at Northeastern analysing 508,192 Steam reviews. Comparing 5,970 titles that disclose generative-AI use against 5,186 titles that use PCG, the recommend rate is 86.3% for PCG versus 68.4% for generative AI — a 17.9-point gap — and generative-AI reviews split almost evenly at 53.0% positive to 47.0% negative. A thematic analysis of 600 reviews raises five themes: signals of low developer investment, ideological rejection, conditional acceptance, mismatch between disclosure and evidence, and criticism of not using AI where it should be. arXiv:2608.11539, ACM DOI 10.1145/3831347 assigned.

  49. Ep. 56
    Cai et al.: Bringing the Authoritative Server into Learned World Models — Fukai Reads
    2026-08-12

    A paper on multiplayer world models by Cai and eight colleagues at Alaya Lab, Peking University and Institute of Science Tokyo. It ports the authoritative-server contract of online games into a learned model, splitting it into a Logic Engine that advances a typed shared state and a Rendering Engine that draws each camera from it. On matched multiplayer Snake it reaches 0.764 state recovery against 0.128 for the best video-based baseline, with cross-view disagreement of 0.000 by construction, and advances 1,024 player entities for 10,000 ticks. arXiv preprint, submitted 6 August 2026, not yet peer-reviewed.

  50. Ep. 55
    Han et al.: Sorting Out When Learning Order Matters, by Computational Complexity — Fukai Reads
    2026-08-10

    A paper by Han and four colleagues at UC Davis and partner institutions on the computational complexity of instructional sequencing. They formalise the ordering of prerequisite-linked concepts as a stochastic shortest-path problem, prove that the stochasticity of retry-after-failure collapses exactly by dividing cost by success probability, show that optimal ordering nonetheless remains NP-hard, and give a cheap diagnostic that upper-bounds the value of sequencing before any optimisation. On 70,893 real interactions from an introductory CS course that headroom was under 0.2%, while on a constructed trap greedy sequencing lost 28.3-45.1%. arXiv preprint, submitted 5 August 2026, not peer reviewed.

  51. Ep. 54
    Honda et al.: Measuring Which Options Are Worth Trying by How Much Uncertainty They Remove — Fukai Reads
    2026-08-09

    A paper by Honda and five co-authors at the University of Tokyo proposing the B-EUR model, which formalises the value of trying a candidate option as the uncertainty about action-outcome relations it is expected to remove. Tested through simulation and human experiments (44 participants) on a graph-shape guessing task, the value of trying, enjoyment and choice frequency all followed an inverted U peaking at intermediate generalizability. Outcome discriminability affected choice behaviour but showed no significant effect on subjective ratings. An arXiv preprint posted 6 August 2026, not yet peer reviewed.

  52. Ep. 53
    Tarun Kumar S: What Happens When You Tell a Human-Move Predictor the Last 20 Moves and the Clock — Fukai Reads
    2026-08-08

    A paper by Tarun Kumar S of Peargent Labs on Otter, a chess AI that predicts human moves. Where earlier models treated each position independently, Otter conditions on the last 20 moves and on remaining clock time, reaching 55.23% top-1 accuracy with 15.3M parameters — 1.98 points above Maia 2. Of the +7.62 point gain over a board-only baseline, history contributes +5.24 and the clock +2.38. An arXiv preprint posted 5 August 2026, not yet peer reviewed.

  53. Ep. 52
    Geheeb et al.: Let an LLM Poke at Your Game Design Pillars — Fukai Reads
    2026-08-07

    A paper on game design pillars and LLMs by Julian Geheeb and colleagues at the Technical University of Munich. Pillars are heavily used in industry but almost unexamined academically; the paper gives them a formal definition and quality criteria, then hands structural checking, contradiction detection and feature evaluation to an LLM in a prototype called SPINE. A 42-hour game jam and interviews with four developers produce a consistent picture: useful at the moment of putting a pillar into words, thinning with each rewrite iteration, and unable to recognise deliberate juxtaposition when flagging contradictions. Peer-reviewed at FDG '26.

  54. Ep. 51
    Chen: Reconstruct the Persistent World First, Then Build Something Playable — Fukai Reads
    2026-08-06

    A narrative-to-game paper by Yi-Chun Chen. Before generating scenes or gameplay individually, it makes explicit reconstruction of a persistent world — entities, locations, relationships, evolving state — the central objective, maintained as one computational object shared across the pipeline. The prototype builds the world with GPT-5-mini plus constrained world completion and realises it as playable tile-based PyGame environments. An arXiv preprint offering qualitative feasibility across three cases, with no quantitative evaluation.

  55. Ep. 50
    Huang et al.: Letting an AI Play the Generated Game, Then Fix It — Fukai Reads
    2026-08-05

    A game-generation paper by Yixu Huang and colleagues (Fudan University, Xiaohongshu and others). Play2Code puts a screen-driving GUI agent into the generation loop as a playtester, evaluated on PlaytestArena, a new environment of 200 tasks and 1,548 rubric criteria. Averaged over three backbones, rubric pass-rate goes from 29.7% for single-pass generation and 52.2% for a code-inspection-only pipeline to 66.8%.

  56. Ep. 49
    Wu et al.: Rebuilding a Case Report Into a Chain of Decisions — Fukai Reads
    2026-08-04

    A medical-education gamification paper by Qian Wu and colleagues (CUHK and others). MedGame is a dual-engine framework that converts static case reports into a three-level Act / Scene / Decision Node script and then into a dependency graph of multimodal generation tasks. Fine-tuning on a 5,000-case benchmark lifts structural validity from 79.4% to 99.1%, while medical accuracy plateaus around 7 out of 10.

  57. Ep. 48
    Wang et al.: Making a Puzzle Solver the Teacher for Every Single Move — Fukai Reads
    2026-08-03

    A game-AI paper by Yu Wang and colleagues. Where long-horizon puzzles reward only the final win, they convert the drop in a solver's remaining-distance-to-goal into a per-move score and mix it into training. Averaged over Sokoban, Minesweeper and Rush Hour, success rises from 16.6% to 62.1%, and on unseen difficulty from 5.9% to 28.4%. Querying the solver costs about 73 parts per million of training wall clock.

  58. Ep. 47
    Ponnock & Ho: The Order of Mario 1-1 Has a Measurable Teaching Effect — Fukai Reads
    2026-08-02

    A reinforcement learning and level design paper by Jesse Ponnock and Lucas Ho (arXiv preprint, not peer-reviewed). Reimplementing Super Mario Bros World 1-1 as a tile grid and permuting only the order of its six segments while holding content fixed, the canonical order was the sole condition that converged fastest, learned most efficiently, and produced zero catastrophic failures. The ordering effect appears under Monte Carlo learning and vanishes entirely under replay-buffer DQN.

  59. Ep. 46
    Jeong et al. : à énigme identique, changer le mode de réponse change la difficulté — lu par Fukai
    2026-08-01

    Un article évalué par les pairs sur la charge cognitive et la conception de l'interaction, par Harim Jeong et ses collègues (JMIR Serious Games). En maintenant fixe le stimulus d'une tâche de Stroop sur tablette et en changeant uniquement les options de réponse — étiquettes textuelles remplacées par des pastilles de couleur — le taux de bonnes réponses est passé de 0,86 à 0,91 et le temps de réaction a diminué de 85,4 ms, chez 127 enfants de 6 à 12 ans. Les indicateurs neuronaux au niveau du cortex préfrontal n'ont montré aucune différence significative.

  60. Ep. 45
    Zhou et al.: The Verifier is the Curriculum — Training Game Generation on a Launch Check Alone — Fukai Reads
    2026-07-31

    A game-generation paper by Chenyu Zhou and colleagues. Starting from a diagnosis that the learned judge is gameable, they gate self-distillation on a single binary signal — does the generated Godot project launch cleanly — and over three rounds lift clean generation on four unseen families from 8.8% to 42.2%, with best-of-16 coverage going 18/25 to 25/25. Loosen the gate and the gain disappears.

  61. Ep. 44
    Gould & Ward et al. : mesurer la difficulté des puzzles par le « temps de résolution humain » — lu par Fukai
    2026-07-29

    Un article d'évaluation de l'IA par Gould, Ward et leurs collègues. En attribuant un « temps de résolution humain » à 43 bancs d'essai et plus de 30 000 problèmes, puis en mesurant le temps des tâches qu'un modèle réussit à 50 % sans exposer son raisonnement, ils constatent un doublement tous les 373 jours environ sur six ans, atteignant environ trois minutes pour GPT-5.5. Leur méthode de mesure de la difficulté, qui s'appuie notamment sur le sudoku et les mots croisés, est directement réutilisable en conception de puzzles.

  62. Ep. 43
    Li et al.: Rereading Video World Models as Game Engines — the Unsolved Problem Called State — Fukai Reads
    2026-07-28

    A survey of interactive world models by Zhen Li and colleagues. It reorganizes research on generating game worlds with video models along four dimensions drawn from the engine's action-state-observation loop, and argues that the remaining hard problems all revolve around explicit game state. It also contributes a 90+ hour Black Myth: Wukong dataset.

  63. Ep. 42
    Teo et al.: AI Assistants Overassist — Int-Bench Measures Intervention in Problem-Solving — Fukai Reads
    2026-07-27

    Teo et al. on LLM intervention behavior. Using Int-Bench, a simulated setting, they measure when and how much AI assistants help during problem-solving, finding AI intervenes earlier and more often than humans, tends to leak the answer, and does not improve transfer. A useful read for puzzle hint design.

  64. Ep. 41
    Halina & Guzdial: Generating Levels as a "Cake of Time" — Fukai Reads
    2026-07-26

    A procedural level-generation paper by Halina and Guzdial. It represents a level as a "cake" of board states stacked over time, and generates a level and its solution together with PRP, which recombines play traces. In Sokoban, against six existing methods, it reached 100% playability with high diversity, without hand-authored constraints or rewards.

  65. Ep. 40
    Hsu et al.: LLM-Voiced NPCs Make Players' Heads Heavier -- A 'Double-Edged Sword' Experiment — Fukai Reads
    2026-07-24

    An empirical LLM-NPC paper by Hsu et al. (Communication University of China and others). They built a scripted-NPC version and a GPT-4.1 LLM-NPC version of the same game and ran a between-subjects test with 130 players. LLM-NPCs significantly raised cognitive load (p<.001), did not significantly improve overall enjoyment (p=.195), and increased autonomy while lowering usability and trust.

  66. Ep. 39
    Wang et al.: Gauging Tetris Block Puzzle Difficulty by How Fast a Strong AI Learns — Fukai Reads
    2026-07-23

    An arXiv preprint from a National Yang Ming Chiao Tung University and Academia Sinica group that measures the difficulty of the popular mobile game Tetris Block Puzzle. It rates rule variants by how fast and high a strong AI (Stochastic Gumbel AlphaZero) can learn to play, finding that more holding/preview blocks make the game easier while adding block shapes makes it harder (the T-pentomino most of all).

  67. Ep. 38
    Johnson et al. : comment les jeux changent-ils lorsqu'on y intègre un grand modèle de langage — lu par Fukai
    2026-07-22

    Une étude qualitative de Johnson et ses collègues de l'Université de Calgary, portant sur le développement de deux jeux intégrant un LLM dans leur structure même. À partir des introspections des développeurs, l'étude analyse comment le gameplay, la jouabilité et l'expérience du joueur évoluent lorsque le LLM est intégré non comme un « ornement » mais comme un « composant ». Si la variété et la personnalisation augmentent, de nouvelles contraintes apparaissent — l'exactitude, la calibration de la difficulté et la cohérence — et l'étude rapporte que l'application stricte d'un schéma et la vérification en sont la clé.

  68. Ep. 37
    Earle et al.: Recasting Level Design from a One-Person Job to a Multi-Agent Collaboration — Fukai Reads
    2026-07-21

    A paper on reinforcement-learning level generation (PCGRL) by Earle et al. It recasts the traditional single-agent, tile-by-tile method as a multi-agent problem in which several agents divide the work and edit in parallel, showing across maze and dungeon domains that more agents improve generation quality, generalization to unseen boards, and computational efficiency.

  69. Ep. 36
    Bhaumik et al.: Stitching WFC and Reinforcement Learning for Playable, Good-looking Levels — Fukai Reads
    2026-07-20

    A procedural level generation paper by Bhaumik et al. It tackles the weaknesses of WFC (good-looking but unplayable) and reinforcement learning (playable but ugly) with WCRL, which narrows the RL agent's actions using WFC's local rules, generating Lode Runner levels that are both example-like and playable.

  70. Ep. 35
    Shyne et al.: How Far Do Puzzle Solver Loops Match Human Felt Difficulty — Fukai Reads
    2026-07-19

    A logic-grid-puzzle difficulty study by Shyne, Facey & Cooper. Using solver loops (the pass count of a human-style solver) as a difficulty proxy, they generate difficulty-varied puzzles with a quality-diversity algorithm and, in a 63-player study, show solver loops correlate significantly with subjective difficulty (c=0.30, p=0.015).

  71. Ep. 34
    Nath et al.: Training Game AI When Streaming Dirties the Video — Fukai Reads
    2026-07-18

    A paper by a Microsoft team (Nath et al.) on imitation-learning agents for streamed video games. It proposes streaming augmentations that artificially manufacture the temporally connected noise of cloud gaming and mix it into training. Even from five demonstrations, completion rises by up to ~40%, and performance loss under network lag drops from 49.82% to 7.45%.

  72. Ep. 33
    Ye et al.: Measuring Image-Capable AI (MLLMs) with Children’s Intelligence Tests — Fukai Reads
    2026-07-17

    A paper (arXiv preprint) by Hengwei Ye and colleagues at ShanghaiTech University on KidGym, an MLLM evaluation benchmark inspired by children’s intelligence tests (the Wechsler scales). It measures five abilities — Execution, Perception Reasoning, Memory, Learning, Planning — across 12 tasks on a 2D grid at three difficulty levels, evaluating nine models. Even top models reached only 0.30 on abstract-shape puzzles and 0.72 on counting against a human 1.00.

  73. Ep. 32
    Zeng et al.: Automating Game Balancing with LLM-vs-LLM Self-Play — Fukai Reads
    2026-07-16

    A paper by Zeng et al. on automated game balancing. It tackles balancing asymmetric strategy games by using multi-agent LLM self-play as an evaluator and Bayesian optimization to search rule parameters, reporting convergence to near-0% win-rate gaps on their own game, CivMini.

  74. Ep. 31
    Waugh : mesurer le raisonnement de l'IA avec le sudoku et le slitherlink — lu par Fukai
    2026-07-15

    Un article (préprint arXiv) de Justin Waugh d'Approximate Labs sur Pencil Puzzle Bench, un benchmark qui mesure le raisonnement des LLM à l'aide de puzzles papier-crayon. Sur 62 231 puzzles répartis en 94 types, 300 sont sélectionnés ; le cœur du dispositif est qu'une machine peut vérifier chaque coup au regard des règles. 51 modèles ont été évalués. Même le plus fort, GPT-5.2, n'atteint que 56,0 % en mode agentique, avec environ la moitié des puzzles non résolus.

  75. Ep. 30
    Ahn et al. : la difficulté d'un puzzle ne tient pas à son « apparence » mais à son « concept » — lecture de Fukai
    2026-07-14

    Un article de Caroline Ahn et de ses collègues de l'Université de Boston (prépublication arXiv) présente CogARC, une refonte pour l'humain du benchmark de raisonnement abstrait ARC. En enregistrant coup par coup les réponses de 260 participants au total à des puzzles en grille, l'étude montre que la difficulté ne dépend pas de la taille du plateau ni du nombre de couleurs, mais de la complexité conceptuelle des règles, et que même leurs erreurs convergent vers les mêmes réponses erronées.

  76. Ep. 29
    Luo et al. : la « manière d'aider » de l'IA compte autant que le contenu de l'aide — lecture de Fukai
    2026-07-13

    Un article de Yunhao Luo et de ses collègues de UC Santa Barbara sur la « manière d'aider » d'une IA à initiative mixte. À partir du puzzle Rush Hour, l'étude compare une aide à la demande déclenchée par un bouton et une aide automatique déclenchée par un minuteur d'inactivité : les performances sont quasi identiques, mais l'aide à minuteur obtient une évaluation nettement plus favorable de l'IA. Accepté à IUI '26.

  77. Ep. 28
    Ying et al.: Measuring AI's General Intelligence Through Every 'Human Game' — Fukai Reads
    2026-07-12

    A preprint from a team at MIT, Harvard and others that measures AI's general intelligence through games humans made. Rebuilding 100 popular App Store and Steam titles with an LLM and having seven frontier vision-language models play them, the best reached only 8.5 against a human median of 100, falling far short on memory, planning and inferring rules.

  78. Ep. 27
    Triebel et al.: Does AI Have Both a Head and a Hand on a Classic Physics Puzzle? — Fukai Reads
    2026-07-11

    A paper by Triebel et al. evaluating VLMs on the classic physics puzzle The Incredible Machine 2. Using VLATIM, a five-stage benchmark, it asks whether screen-operating AI can solve problems like humans; the cleverer large models can plan but cannot click precisely, and no model solved even one puzzle to completion.

  79. Ep. 26
    Nasvytis et Fan : l'insight et le « transfert » se manifestent dans la manière de parler — lu par Fukai
    2026-07-10

    Un article de Nasvytis et Fan, de Stanford, qui saisit l'insight et le transfert à travers la verbalisation de la pensée. Cent quatre-vingt-neuf participants ont résolu cinq énigmes d'équations d'allumettes ; le groupe qui répétait le même type est devenu plus rapide et plus précis après sa première bonne réponse (taux de réussite de 0,75 à l'essai 5), et la proportion de personnes verbalisant le type du problème a été multipliée par environ sept. On peut lire que le signe du transfert est « la capacité à mettre l'astuce en mots ».

  80. Ep. 25
    Li et al.: Making Geometry Problem Solving Verifiable with a Solver as Referee — Fukai Reads
    2026-07-09

    An arXiv preprint by Can Li et al. on geometry problem solving (GPS). Their SD-GPS translates diagram-and-text problems into a form a symbolic solver can execute, and at impasses proposes helper lemmas verified by the solver itself. The abstract reports it consistently outperforms existing methods on Geometry3K and PGPS9K. Fukai reads it for its use in solvability-guaranteed puzzle generation.

  81. Ep. 24
    Sestini et al.: Making AAA Game NPCs Feel Authentic with Reinforcement Learning — Fukai Reads
    2026-07-08

    A vision paper from the research team at Electronic Arts. It tests whether AAA game NPCs can be improved with reinforcement learning, through two real cases — goalkeeper positioning in EA SPORTS FC 25 and infantry locomotion in Battlefield 6 — and lays out seven requirements RL must meet in production. Its conclusion: RL is a tool to augment, not replace, existing game AI.

  82. Ep. 23
    Xu et al. : quand l'IA générative devient le « cœur du jeu » — Fukai décrypte une étude sur les jeux nativement IA
    2026-07-07

    Un article de synthèse (prépublication arXiv) signé par Zhiyue Xu et cinq coauteurs, consacré aux « jeux nativement IA », où l'IA générative constitue la boucle centrale elle-même. Le concept est défini par un contrefactuel — le jeu tiendrait-il debout si l'on retirait l'IA ? — et 53 œuvres réelles sont classées selon deux axes : le type de jeu (G) et le mécanisme d'IA dominant (N). L'étude montre un biais marqué vers les genres narratifs et un usage encore ténu de l'IA pour arbitrer les règles.

  83. Ep. 22
    Wermann et al.: How In-Game AI 'Words' vs 'Demonstration' Change Learning and Cognitive Load — Fukai Reads
    2026-07-06

    A pre-registered experiment by LMU Munich and colleagues comparing 'verbal' and 'demonstration' support from an in-game AI NPC. Splitting 152 people into three groups in Qookies, a quantum-technology learning game, they found no difference in learning gains between conditions, but the verbal-plus-visual group reported significantly lower intrinsic cognitive load than the verbal-only group (d=0.60).

  84. Ep. 21
    Aryan et al.: When You Stall, the World Changes — AbideGym Turns Static RL Worlds into Adaptivity Tests — Fukai Reads
    2026-07-05

    A preprint by Aryan et al. (Abide AI) on RL environment design. To fight the brittleness that comes from training in fully static worlds, AbideGym rewrites the rules and grows the map mid-episode, triggered by the agent's own inactivity, forcing it to abandon memorized policies and re-plan. The paper presents the design and a comparison to prior work; no experimental results yet.

  85. Ep. 20
    Wang et al. : un agent LLM qui lit à quel point la tête est « occupée » à partir du regard — lu par Fukai
    2026-07-04

    Un article de Meta Reality Labs et al. proposant d'estimer la charge cognitive (à quel point la tête est occupée) à partir des données de regard. Face à la faible généralisation et à la difficile interprétabilité des méthodes existantes, les auteurs structurent le regard et le font raisonner par un LLM, muni de contexte, de différences individuelles et d'exemples, via le cadre GazeMind, qui rapporte une précision de 62,73 % en classification à 3 niveaux (plus de 20 points au-dessus des méthodes existantes).

  86. Ep. 19
    Mirowski et al. : faire passer le récit de « l'écrire » au « le trouver » — Fabula, une IA d'écriture cultivée avec une communauté d'auteurs — lu par Fukai
    2026-07-03

    Un article sur Fabula, une IA d'aide à l'écriture développée par Google DeepMind. Son « gestionnaire de drame », qui planifie et génère un récit de façon hiérarchique, a été façonné de manière critique avec 42 experts en co-conception participative : le système s'est révélé solide pour la structure, mais faible sur le style et la surprise. Fukai en tire des enseignements qui s'appliquent directement au récit interactif dans le jeu vidéo.

  87. Ep. 18
    Özkan : entraîner ensemble l'IA qui génère les niveaux et l'IA qui les résout — lu par Fukai
    2026-07-02

    Un article de Miraç Buğra Özkan qui fait apprendre simultanément, par apprentissage par renforcement, la génération de niveaux et leur résolution. Sur Unity, un colibri (le solveur) et une île flottante (le générateur) apprennent en observant mutuellement leurs résultats, atteignant environ 90,2 % de réussite sur 100 configurations inédites.

  88. Ep. 17
    Liu et al.: More Memory Makes AI Agents Less Cooperative — Fukai Reads
    2026-07-01

    An arXiv paper from a Carnegie Mellon-led team studying how an LLM agent's memory length affects cooperation. Across 7 models, 4 repeated social-dilemma games, history windows up to 80 rounds and 500-round matches, longer history degrades cooperation in 18 of 28 settings — a 'memory curse.' The cause is the content of accumulated defection records, not context length, and forward-looking reasoning partly fixes it.

  89. Ep. 16
    Feng et al. : les agents LLM peuvent-ils bien marchander dans un jeu commercial ? — Fukai lit
    2026-06-30

    Le benchmark SidConArena d'une équipe de l'Université Tsinghua, conçu pour évaluer les agents LLM dans un jeu commercial coopératif-et-compétitif. Construit sur le jeu de plateau Sidereal Confluence, il évalue les agents sur trois phases (négociation, production, enchère à pli cacheté) et constate que les modèles de pointe sont plus forts mais continuent à mal évaluer les ressources, à négocier passivement et à mal planifier sur de longues durées.

  90. Ep. 15
    Bazzaz et al. : Penser que c'est de l'IA change l'experience — lu par Fukai
    2026-06-29

    Un article CHI '26 de Bazzaz et Cooper sur les biais de perception du contenu genere. 142 participants ont joue a des niveaux melanges dans Super Mario Bros. et Sokoban. Les joueurs ne parviennent presque pas a identifier l'auteur, mais les niveaux qu'ils croient etre generes par IA sont juges moins amusants, plus difficiles et plus frustrants.

  91. Ep. 14
    Liu et al.: AI Assistance Erodes Persistence — A Warning for Hint Design — Fukai Reads
    2026-06-28

    A paper by Grace Liu and colleagues on how AI assistance affects independent problem-solving and persistence. Across RCTs with 1,222 participants, AI raised in-session performance but, once removed, left people solving less and giving up more. Those who got direct answers declined most while hint-users did not, a result that speaks directly to game hint design.

  92. Ep. 13
    Jara Gonzalez & Guzdial: Generating Enemy Shapes as Gates You Need a Mechanic to Beat — Fukai Reads
    2026-06-27

    A paper by Jara Gonzalez and Guzdial on generating enemy morphologies (collision shapes). They frame 'enemies defeatable only with a specific mechanic' as a 4x4 grid generation problem, compare reinforcement learning, A* search, and neural generation, and find a simple A* reachability rule yields the best gating and most diverse shapes at the lowest cost.

  93. Ep. 12
    Munk et al.: Generating Dynamic Game Text with Small Language Models — Fukai Reads
    2026-06-25

    A paper by Munk et al. (IT University of Copenhagen) on generating in-game text dynamically with small language models (SLMs). It tackles the offline, cost and consistency walls of cloud LLMs using small models aggressively fine-tuned for narrow jobs. Their proof of concept, DefameLM, runs a medieval-RPG smear-poster loop, showing a one-billion-parameter-class model reaches high quality in a few seconds on a consumer PC.

  94. Ep. 11
    Zeytuncu : la difficulté d'un puzzle est déterminée par « le nombre de chiffres utilisés » — lu par Fukai
    2026-06-24

    Article de Yunus E. Zeytuncu sur la modélisation de la difficulté dans les puzzles arithmétiques entiers (puzzles de type Numbers où l'on combine des chiffres pour atteindre un nombre cible). Un solveur exact génère environ 3,47 millions d'instances, la difficulté est définie par le nombre minimal d'opérations, et l'article démontre que le seul nombre de chiffres utilisés dans la solution minimale constitue une « statistique suffisante minimale » permettant de prédire parfaitement la difficulté.

  95. Ep. 10
    Chao et al.: Insight Is About Searching Far — Fukai Reads
    2026-06-23

    A paper on insightful problem-solving by Chao, Hsieh & Wu. Using a Japanese RAT and a simulation to quantify the search path to a solution, it shows that de-fixation is necessary for solving but is not what determines insight; the hallmark of insight is exploring the solution space over greater distances.

  96. Ep. 9
    Monti et al.: Measuring AI's Planning Power on a Single-Corridor Sokoban — Fukai Reads
    2026-06-22

    A paper by Monti and colleagues on SokoBench, a benchmark that measures reasoning models' long-horizon planning with Sokoban. By lining up only single-box straight corridors and narrowing difficulty to a single axis (corridor length), it shows that even state-of-the-art reasoning models break down once more than 25-30 moves of lookahead are needed. The authors locate the cause in accumulated miscounting.

  97. Ep. 8
    Luo et al. : Les agents IA peuvent-ils créer un jeu jouable de bout en bout dans un moteur réel ? — Lu par Fukai
    2026-06-21

    La base d'évaluation GameCraft-Bench de Luo, Wang et al. mesure si des agents de programmation peuvent générer des jeux de bout en bout. Le benchmark demande aux agents de créer, à partir de spécifications en langue naturelle, un jeu jouable sur Godot, et l'évalue selon trois critères : lancement, replay d'actions et notation vidéo — 140 tâches, 15 genres. Même la meilleure configuration n'atteint que 41,46 %, montrant que les agents peuvent créer des mécaniques, mais pas encore des produits finis dotés de profondeur, de lisibilité et de finition.

  98. Ep. 7
    Li et al. : AutoBG, une IA qui accompagne la conception de jeux de plateau de l’idée à l’achèvement — lu par Fukai
    2026-06-20

    Article (preprint arXiv) sur AutoBG, une IA d’assistance à la conception de jeux de plateau par Zizhen Li et al. L’ensemble du processus de conception, de l’idée à la génération du livret de règles et aux retours individualisés, est traité par une Verifier-Gated Iteration séparant rôle de génération et rôle d’évaluation. BG-Critic, le module évaluateur, dépasse GPT-5.4 en qualité de diagnostic, selon les auteurs.

  99. Ep. 6
    Nasir et al. : faire évoluer les « règles mêmes » du jeu — Fukai lit MORTAR
    2026-06-19

    Article de Nasir, Togelius et al. sur la conception automatique de jeux. MORTAR propose de faire évoluer les « mécaniques (règles du jeu) » elles-mêmes par un algorithme de qualité-diversité et un grand modèle de langage, en mesurant la qualité par les victoires et défaites d'IA de niveaux différents. GPT-4o-mini génère des jeux variés et jouables, et la contribution de chaque mécanique est quantifiée.

  100. Ep. 5
    Jiang et al. : peut-on créer un « jeu jouable » rien qu'avec des mots ? — Fukai lit OpenGame
    2026-06-18

    Article de Yilei Jiang et al. (Université chinoise de Hong Kong) présentant OpenGame, un agent open source qui génère automatiquement des jeux web 2D complets à partir de descriptions en langage naturel. En s'appuyant sur un squelette réutilisable et un « manuel de débogage vivant », il limite les erreurs d'intégration et atteint l'état de l'art sur 150 tâches. Les puzzles restent néanmoins le genre le plus difficile à traiter.

  101. Ep. 4
    McConnell & Zhao : générer en temps réel des puzzles au niveau de difficulté « idéal » par algorithme génétique — Lu par Fukai
    2026-06-17

    Article de McConnell et Zhao sur la génération adaptative de puzzles à l'aide d'un algorithme génétique. Des puzzles de tracé de chemin proches de Cosmic Express sont générés en temps réel en environ 7 secondes par puzzle, adaptés à un modèle de joueur enregistrant la façon de résoudre de chaque joueur. Une expérience avec 18 participants montre que la version reposant « uniquement sur le temps » est inférieure en termes de difficulté perçue et de sentiment de progression.

  102. Ep. 3
    Li et al. : Les LLM peuvent-ils « jouer et gagner » à des jeux 2D ? — Fukai lit GVGAI-LLM
    2026-06-16

    L'article GVGAI-LLM de Li et al. (NYU et autres). Il propose un benchmark qui fait jouer des modèles de langage à 118 jeux 2D pour mesurer leur capacité de raisonnement et de compréhension spatiale. En traduisant les grilles en cartes ASCII et en résolvant en zero-shot, GPT-4o-mini a obtenu un taux de victoire de 0 % sur 477 des 540 niveaux, et un taux global de 10,27 %, restant en deçà des algorithmes de recherche classiques. Je l'analyse dans l'ordre : problème, méthode, résultats, utilisations, limites.

  103. Ep. 2
    Kar : vérifier en temps réel si un niveau généré est praticable — lecture par Fukai
    2026-06-14

    Article de Rishabh Kar (King's College London) sur la PCG. Il propose Momentum, un système qui valide la praticabilité des niveaux générés dans la même boucle de jeu, sans jamais stopper l'exécution. Deux agents autonomes courent devant le joueur et inspectent la route en avance, l'un par détection géométrique aérienne, l'autre par NavMesh au sol. L'évaluation repose sur une estimation structurelle tirée du code.

  104. Ep. 1
    Xu et al. : Élever les « mécanismes » du jeu au rang de coordonnées pour générer automatiquement des niveaux solubles — Fukai lit
    2026-06-13

    Un article de Xu et Verbrugge de l'Université McGill sur la PCG. Face aux méthodes traditionnelles qui donnent la priorité au terrain, ils proposent HDPCG : un cadre qui effectue des recherches de chemin sur un graphe à dimensions étendues où les « mécanismes » tels que l'inversion de gravité et les sols mobiles sont élevés au rang de coordonnées, garantissant la solubilité pendant la génération. Les niveaux ont été reproduits sous forme jouable dans Unity.