TAG
#arxiv
0 篇评论 · 4 篇随笔
相关随笔
How do you measure a "good mechanic"? A paper on automatic game design, and a talk on modelling puzzles as constraint problems
Two pieces today: a preprint paper and a talk from last autumn's puzzle-game developer conference. First, I read in full "MORTAR: Evolving Mechanics for Automatic Game Design" (arXiv, submitted 31 December 2025) by researchers at the University of the Witwatersrand and New York University. It evolves a game's underlying rules and interactions — its "mechanics" — using a quality-diversity algorithm plus an LLM, then measures whether stronger AI agents consistently beat weaker ones (a "skill gradient") via Kendall's Tau. Second, I looked at Alastair Aitchison's (Playful Technology) talk "The Rules of the Game: Modelling Puzzles as Constraint Satisfaction Problems" from ThinkyCon 2025 (November 2025), which models puzzles as constraint satisfaction problems and cites recent games like Lingo, Blue Prince, and Is This Seat Taken? Both pieces try to bring external, measurable structure to design work that usually stays intuitive.
让 LLM 造出一整款游戏,再让 AI 去试玩——ScriptDoctor 呈现的自动游戏设计现状
今天一篇。我通读了 NYU 的 Sam Earle、Julian Togelius 等人的论文《ScriptDoctor: Automatic Generation of PuzzleScript Games via Large Language Models and Tree Search》(英文,arXiv:2506.06524,投稿至 IEEE Conference on Games 的短论文)原文。他们把 PuzzleScript——由 increpare(Stephen Lavelle)创造、专用于 2D 网格回合制解谜游戏的描述语言——当作"模式生物",让 LLM 生成包含规则、美术与关卡的一整款游戏,并借助编译器报错与宽度优先搜索(BFS)试玩代理的反馈反复修正。给它几款人类制作的游戏作范例,产出质量明显提升;推理模型(o1、o3-mini)优于 GPT-4o。但最深的启示在失败一侧:看似最复杂的游戏,往往只是因为机制"坏掉了"才复杂——可解并不等于好玩。
"最强的玩家"并非"最好的测试者"——用 LLM 测量游戏难度的框架揭示的悖论
今天只有一篇。我通读了 Adobe Research 的 Chang Xiao 与哥伦比亚大学的 Brenda Z. Yang 合著的论文《LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents》(英文,arXiv:2410.02829)原文。这项研究探讨能否让现成的 LLM 游玩游戏,并将其成绩用作难度的代理指标,在 Wordle(猜词解谜)与 Slay the Spire(卡牌构筑 roguelike)上进行了验证。核心发现颇为悖论:LLM 的游玩水平不及普通人类,但"哪些关卡更难"这一相对难度,却与人类数据高度相关。更进一步,一个信息论意义上接近最优的 Wordle 求解器(比人类用更少的步数解出)却与人类感知的难度几乎不相关。也就是说,"解得最强的一方"并不等于"最好的难度测试者"。对于思考如何验证难度曲线的设计者而言,这是一篇启发颇多的论文。
「可解」与「线索可见」——GenEscape 所言明的密室逃脱设计二条件
今日一篇。通读了美国华盛顿大学 Mengyi Shan、Brian Curless、Ira Kemelmacher-Shlizerman、Steve Seitz 的论文《GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles》(英文,arXiv:2506.21839)原文。这是一项让文字→图像模型将密室逃脱谜题以"图像"形式生成的研究,但值得关注的是它将设计论的核心切分为两个条件——谜题须(1)可解(物体的可供性构成合理的行动序列),(2)具备引导玩家通向该解法的充分视觉线索。作者们让 Designer / Player / Examiner / Builder 四个智能体反复迭代,尤其是 Examiner 逐一消除"意外捷径"。虽是 AI 研究的形式,但将设计者在测试游玩时通常进行的作业明文化,可作为解谜设计的议论来阅读。

