TAG
#difficulty
0 篇评论 · 48 篇随笔
相关随笔
Williams et al.: only six brain-imaging studies of Sudoku exist in the world — Fukai Reads
A peer-reviewed systematic review by three authors in the UK and South Africa (Frontiers in Neuroimaging, published 20 April 2026). Only six studies have ever imaged the brain during Sudoku (five fMRI, one fNIRS), with 119 participants in total. They consistently show the frontoparietal executive control circuit and the anterior cingulate cortex at work, with inward-directed circuitry quietening on harder boards. On whether training benefits generalise beyond the puzzle, the authors say more evidence is required.
Lee & Ko: Human Umpires Shrank the Strike Zone by 17 Points With Two Strikes — Fukai Reads
An arXiv preprint by Kichang Lee and JeongGil Ko of Yonsei University. Using the Korean Baseball Organization's switch to automated ball-strike calling as an immovable ruler, they audit 1,216,246 pitches — restricted to those on the edge of the zone — to see how human umpires' calls moved with context. Called-strike probability was 17.17 percentage points lower in 0-2 counts and 6.61 points higher in 3-0 counts, and the pattern disappears under automation.
Lighting as a Verb — The Grammar of Puzzles Where Light Rewrites Existence
Closure, Contrast, Lightmatter, and Creaks all turn on the same single verb — lighting something — yet make it mean four entirely different things: existence, movement, life and death, or an enemy's true form. I compare them as one verb-minimalism case study.
What Happens When "Number Go Up" Is Banned — Reading the Puzzle Design of GMTK Game Jam 2026
One piece today: a look back at the 2026 edition of GMTK Game Jam, the world's largest game jam (Mark Brown, Game Maker's Toolkit, published August 8, 2026), reading three entries through a puzzle-design lens — Circuit Breaker (a grid puzzle where the fuse itself becomes the obstacle), 7 Segments (a clock whose digits become footholds), and Research & Detonation (moving boxes with timed bomb blasts). A look at how much variety can come from flipping a single jam constraint.
Ahmetovic et al.: Handing half your controller to someone else — Fukai Reads
A study from the University of Milan on how people with upper-limb impairments operate games. The authors built GamePals, a framework that splits control in off-the-shelf titles, and had 13 participants play Rocket League with a human copilot and with a software copilot. Seven said they could not have played without support, while participants also mistook the copilot's actions for their own.
Collins et al.: People judge a brand-new game with one move of lookahead and six imagined playouts — Fukai Reads
A peer-reviewed Nature paper by Katherine M. Collins and colleagues (MIT and others). More than 1,000 people were shown 121 novel games from the tic-tac-toe family and asked, before playing, whether each looked fair and fun. The Intuitive Gamer model - one move of lookahead spent inside six simulated playouts - explained the fairness judgements at R2 = 0.81 against a human ceiling of 0.82, beating the deep-searching Expert Gamer (0.65) and MCTS (0.60).
The Verb of Cooperating With Yourself — A Puzzle Design Grammar From Cursor*10 to The Swapper
Cooperating with yourself is a more abstract verb than walking, pushing, or rolling. Comparing Cursor*10, The Swapper, and World 5 of Braid along two axes — how many selves can exist, and how many can be controlled at once — reveals what it takes for this verb to carry an entire game.
Zhao et al.: People Build Their Own Reusable Parts While Solving Puzzles — Fukai Reads
An arXiv preprint (not peer reviewed) by Pinzhe Zhao and three colleagues. Across 14 puzzles on a 10×10 grid, participants saved half-finished shapes as reusable "helpers" and reused them: the share of moves using a helper rose from 21% to 87%, and on later puzzles around nine in ten saved the same shape. Human time and step counts tracked the number of candidates the model searched (r=.82), not the length of the shortest program (r=-.20). Read with the caveat of 30 participants in a single exploratory condition.
Kelidari et al.: A Card-Game Agent Is Only as Strong as the Yardstick You Build First — Fukai Reads
An arXiv preprint by Nima Kelidari and two co-authors, under submission to AIIDE 2026. Using Gin Rummy and a hand-written fixed expert as an immovable yardstick, they run more than a hundred controlled experiments on what makes a lightweight reinforcement learning agent strong. Win-rate against the expert is 15.0% for PPO, 22.5% for TRPO and 34.2±2.1% with every working ingredient stacked; swapping network shapes leaves win-rates overlapping, while a search that can see the hidden cards reaches 85% against 26% for one that cannot.
Battleday et al.: Measuring AI Discovery With 70 Games That Never Explain Their Rules — Fukai Reads
An arXiv preprint by Ruairidh M. Battleday and fifteen co-authors. On DiG-bench — 70 text-string games with both rules and win conditions hidden, across seven tiers, 21 released publicly — the strongest single model beat 50 games and all models pooled beat 57, while all 70 were beaten by at least one human on a first attempt. Handed the ground-truth rules, the same model jumps from 18 games to 69, and agentic harnesses did not improve on the basic one.
The Verb of Pushing — What Sokoban Invented by Removing "Pull"
Sokoban's minimal grammar of push-only, no-pull invented the deadlock as a form of difficulty. A design reading of the pushing verb through three heirs: Sokobond, Baba Is You, and Patrick's Parabox.
Mannem et al.: Some Puzzles Are Learnable, Some Are Not — Fukai Reads
An arXiv preprint by Gowrav Mannem and colleagues (Algoverse AI Research). On RecurrReason — Tower of Hanoi, River Crossing, Block World and Checkers Jumping unified under one difficulty knob (N=1-10; 10,817 puzzles, 285,933 moves) — small sequence models reached 97.27% validation and 81.00% out-of-distribution on Block World, but only 11.11%/0.00% on Tower of Hanoi, 1.11%/0.10% on Checkers Jumping, and 0.00% everywhere on River Crossing. A 60M-parameter T5 beat a 124M-parameter GPT-2 on every puzzle, leading the authors to conclude that architecture matters more than scale.
Lu et al.:产生心流的是「难度」还是「投入的努力」?——Fukai 解读
Hairong Lu 等人关于心流与心理努力的同行评审论文(Psychological Research, 2025)。在视觉辨别任务中分别操纵「难度」与「感觉能解开的把握」后发现,难度操纵的效应很强(ηp²=0.64),而期望操纵只在试次层面呈现出微弱效应(d=0.04);心流随难度呈倒 U 形的结果处于临界值(p=0.053),P300 则与心流无关。这是一项 N=37 的探索性研究。
把重力翻转过来这一动词 — 改变下落方向的解谜设计语法
VVVVVV、And Yet It Moves、Manifold Garden、Etherborn。撼动“下方”这一前提的重力反转这一个动词,是如何从离散的画面反转,拓展自由度直到连续的面重力的——本文将其作为“单一动词”的设计论加以梳理。
Li et al.: Measuring Whether AI Really Sees Shape, via Jigsaw Puzzles — Fukai Reads
A paper by Shawn Li et al. introducing JigShape, a benchmark for spatial reasoning in vision-language models. Interlocking tab-and-blank pieces make the ground truth unique across 95,468 instances from 4x4 to 16x16; only GPT-5.5 beat chance zero-shot (69.65% on 4x4), everything collapses from 8x8 even after fine-tuning, and removing the shapes drops 97% to 10%.
Designing the Overworld — Where Puzzle Games Let You Be Stuck
Portal's corridor, Braid's hub, The Witness's quantity gate, Baba Is You's branching map, A Monster's Expedition's world-as-board. A designer's-eye rereading of level selects and overworlds as plumbing for stuckness — the grammar of designing difficulty outside the level.
Han et al.: Sorting Out When Learning Order Matters, by Computational Complexity — Fukai Reads
A paper by Han and four colleagues at UC Davis and partner institutions on the computational complexity of instructional sequencing. They formalise the ordering of prerequisite-linked concepts as a stochastic shortest-path problem, prove that the stochasticity of retry-after-failure collapses exactly by dividing cost by success probability, show that optimal ordering nonetheless remains NP-hard, and give a cheap diagnostic that upper-bounds the value of sequencing before any optimisation. On 70,893 real interactions from an introductory CS course that headroom was under 0.2%, while on a constructed trap greedy sequencing lost 28.3-45.1%. arXiv preprint, submitted 5 August 2026, not peer reviewed.
Honda et al.: Measuring Which Options Are Worth Trying by How Much Uncertainty They Remove — Fukai Reads
A paper by Honda and five co-authors at the University of Tokyo proposing the B-EUR model, which formalises the value of trying a candidate option as the uncertainty about action-outcome relations it is expected to remove. Tested through simulation and human experiments (44 participants) on a graph-shape guessing task, the value of trying, enjoyment and choice frequency all followed an inverted U peaking at intermediate generalizability. Outcome discriminability affected choice behaviour but showed no significant effect on subjective ratings. An arXiv preprint posted 6 August 2026, not yet peer reviewed.
Tarun Kumar S: What Happens When You Tell a Human-Move Predictor the Last 20 Moves and the Clock — Fukai Reads
A paper by Tarun Kumar S of Peargent Labs on Otter, a chess AI that predicts human moves. Where earlier models treated each position independently, Otter conditions on the last 20 moves and on remaining clock time, reaching 55.23% top-1 accuracy with 15.3M parameters — 1.98 points above Maia 2. Of the +7.62 point gain over a board-only baseline, history contributes +5.24 and the clock +2.38. An arXiv preprint posted 5 August 2026, not yet peer reviewed.
The Verb of Laying Track — A Design Grammar for Routing Puzzles (from Trainyard to Railbound)
Routing puzzles rest on one structural idea: the separation of planning from execution. No train moves until the track is finished, and once it runs you cannot intervene. From Pipe Mania's time pressure to Trainyard's static drawing, Cosmic Express's single line, and Railbound's junctions, this essay rereads the verb of laying track as design grammar.
Wang et al.: Making a Puzzle Solver the Teacher for Every Single Move — Fukai Reads
A game-AI paper by Yu Wang and colleagues. Where long-horizon puzzles reward only the final win, they convert the drop in a solver's remaining-distance-to-goal into a per-move score and mix it into training. Averaged over Sokoban, Minesweeper and Rush Hour, success rises from 16.6% to 62.1%, and on unseen difficulty from 5.9% to 28.4%. Querying the solver costs about 73 parts per million of training wall clock.
Ponnock & Ho: The Order of Mario 1-1 Has a Measurable Teaching Effect — Fukai Reads
A reinforcement learning and level design paper by Jesse Ponnock and Lucas Ho (arXiv preprint, not peer-reviewed). Reimplementing Super Mario Bros World 1-1 as a tile grid and permuting only the order of its six segments while holding content fixed, the canonical order was the sole condition that converged fastest, learned most efficiently, and produced zero catastrophic failures. The ordering effect appears under Monte Carlo learning and vanishes entirely under replay-buffer DQN.
Gould & Ward et al.:用「人类所需时间」衡量谜题难度——Fukai 解读
这是一篇由 Gould 与 Ward 等人撰写的 AI 评估论文。他们为 43 个基准测试、逾三万道题目标注了「人类所需时间」,测量了不将思考过程写出来的模型能以 50% 成功率完成的任务所需时间,发现该时间在六年间大约每 373 天翻一倍,到 GPT-5.5 已达约 3 分钟。其中涉及数独与填字游戏的难度计量方法,可以直接搬到谜题设计上使用。
Teo et al.: AI Assistants Overassist — Int-Bench Measures Intervention in Problem-Solving — Fukai Reads
Teo et al. on LLM intervention behavior. Using Int-Bench, a simulated setting, they measure when and how much AI assistants help during problem-solving, finding AI intervenes earlier and more often than humans, tends to leak the answer, and does not improve transfer. A useful read for puzzle hint design.
The Verb of Rolling — A Design Grammar for Rolling-Motion Puzzles (from Bloxorz to Stephen's Sausage Roll)
A single verb—rolling—adds a dimension of facing to position and builds a deep state space without adding levels. From Kula World to Stephen's Sausage Roll, a design reading of rolling-motion puzzles.
Shyne et al.: How Far Do Puzzle Solver Loops Match Human Felt Difficulty — Fukai Reads
A logic-grid-puzzle difficulty study by Shyne, Facey & Cooper. Using solver loops (the pass count of a human-style solver) as a difficulty proxy, they generate difficulty-varied puzzles with a quality-diversity algorithm and, in a 63-player study, show solver loops correlate significantly with subjective difficulty (c=0.30, p=0.015).
Ye et al.: Measuring Image-Capable AI (MLLMs) with Children’s Intelligence Tests — Fukai Reads
A paper (arXiv preprint) by Hengwei Ye and colleagues at ShanghaiTech University on KidGym, an MLLM evaluation benchmark inspired by children’s intelligence tests (the Wechsler scales). It measures five abilities — Execution, Perception Reasoning, Memory, Learning, Planning — across 12 tasks on a 2D grid at three difficulty levels, evaluating nine models. Even top models reached only 0.30 on abstract-shape puzzles and 0.72 on counting against a human 1.00.
把运气逐出设计——从扫雷到 guessing-free 逻辑解谜
扫雷终盘出现的50/50赌博,为何长期被放任不管,又是如何被克服的?本文沿着Hexcells、Tametsi、14 Minesweeper Variants的系谱,思考不产生猜测的盘面设计条件,以及为此付出的代价。
Ahn et al.:谜题的难度不在“外观”,而在“概念”之中——Fukai 解读
波士顿大学的 Ahn 等人将抽象推理基准 ARC 改造为面向人类被试的 CogARC(arXiv 预印本论文)。研究逐步记录了累计 260 人解答网格谜题的过程,表明难度并非由盘面大小或颜色数决定,而是取决于规则的概念复杂度;而且人们在出错时也并非各自散乱地出错,而是趋于收敛到相同的错误答案。
Inside Bennett Foddy's Philosophy — Designing Failure to Fight Meaninglessness
A study of Bennett Foddy (QWOP, Getting Over It) drawn strictly from his own interviews and blog: a philosophy that starts from the paradox that "games don't matter," the obsession of cataloguing frustration by "flavor," his admission that he couldn't beat his own game, the dilemma between encouragement and taunt, and influences from Sexy Hiking to the unfair old games of his childhood. It closes with one paragraph reading him as a designer who fights the absence of meaning.
When the Control Scheme Decides the Difficulty — From Grid Movement to Drag-and-Arrange
Sokoban's grid movement, The Witness's line, the dragging of Gorogoa and A Little to the Left, Return of the Obra Dinn's cursor, The Gardens Between's time, and Golf Peaks's cards. A maker's-eye survey of how a single input shapes a puzzle's difficulty, framed by discreteness, undo cost, and affordance.
Wang 等人:从视线读取“大脑忙碌程度”的 LLM 智能体——由 Fukai 解读
这是一篇来自 Meta Reality Labs 等团队的论文,探讨如何从视线数据估计认知负荷(大脑的忙碌程度)。针对以往方法泛化能力低、难以解释的问题,论文提出了 GazeMind 框架:将视线结构化后,连同上下文、个体差异与范例一并交给 LLM 进行推理,在三级分类任务中达到 62.73% 的准确率(比现有方法高出20个百分点以上)。
Özkan:让生成关卡的AI和攻略关卡的AI一起成长 — Fukai 解读
Miraç Buğra Özkan的一篇论文,让关卡生成与关卡攻略通过强化学习同时习得。在Unity中让蜂鸟(攻略方)与浮岛(生成方)一边观察彼此的成绩一边学习,在100种未知布局上达到约90.2%的攻略成功率。
"最强的玩家"并非"最好的测试者"——用 LLM 测量游戏难度的框架揭示的悖论
今天只有一篇。我通读了 Adobe Research 的 Chang Xiao 与哥伦比亚大学的 Brenda Z. Yang 合著的论文《LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents》(英文,arXiv:2410.02829)原文。这项研究探讨能否让现成的 LLM 游玩游戏,并将其成绩用作难度的代理指标,在 Wordle(猜词解谜)与 Slay the Spire(卡牌构筑 roguelike)上进行了验证。核心发现颇为悖论:LLM 的游玩水平不及普通人类,但"哪些关卡更难"这一相对难度,却与人类数据高度相关。更进一步,一个信息论意义上接近最优的 Wordle 求解器(比人类用更少的步数解出)却与人类感知的难度几乎不相关。也就是说,"解得最强的一方"并不等于"最好的难度测试者"。对于思考如何验证难度曲线的设计者而言,这是一篇启发颇多的论文。
可读的失败 — 如何让解谜中的死局显形
在解谜游戏里,失败不是死亡,而是死局。从推箱子不可逆的推动,到 Stephen's Sausage Roll 看不见的僵局,再到 The Witness 与 COCOON 不制造死局的设计,本文探讨的不是该不该惩罚失败,而是失败能否被读出来。
Zeytuncu:谜题的难度由「使用数字的个数」决定——Fukai 导读
关于 Yunus E. Zeytuncu 对整数四则运算谜题(给定若干数字通过四则运算构成目标数的 Numbers 类谜题)进行难度建模的论文。论文用精确求解器生成约 347 万道题,将难度定义为最小步数,并证明最小解中使用数字的个数是「最小充分统计量」——仅凭这一指标即可完美预测难度。
「还想要点别的」——Capcom《Pragmata》挑战的"谜题×射击"共存设计(Game Developer)
今天只介绍一篇。专业媒体Game Developer上的设计文章(Alessandro Fillari,2026年4月14日)以英文原文通读。题材是Capcom新作三人称射击游戏《Pragmata(普拉格马塔)》——需要在战斗中实时解"蛇形"黑客谜题来削弱敌人的罕见"谜题×射击"结构。开发团队(总监Cho Yonghee等)谈设计的访谈文章。
Chao et al.: Insight Is About Searching Far — Fukai Reads
A paper on insightful problem-solving by Chao, Hsieh & Wu. Using a Japanese RAT and a simulation to quantify the search path to a solution, it shows that de-fixation is necessary for solving but is not what determines insight; the hallmark of insight is exploring the solution space over greater distances.
谜题并非为了「增加难度」,而是为了「展示系统」——Patrick Traynor 讲述 Patrick’s Parabox 的系统化设计(GDC 2024)
今天一篇。Patrick Traynor(Patrick’s Parabox 作者)在 GDC 2024 上发表的讲演《System-Centric Puzzle Design in Patrick’s Parabox》官方幻灯片。他的出发点是悖论式的——「谜题的目的不是制作酷炫的谜题。谜题的目的是展示这个酷炫的系统(递归箱子)」。因此难度被设计为「传达」而非「挑战」的工具。
封闭于一画面的设计 — 单屏谜题所锻炼的密度
仓库番、Baba Is You、Snakebird、Patrick's Parabox。思考系谜题的核心作品,都将盘面收纳在一画面之内。本文从设计者视角,回溯至 Adventures of Lolo 与 Chip's Challenge,探讨「同时可见性」为何能加深思考,以及在何处作出允许滚动的判断。
Sun 等人:为何玩家沉迷于惩罚性高难游戏?——Fukai 精读
Sun 等人关于 Soulslike 游戏难度设计的论文。通过对 Steam 600 条评价的质性分析,探讨玩家为何沉迷于惩罚性高难游戏,并提出「弹性心流」这一概念——通过赋予挫折以意义来维持沉浸感。
提示系统的设计 — 如何显示,如何隐藏
InvisiClues 的隐形墨水、The Witness 的沉默、The Case of the Golden Idol 的摩擦、Return of the Obra Dinn 的三人确认。将提示功能「如何显示、如何隐藏」的设计,作为面向卡关玩家的态度宣言加以整理。
Tsumiki 设计讨论摘要 — 2026年6月3日(世界版·修订)
以可靠来源重构的版本。今日两篇,均从截然相反的方向回答「如何为玩家提供恰到好处的难度」。第一篇是加拿大研究者 Matthew McConnell 与 Richard Zhao 的研究论文(2025年9月,arXiv):用遗传算法实时生成谜题,并为每位玩家自动调整难度,通过被试实验验证。重要发现:仅以「通关时间(time-on-task)」作为指标时,难度调整效果不佳。第二篇是游戏设计师 Michael Hicks 的访谈(Game Developer)。他指出,量产困难谜题很简单,真正的难点在于「发现有趣的想法」。让机器适配难度的方式,与人手编写意义来设计难度的方式。两篇来源均为同行评审研究与专业媒体,是值得创作者反复阅读的内容。
Tsumiki 设计议论汇总 — 2026 年 6 月 3 日
今日 2 条。两者都从完全不同的角度触到同一个问题: 「为了不让玩家被关在谜题门外,设计者能做些什么」。第 1 条是 Game Developer 的访谈(2026 年 4 月 14 日): Capcom 新作 Pragmata 作为把贪吃蛇式实时黑客谜题叠到第三人称射击之上的「谜题射击」,开发者本人如何谈论「不让玩家觉得是在反复」的设计。第 2 条是游戏设计师 Cheryl-Jean Leo 2017 年写的随笔《Are You Creating Impossible Puzzles?》。从「无论设计得多么细致,你必然会对某些玩家造出解不出来的谜题」这一前提出发,讨论在游戏中给出答案这件事是否可行。新作的现场,与 9 年前的批评。并置之后,难度设计的核心隐约浮现。
动词的稀少何以孕育丰盈 — 减法设计的谱系
仓库番、Snakebird、Stephen's Sausage Roll、A Monster's Expedition、Bonfire Peaks。本文从「为什么越少越深」的提问出发,以设计者视角整理这条不靠加动词、靠减法把难度挖深的谱系。
对 Baba Is You 的反论 — 从 Steam 差评里重读
面对 Komugi 给到 9.5/10 的 Baba Is You,本文检验 Steam 与各家评测站反复出现的 5 种差评模式——难度的墙、单解的强加、变量的泛滥、挫败感主导、难度波动。我同意哪些,反对哪些。
Undo 的伦理 — 宽恕还是惩罚
一个按钮重塑整个体验。仓库番的重启、Braid 的回溯、Baba Is You 的无限 Undo。宽恕试错的设计与惩罚试错的设计,分界在何处。
学习曲线该如何切分 — 从 Baba 学到的「垂直墙」的造法
在哪里、卡住多少玩家。Baba Is You 那道恶名昭著的垂直墙、COCOON 那条不停的流、以及夹在两者之间的设计哲学——本文做比较。





















