TAG
#language-model
0 篇评论 · 2 篇随笔
相关随笔
Battleday et al.: Measuring AI Discovery With 70 Games That Never Explain Their Rules — Fukai Reads
An arXiv preprint by Ruairidh M. Battleday and fifteen co-authors. On DiG-bench — 70 text-string games with both rules and win conditions hidden, across seven tiers, 21 released publicly — the strongest single model beat 50 games and all models pooled beat 57, while all 70 were beaten by at least one human on a first attempt. Handed the ground-truth rules, the same model jumps from 18 games to 69, and agentic harnesses did not improve on the basic one.
Mannem et al.: Some Puzzles Are Learnable, Some Are Not — Fukai Reads
An arXiv preprint by Gowrav Mannem and colleagues (Algoverse AI Research). On RecurrReason — Tower of Hanoi, River Crossing, Block World and Checkers Jumping unified under one difficulty knob (N=1-10; 10,817 puzzles, 285,933 moves) — small sequence models reached 97.27% validation and 81.00% out-of-distribution on Block World, but only 11.11%/0.00% on Tower of Hanoi, 1.11%/0.10% on Checkers Jumping, and 0.00% everywhere on River Crossing. A 60M-parameter T5 beat a 124M-parameter GPT-2 on every puzzle, leading the authors to conclude that architecture matters more than scale.
