TAG
#llm-agents
0 reviews · 3 essays
Related essays
Tudor et al.: The scoreboard was one query away, and the agent never opened it — Fukai Reads
An arXiv preprint (submitted 2 September 2026) by seven authors from Oxford and elsewhere. They wired 76 tool endpoints into Sid Meier's Civilization VI and had language-model agents play whole games of 300+ turns. Agents queried victory progress only once every 30-75 turns (the supplied playbook recommended every 20), and in 7 of 20 losses that were foreseeable they never checked it in the final 20 turns. Commitments the agents wrote down for themselves were carried out within ten turns only 48.2%-65.8% of the time.
Liu et al.: More Memory Makes AI Agents Less Cooperative — Fukai Reads
An arXiv paper from a Carnegie Mellon-led team studying how an LLM agent's memory length affects cooperation. Across 7 models, 4 repeated social-dilemma games, history windows up to 80 rounds and 500-round matches, longer history degrades cooperation in 18 of 28 settings — a 'memory curse.' The cause is the content of accumulated defection records, not context length, and forward-looking reasoning partly fixes it.
Feng et al.: Can LLM Agents Bargain Well in a Trading Game? — Fukai Reads
A Tsinghua University team's benchmark, SidConArena, for evaluating LLM agents in a cooperative-yet-competitive trading game. Built on the board game Sidereal Confluence, it scores agents across negotiation, production, and sealed-bid auction phases, finding that frontier models are stronger but still misprice resources, bargain passively, and plan poorly over long horizons.
