Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM News Intel Wants a Piece of the Custom AI Chip Boom, and Its New Secret Weapon Is Marvell's Ex-Sales Chief News Autonomous Trucks Are Now Hauling Freight to America's Busiest Land Border. A Human Is Still in the Cab. Business GlobalFoundries Will Make TSMC's AI Chip Interposers in New York Under a $2 Billion, Five-Year Deal Productivity 90% of UK Lawyers Now Use AI, and the Billable Hour Is Starting to Crack Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM News Intel Wants a Piece of the Custom AI Chip Boom, and Its New Secret Weapon Is Marvell's Ex-Sales Chief News Autonomous Trucks Are Now Hauling Freight to America's Busiest Land Border. A Human Is Still in the Cab. Business GlobalFoundries Will Make TSMC's AI Chip Interposers in New York Under a $2 Billion, Five-Year Deal Productivity 90% of UK Lawyers Now Use AI, and the Billable Hour Is Starting to Crack

Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM

Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM

Researchers affiliated with University College London and Huawei Noah's Ark Lab report that Memento 3, a frozen-LLM agent that maintains a natural-language rulebook compiled into executable world-model code, cleared every level of all 25 public ARC-AGI-3 games with a mean relative human action efficiency of 100.0, using 7,518 actions against a 17,135-action human baseline. A Claude Opus 5 run on the same benchmark scored 40.7.

A preprint posted to arXiv on October 8 argues that the fastest way to a better agent may be to leave the model exactly where it is. Memento 3 (arXiv:2610.11794), from authors including Haoyu Zhao and Jun Wang with affiliations the paper and coverage tie to University College London and Huawei's Noah's Ark Lab, describes a system in which a frozen LLM learns to master unfamiliar game environments by continuously rewriting a plain-language rulebook and compiling it into executable world-model code. The underlying language model never receives a single gradient update.

The mechanism is a loop of observation, reflection, rule revision, compilation and verification. The agent records its hypotheses about how the environment works in a Markdown rulebook kept as persistent semantic memory, deliberately leaving unknown dynamics underspecified. When its predictions fail, it revises the rules; revised code is accepted only after two gates pass — the LLM itself must judge the code faithful to the rulebook, and a cell-exact replay must reproduce the observed state transitions. The authors frame the process as a model-based route to recursive self-improvement, in which exploration, model revision and planning all improve while the base model stays fixed.

The headline result is on ARC-AGI-3, the game-based benchmark from the ARC Prize Foundation designed to test learning in unfamiliar environments. The single-model agent cleared every level of all 25 public games and recorded a mean Relative Human Action Efficiency of 100.0, the metric's ceiling, while consuming 7,518 actions in total — 44% of the 17,135-action human baseline. For comparison, the paper reports a Claude Opus 5 run on the same benchmark harness at 40.7 mean RHAE, 59.3 points lower.

An Atari Pong case study extends the pattern beyond puzzle grids. A learned feedback controller distilled from the agent's world model won 21:0 in each of three evaluated episodes with different openings, using 9,504 emulator frames of learning and no LLM calls at evaluation time. The paper claims this is roughly 42 times more sample-efficient than model-based RL baselines such as EfficientZero V2, which it lists at around 400,000 frames; that efficiency figure is the authors' own comparison and has not been independently replicated.

Ablations in the paper suggest the rulebook itself carries real weight: removing its contribution costs 9% more actions and 18% more agent turns. A population variant that maintains multiple world models in parallel, sharing interaction evidence, cut the action count on one game from 899 to 597 — a 33% reduction — while holding RHAE at 100.

The result lands in the middle of a live argument about where agent capability comes from. Labs are betting enormous compute budgets on RL and weight updates; Memento 3 makes the case that an external, inspectable, continuously revised memory can beat a much larger inference-only baseline on a benchmark explicitly built to expose the gap between machine and human learning. Because the rulebook is plain text, the agent's accumulated knowledge is auditable in a way that tuned weights are not.

The caveats are the usual ones for a benchmark paper. The 25 games are the public set; ARC-AGI-3 also maintains private and semi-private holdings, so the result does not measure generalization to unseen games. All figures are the authors' self-reported evaluations, and independent replications have not yet appeared. Still, the direction of travel is notable: the paper joins a fast-growing cluster of work treating external memory and explicit world models as a credible rival to brute-force training, and it does so with the kind of complete, verifiable sweep — every level of every public game — that is hard to dismiss as cherry-picking.

Comments (0)

Log in to join the discussion

Log In

No comments yet