For anyone that hasn’t seen it yet, I started a YouTube channel where I’ll upload explainer-style videos of both papers discussed in my newsletter and introductory content about important AI concepts.
This week in LLM Watch:
Base models with a light harness reach more agentic tasks than their post-trained versions once given enough rollouts, and an 8-rollout estimate predicts how much coverage post-training cost (Sharpening Tax).
In ten real companion relationships, only 3.4% of sampled user messages needed anything beyond the current thread, so pooled memory benchmarks mostly measure messages that need no memory (RealCompanion).
With the backbone fixed, switching from OpenCode to Claude Code raised the best score on unfamiliar games from 74.5 to 81.8 (Learn2Play Bench).
Let’s go!
Sharpening Tax in Post-Training

What problem does it solve?
A common hypothesis holds that RL post-training sharpens behaviors the base model already has: single-shot accuracy (pass@1) rises while the number of tasks solved across many attempts (pass@K) falls. The evidence comes mostly from math and coding. Agentic tasks could behave differently, because multi-turn tool use is widely assumed to be learned during post-training. Researchers at Meta Superintelligence Labs, the University of Wisconsin–Madison and Stanford test this on 14 open base/post-trained checkpoint pairs from Gemma-4, Ministral-3, Qwen2.5 and Qwen3.5 (3B to 35B effective parameters), across BFCL v4 multi-turn, WebShop and ACEBench, which gives 42 combinations (paper, code).
How does it solve the problem?
Base models run inside a model-agnostic harness: a plain-text prompt, a tool catalog written as Python signatures, and JSON tool-call blocks. Post-trained models use the default scaffolding with thinking mode off. Each task gets up to 128 rollouts. The Sharpening Tax sums the gap between pass@K and pass@k for every k below K, divides by the single-shot failure rate, and subtracts the post-trained score from the base score. The authors prove this area equals the expected number of failed attempts before the first success, so a positive tax means post-training made successes depend less on retries.
Raising the sampling temperature is the obvious fix for lost coverage, and the paper reports that it improves pass@K but lowers pass@1. Its alternative, posterior-tempered group sampling (PTGS), keeps discounted success and failure counts for each prompt during RL training, draws a difficulty estimate from a Beta posterior, and sets that prompt’s temperature between 1/τ and τ: hard prompts are sampled hotter and easy prompts cooler. The RL update itself is unchanged.
What are the key findings?
On WebShop, the gemma-4-31B base model passes over 85% of tasks at 128 rollouts, against 56% for gemma-4-31B-it. The rollout count at which the base model overtakes falls from above 128 for the 4B Gemma to about 3 for the 31B. Post-training cut the share of WebShop tasks solved sometimes but not always from 87.6% to 30.0% for gemma-4-31B, while never-solved tasks rose from 12.4% to 44.0%. The calibrated tax at 128 rollouts is positive in 36 of the 42 combinations, and its value from 8 rollouts on half the tasks predicts the 32-rollout tax on the other half with Spearman ρ=0.85. With PTGS, Qwen2.5-7B-Instruct trained with PPO on Sokoban goes from 46.5 to 61.1 pass@1 and from 55.0 to 69.7 pass@128. With GRPO, Sokoban pass@128 goes from 55.3 to 72.5.
Why does it matter?
If your product samples many candidates and has a verifier to pick one, the post-trained checkpoint is not the default choice at large model sizes. Run 8 rollouts of both base and post-trained versions on your own tasks and compute the tax before deciding. The paper also reports that routing tasks between base and post-trained policies by tax often beats either one alone. For clients judging fine-tuned agents, ask for pass@K and the tax next to pass@1.
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

What problem does it solve?
Memory benchmarks such as LoCoMo and LongMemEval generate the conversation and write the questions, so nearly every probe needs the past and the benchmark never tests whether retrieval was warranted. Quis Lab releases ten real relationships with an AI companion: 27,218 messages over 36 to 120 days, released with consent and anonymized (paper).
How does it solve the problem?
Each participant ships with five files: the conversation and four files derived from it (profile, persona, chat ground truth, question set), with every claim citing the message identifiers it rests on. Chat probes are messages users actually sent. Labels come from a five-stage pipeline (referent, verification against the source, rule-based tier, grounded reference reply, validation) that records its reasoning at each stage. Verification can remove evidence but never add it, so labeling errors can only make memory look rarer. A random proportional sample of 1,169 scoreable probes supports rates. An enriched sweep supplies the memory-bearing cases, and 434 session openers test whether a system holds back.
What are the key findings?
On the proportional sample, 3.4% of messages need anything outside the current thread (1.3% under a strict reading). When memory is needed, the furthest required message sits a median of 2,157 messages back. Recency finds a required message in the top five for 95.9% of pooled probes and 2.2% of the distant ones, while BM25 finds one for 24.8% of distant probes. Replacing the last three messages with the recorded evidence raises content match by 0.108 on the proportional sample, and 96% of that gain lands on probes that need no memory. On the 167 items whose reference reply depends on the evidence, the gain is 0.455. A detector shown the probe and three prior messages reaches AUROC 0.530, near chance, while authored questions reach 0.802. Labeling the same ten messages as retrieved memories rather than earlier turns raises memory use by 10 to 14 points. Misfires are zero with no context and rise to 60.8% when ten unselected messages are supplied. Three agent systems reconstruct personas at F1 0.69 to 0.70 while processing between 18M tokens (Codex GPT-5.6-sol) and 546M (Claude Opus 5.5).
Limitations: Ten people on one app is not a population sample, and under the strict reading only five contribute memory-bearing probes. A stylometric attack re-identified authors at 83% (statistical) and 98% (language model), against 10% chance.
Why does it matter?
If you build or sell assistant memory, a gain measured on authored questions mostly reflects messages that never needed retrieval. Report misfire rate on probes that need nothing, split ablations by whether memory was required, and drop the retrieved-memories header from injected context. The cost spread in persona reconstruction is a direct line item for enterprise pricing.
AgentGarten: Code Worlds for Evolving Agents

What problem does it solve?
Simulators keep consistent state and rules but need costly 3D assets for realistic visuals. Video world models produce realistic frames but keep state implicit in generated history, where rendering errors accumulate and nothing can be inspected or edited. MirroS, Tsinghua University and Peking University split the two (paper, code).
How does it solve the problem?
A scene program runs in a simulator or game engine that holds object poses, articulations and task variables. The engine exports depth or surface-normal maps, and one shared neural renderer turns them, plus a reference image and text, into frames. The state transition never reads a rendered frame, so renderer mistakes cannot change what happened. Coding agents write the scene programs from an image or text.
The renderer starts from Cosmos 3-Nano and is distilled with Adversarial Forcing: distribution matching on the model’s own rollouts, a second pass that recomputes history block by block so later losses reach the parameters that encoded earlier frames, and a real-video GAN loss with an exact R1/R2 penalty computed without differentiating through fused attention kernels. Agents act through short Python programs, see only rendered frames, and after each round write Markdown playbooks that the next round’s agents read.
What are the key findings?
Block-by-block replay matches the rollout bitwise, against 3.99% relative L2 error for a full-sequence FlexAttention replay, with a 5.4% faster forward pass. The renderer runs at 36.5 frames per second at 480x832 on one H100, 438.8 ms per 16-frame block. Against Self Forcing, the comparison over 30-second rollouts is qualitative only. In hide-and-seek, hiders built shelters by round 4 and seekers crossed walls with ramps by round 10. The authors cite roughly 25 million and 100 million episodes for the same milestones in OpenAI’s 2019 self-play study and state that the two paradigms differ, so this is not a sample-efficiency ratio. In four more worlds over four rounds, a companion-dog score rose from 13 to 19, a bridge swap fell from 71 to 41 seconds, and a quarry loader went from 30 to 90.9 after scoring 0 in round 2.
Why does it matter?
If you build training or evaluation environments for computer-use or embodied agents, keep state in code you can inspect and replay, and let the learned model handle only appearance. A recorded rollout can then be re-rendered in a different style without changing events. Treat the playbook results as a demonstration rather than a measured learning rate.
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
What problem does it solve?
Group-relative RL with verifiable rewards (RLVR), such as GRPO, learns from differences in reward within a group of rollouts. When every rollout in a group gets the same reward, the advantage is zero and the group contributes no gradient, even though the trajectories show what the task required and how the agent failed. The work comes from a team including Chien-Sheng Wu, Julian McAuley and Silvio Savarese (paper). Only the abstract was available to us, so this section stays at that level of detail.
How does it solve the problem?
The authors call the approach prospective learning: experience gathered after an attempt supervises predictions the agent makes before acting. Self-Retrospection Distillation (SRD) uses each completed trajectory to identify knowledge that would have helped and pitfalls to avoid. That hindsight is privileged, since it depends on having seen the trajectory. SRD trains the same policy, seeing only the pre-interaction view, to predict it. The foresight is a training target only and need not be generated at inference, so deployment cost stays the same. SRD is added on top of RLVR, so it supplies a signal in groups where the reward contrast is zero.
What are the key findings?
Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD adds up to 24.2 percentage points over RLVR and self-distillation baselines. Between 37% and 98% of rollout groups were reward-uniform depending on model scale. In the 2B setting, where 98% of groups were all-failure, RLVR training ended at 0.0% success while RLVR plus SRD reached 60.6% under the same rollout budget.
Why does it matter?
If you run GRPO-style training on hard agentic tasks with small models, log the fraction of reward-uniform groups first. If most groups are all-fail, outcome reward alone may be teaching nothing, as the 2B result shows, and an auxiliary signal built from the trajectories you already collected is worth testing before buying more rollouts. Read the full paper for construction cost before committing a training budget.
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

What problem does it solve?
Benchmarks for self-improving agents usually use familiar tasks, so a score gain may come from recalled pretraining knowledge rather than learning from interaction. Researchers at the National University of Singapore built 20 new text games whose hidden rules are novel or counterintuitive. In one, stealing a colleague’s lunch gets the player promoted while hard work does not (paper).
How does it solve the problem?
Ten games repeat the same instance across episodes, and ten re-shuffle characters, menus or layouts while keeping the rules, to test transfer. Feedback is deterministic and scoring automatic. Agents play 5 episodes per game (10 for five challenge games) over three independent trials. Four metrics separate peak performance (Max, Mean) from improvement (Learning Gain, Learning Slope). The authors compare nine backbones in the OpenCode harness, six experience-retention methods (Naive, Memory, Reflexion, EvoTest, ReasoningBank, AWM) at fixed backbones, two harness swaps, and 21 to 30 college students per game.
What are the key findings?
In OpenCode, Claude Opus 5 leads Max (74.5) and Mean (53.8), while Qwen 3.8 Max leads Learning Gain (+25.2) and Learning Slope (+6.30), so the best scorer is not the fastest learner. With Kimi K3, Memory, which keeps raw interaction histories, leads all four metrics (Max 63.3, against 62.4 for EvoTest and 54.7 for Naive). With Opus 5 the order differs: EvoTest has the higher Max, 66.1 against 60.9. In one GemForge run, EvoTest turned a single failed fusion into a prompt rule banning combinations it never tried. Claude Code with Opus 5 reaches Max 81.8 against 74.5 in OpenCode, and Codex with GPT-5.6-SOL reaches 74.3 against 57.0. The top human reaches 84.3. Humans repeat action sequences less (similarity 0.54 against 0.66 for agents) and set a new personal best after a setback more often (33% against 22%). OpenCode with Kimi K3 uses 56% fewer input tokens per episode than Memory and scores higher Max.
Why does it matter?
Before building a reflection or rule-distillation layer, test plain retained history on your backbone. Distilled rules can lock in conclusions drawn from one observation. Evaluate harness and model together: on this benchmark a harness swap moved Max by 7 to 17 points at a fixed model, and in some cases also lowered cost.
MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

What problem does it solve?
Vision-language-action models transfer poorly to new tasks without more robot data and cannot directly benefit from newer general VLMs. Agentic alternatives add segmentation models, coding agents, skill libraries or motion planners. Researchers at the University of Illinois Urbana-Champaign test whether one general VLM can drive a robot the way a human teleoperator does, with none of those extra components (paper).
How does it solve the problem?
A 240-question diagnostic built from LIBERO demonstrations shows action selection is the weakest VLM skill: Qwen3.8-Flash-Next scores 36.25% on it, against 55.00% on progress assessment and 65.00% on completion. That motivates short, revisable action proposals with repeated checks. One frozen VLM plays five roles. A Planner writes subgoals with success criteria. An Executor proposes short batches of move, rotate and gripper commands in the robot base frame with magnitudes it picks, such as moving forward 12.20 mm. A deterministic Controller validates and executes them and reports what the robot actually did. A Monitor on a background thread can cancel queued commands at the next action boundary, a Verifier checks outcomes against observations and measured state, and Memory writes summaries without blocking execution.
What are the key findings?
On LIBERO-PRO, MotorMind reaches 66.7% success on base suites and 53.8% under perturbations, against at most 13.3% and 19.2% for the evaluated zero-shot methods. Under perturbations it edges task-fine-tuned OpenVLA-OFT at 51.2%, though fine-tuned π0.5 reaches 98.3% on the base suites. On a real xArm6 it succeeds in 95% of placement trials, 98% without interference and 92% with a person moving objects. Swapping in GPT-6 Sol raises base success to 83.3% with longer wall time. Removing replanning drops success to 36.7%, and removing the planner drops it to 0.0%. Most remaining failures are visual grounding errors and premature completion claims.
Limitations: Episodes average 223.4 seconds on base suites, against about 6 seconds for fine-tuned π0.5. Ablations and the backbone swap are single-seed, and all tasks are tabletop.
Why does it matter?
For a client that needs a manipulation pilot without collecting demonstrations, a VLM with mid-level actions, a replanning loop and an interrupting monitor is now a measured option, and it improves when the backbone improves. The cycle time rules it out where throughput counts. Budget for that, and for failures where the model picks the wrong object.
❤️ If you enjoyed this article, give it a like and share it with your peers.
Extended “Papers of the Week” for Additional Reading
nanoMuse: An Open-Source Personal Agent for Every Device You Own Zhejiang University releases a GPL-3.0 personal agent that runs on Android, iOS and desktop, operates phone and computer screens, routes every action through a Sentinel gate, and stores memory as readable files. It reports no success rate for its screen-operating hands.
Long-WAM: Scaling the Context of World-Action Models NVIDIA, MIT and collaborators find longer visual history helps far more after autoregressive video pretraining: RoboCasa GR-1 success rises from 63.3% to 78.7% at 19.2 seconds of context, while a bidirectional initialization shows no net gain.
World Observer: Joint Actor-Observer Generation for Persistent World Modeling KAIST generates panoramic observer videos jointly with the agent’s view, so objects keep moving while unseen. On its synthetic benchmark, out-of-view dynamics consistency reaches 0.531, against 0.387 without the observer stream.
Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness HKU MMLab loops designer, builder, code-writing player and reviewer agents. Three refinement rounds raise GameCraft-Bench from 72.70 to 77.89, and GameASG-Bench task success reaches 25/47 against 9/47 for the same model alone.
From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation Trace2Env builds a worldbook from past interaction traces and lets an agent simulate the environment. On ALFWorld, 85% of actions generated in simulation still succeed in the real environment, against 3% with direct prompting.
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization Waterloo and collaborators mine recurring, success-bearing action spans into a tree with no LLM calls, then train on it. Offline RL on WebArena reaches 22.9% success with Qwen2.5-7B, against 18.4% for full-data SFT.


