Welcome to this week’s issue of LLM Watch!
For anyone that hasn’t seen it yet, I started a YouTube channel where I’ll upload explainer-style videos of both papers discussed in my newsletter and introductory content about important AI concepts.
The first video is about Just-in-Time Memory, a new agent memory technique from Salesforce AI. I’ll embed it below its text section. Cheers!
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

What problem does it solve?
Most agent memory systems decide what to keep when a task ends. The finished trajectory is distilled into a reflection, workflow, skill, or reasoning strategy, and later retrieved by similarity. The paper names two costs of this design. Details dropped at write time cannot be recovered when a later task depends on them, and one fixed summary has to serve every future query, although the same household trajectory might teach one task how to heat an object and another where to place it. Training a write-time curator is also awkward, because a storage decision is graded only when a matching query arrives, possibly many tasks later. SkillOS, the closest learned prior work, groups related tasks to create that delayed signal.
How does it solve the problem?
JitMem keeps complete raw trajectories and admits one to the bank only if the executor model, acting as judge, says the task succeeded. For a new task, BM25 over task descriptions fetches the top three trajectories. A curator (Qwen3-8B, thinking mode off) reads the current task and those traces and writes a short briefing: relevant past experiences, strategies that worked, and guidance for this task. A frozen executor acts with that briefing in its prompt, and the briefing is then thrown away. Write-time memory resembles taking notes before you know the exam questions. JitMem rereads the textbook after seeing the question.
Because the briefing is used on the same task it was written for, the curator’s reward is the task outcome itself. Training runs GRPO for 100 steps with groups of 8 candidate briefings, rewarded by each benchmark’s native success metric, against a fixed bank of base-executor successes.
What are the key findings?
With a Qwen3-8B executor, trained JitMem reaches 77.4 success rate on ALFWorld against 61.2 for SkillOS, and 32.8 on WebShop against 16.5 (paper). The untrained curator is already strong: with Gemini-2.5-Pro as curator and executor, it scores 61.0 on WebShop against 41.0 for SkillOS-gemini. A curator trained with the Qwen3-8B executor transfers to GPT-5.4 without retraining, reaching 86.7 on ALFWorld against 88.1 for a curator trained directly with GPT-5.4. The briefings are compact: on ALFWorld with GPT-5.4, the untrained curator adds 1.9K input tokens over no memory, against 10.7K for ReasoningBank and 13.4K for SkillOS-base. Removing the retrieved trajectories from the trained curator costs up to 14.8 points on ALFWorld and 15.2 on WebShop, so it relies on retrieved experience rather than its own parametric knowledge.
Why does it matter?
If you run agents on repeated task streams, store raw successful trajectories instead of only distilled lessons, and add one curation call that sees the current task. The prompted version needs no training and already beats write-time baselines in several settings. Budget for that extra LLM call per task, and expect the authors’ BM25 retriever to need replacing as the bank grows large and varied.
Prefer video? I got you covered. Let me know what you think, I’m new to YouTube and video creation in general.
Raven: The Harness of Harnesses for Composable Agentic Intelligence

What problem does it solve?
EverMind AI argues that agents now face long workflows that span domains, while each harness (tools, context handling, skills, recovery rules) is tuned for one domain and is costly to design by hand. Chaining specialists through free-form messages loses detail and makes outputs hard to trace back to their sources (paper).
How does it solve the problem?
Raven treats each model plus harness pair as a callable unit. A registry lists native specialists (Raven-Research, Raven-Code, Raven-Design, Raven-Oncall) and third-party agents such as Claude Code, Codex, Hermes Agent, and OpenClaw, connected through Agent Client Protocol, a command-line process, or an OpenAI-compatible API. A Host Agent decomposes a request into a typed DAG where each node is an agent call and each edge is a file-backed artifact handoff. Before any worker starts, the runtime runs five groups of checks: format, graph structure, agent capability, agent status, and environment, such as rejecting a path reference sent to an agent without local file access. A separate model call judges each finished node, and failures return to the host to continue, abandon, or replan. Around this sit harness evolution built on the team’s HarnessBank, EverOS memory, and Skill Forge retrieval built on SkillCorpus. A formal section gives sufficient conditions under which composition covers tasks no single agent solves reliably within the same budget.
What are the key findings?
On MAOB, the team’s new 140-request benchmark, Raven’s Exact Match graph agreement is 0.711 against 0.607 for the strongest baseline with Qwen3.8-27B, and 0.867 against 0.762 with DeepSeek-V4-Flash-0731 (paper). Raven-Research reaches 76.5% on DeepResearch Mixed with DeepSeek-V4-Flash, against 68.9% for DeepSeek-Harness and 67.2% for MiroFlow, at 0.0242 USD per question. Raven-Code resolves 15 more of 731 SWE-bench Pro tasks than Claude Code on the same Qwen3.8-27B backbone.
Why does it matter?
For teams wiring several coding or research agents together, the admission checks and artifact ledger are directly reusable ideas: they reject a bad plan before it spends compute and leave an audit trail of which file each node consumed. Enterprises evaluating orchestration products should treat planning-quality scores as a first filter and still ask for end-to-end success and cost on their own workflows.
In-Context Learning for Robots: Methods and Applications

What problem does it solve?
A robot may already know how to grasp every part and still need evidence about what to do: a demonstration sets the assembly order, history shows which placements are done, contact reveals a tighter fit. This survey (arXiv) studies how such evidence changes deployed behavior without another task-specific update to neural parameters, and asks how that change survives physical execution when objects, environments, or conditions shift.
How does it solve the problem?
The survey organizes the field by the interface that connects contextual evidence to execution. It distinguishes four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution, with correspondence and memory mechanisms shared across them. It places these within six learning horizons, from explicit control programming (S1) and learned policies (S2), through history-based adaptation (S3) and in-context task learning (S4), to physical recursive self-improvement (S5) and knowledge exchange across robot bodies (S6). A shared notation separates supplied evidence, execution intermediates, and retained state, so fine-tuning (task data updates the weights) and in-context learning (a support demonstration enters as context, weights fixed) can be compared on one policy interface.
What are the key findings?
The findings are conceptual. Where task information is stored, in weights or in context, determines acquisition cost, persistence, and what resets between episodes. Broader motor competence from multi-task, multi-embodiment pretraining is a different property from a broader ability to learn through teaching. The authors argue evaluation should separate three things: whether behavior changes appropriately when the task-defining evidence changes under matched execution conditions, whether it transfers across objects and environments, and whether retained experience helps later tasks.
Why does it matter?
If you evaluate a robot policy marketed as learning from demonstration, ask for the matched control the survey describes: swap the demonstration, hold the scene fixed, and check whether behavior changes. A high success rate on scenes that already determine the task says little about teachability. The taxonomy is also a quick map for deciding whether a new task belongs in weights or in context.
Omni-IO Skills: Harnessing Your Agent Omni-Native

What problem does it solve?
General-purpose agents plan and write code well, but producing audio, video, documents, 3D assets, and coordinated media packages depends on outside tools and explicit orchestration. Adding modalities to a foundation model ties every new capability to a model update. Bolting on specialist tools leaves open how procedures, dependencies, intermediate files, and later revisions are coordinated. The authors, from the National University of Singapore and the University of Oxford, propose a harness layer instead (paper).
How does it solve the problem?
The harness has four layers. Skill Entry holds 27 declarative Skills: 19 Atomic (one operation, such as speech generation), 2 Expert (one deliverable, such as a poster), and 6 Scenario (an application context, such as event material). Higher-level Skills expand recursively into terminal tasks. An MCP Tool Service maps each task to understanding, generation, or utility tools behind one call contract. A Provider and Configuration layer binds each capability to a provider, model, credentials, defaults, and fallbacks, so a backend can be swapped without rewriting Skills. An Asset Registry records every output in an append-only JSON file with a stable ID, type, path, parameters, turn, and source asset, with writes serialized by a file lock.
Workflows become Declare Execution Graphs. After checks that every dependency resolves and the graph is acyclic, nodes run in waves by dependency rather than modality, so image and audio jobs run together while an image-conditioned video waits. A failed node cancels its pending descendants while independent branches continue. On a later turn, earlier assets enter the graph as completed source nodes, so a landing-page restyle reruns only that node.
What are the key findings?
On UniM-90, a 90-instance subset of UniM, the harness raises input support from 40.00% to 100% for GPT-5.6 Sol and from 38.89% to 100% for Claude Sonnet 5. Relative Semantic–Quality Coupled Score rises from 26.99 to 74.94 and from 27.82 to 77.78. Strict Structure Score reaches 100.00 and 99.78 (paper). Absolute scores also rise (67.49 to 74.94, and 71.53 to 77.78), though the base agents’ absolute scores cover only the instances they could process.
Why does it matter?
For builders who need multi-asset outputs from an existing coding agent, the useful pieces are the provider binding layer and the ID-based asset registry: they let you swap image or video vendors and support revise-this-one-thing requests without rewriting procedures. Before adopting it, measure the extra calls and wall-clock time on your own requests.
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

What problem does it solve?
Self-evolving search agents build their own curricula: a proposer turns documents into questions and pseudo-labels, a solver trains on them, and the solver’s agreement with the labels rewards the proposer. Researchers from Rutgers, McGill, King Fahd University of Petroleum and Minerals, and other institutions show this loop drifts. A wrong label can train the solver to repeat the error on later questions from the same source, and that agreement is then paid back to the proposer as reward. They call this co-cheating and measure it as false-agreement mass, the fraction of label–response pairs that agree on the same wrong answer (paper). In the standard Dr. Zero loop with Qwen3.5, it grows from about 0.4% in round 1 to 6.1% (4B) and 8.8% (9B) by round 3, while in-loop agreement keeps rising.
How does it solve the problem?
Verifying each proposal before training is the obvious fix, and the authors build it as multi-sample verification (MSV): three answers with the source and three without, admitting a task only when both majorities agree. It costs six extra generations per candidate and only lowers false agreement to 5.7% and 7.2%, because six samples from the same model can share errors and the solver that later scores the proposer still trained on labels from the same source.
CrossFit targets that second path. Each source document is assigned once to fold A or B. Two auxiliary solvers each train on one fold, and questions from fold A are scored by the solver trained only on B, and vice versa, using Dr. Zero’s unchanged frontier reward. The main solver still trains on all admitted questions.
What are the key findings?
CrossFit lowers round-3 false agreement to 3.0% at 4B and 3.7% at 9B, against 6.1% and 8.8% for the coupled loop. Replaying the same 3,000 saved questions with source-excluded feedback drops it to 0.4% and 0.1%, which separates feedback ancestry from curriculum changes. A random question-level split reaches only 0.050 and 0.062, so the split has to follow source documents. Across seven search benchmarks, CrossFit averages 48.8% Cover-EM at 4B against 40.0% for Dr. Zero and 40.1% for Search-R1, and 51.2% at 9B against 42.8% and 43.4% (paper).
Why does it matter?
If you run self-play or self-generated data loops, log which data trained the model that scores your generator, and keep a held-out evidence check outside the loop. Rising in-loop agreement alone is not evidence of progress. Splitting feedback by source document is cheap to add to an existing proposer–solver pipeline.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

What problem does it solve?
Symbolic music models make melody, harmony, and form readable but stop before a recording. Audio models produce full songs but leave composition implicit, so there is nothing to inspect or edit. YuE2, from HKUST, Multimodal Art Projection, New York University, Stanford, MBZUAI, ACE Studio, and others, asks whether one model can write the score first and still match the finished-song quality of proprietary systems (paper, weights).
How does it solve the problem?
A 3.58B-parameter, 28-layer Mixture-of-Transformers generates three layers in order. It first writes an ABC score with tempo, meter, key, chords, sections, and vocal and instrumental melody. It then predicts 25 Hz MERT2 semantic tokens and finally 64-dimensional acoustic latents decoded to 48 kHz stereo. Each layer has separate autoregressive and non-autoregressive experts that share one attention computation: score and semantic tokens are predicted causally, while acoustic latents are refined bidirectionally by flow matching while attending to the full plan. Training mixes four tasks, with and without the score and semantic tokens, so one checkpoint supports a matched with-versus-without planning comparison. Since recordings lack aligned scores, the team built SheetSage2 to transcribe lead sheets and MERT2 to supply the semantic tokens.
What are the key findings?
With the same checkpoint, prompts, and decoder, experts prefer songs generated with symbolic planning for overall quality, 49.3% against 34.6% (p=0.0070). With planning in both systems, they prefer the unified model to a separate language model plus diffusion Transformer, 53.4% against 35.6%. On WildSongBench, YuE2 scores 6.7316 SongBench Global Avg against 6.3247 for LeVo 2, the best public baseline, and best-of-8 reaches 6.9632 against 6.9377 for Mureka 9. In expert listening, best-of-8 is preferred over Suno v4.5 57.3% to 30.5% and is nearly even with Suno v5 at 40.4% to 39.9% (paper). Score edits follow through: melody edits reach 84.17% pitch accuracy while 90% to 94% of unedited content is preserved across local edits.
Suno v6 still wins overall preference, 59.3% against 31.6%. Lyric accuracy (PER 0.0844) trails MiniMax Music 3 (0.0627). The symbolic training targets are SheetSage2 transcriptions, not human-aligned scores, so transcription errors enter supervision.
Why does it matter?
For teams building music or audio products, the score gives users and language-model agents a concrete object to edit, and those edits change targeted melody and harmony while keeping the rest. The open weights make this a practical base for controllable generation. Best-of-8 results cost eight generations per request.
(More) Papers of the Week
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks BAAI fine-tunes Qwen3.8-27B on multi-round improvement trajectories from ML engineering and algorithmic programming. It reaches 81.8 on MLE-bench Lite against 73.7 for Naive-N0.5-Flash, and 84.0 on BrowseComp without new research data.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL GAGAR has a trained grader rank test-passing trajectories within a GRPO group, then shifts credit toward cleaner implementations while preserving the advantage sum. On MiMo-V2.6-Flash it reports better code agent performance and slower trajectory-length growth.
LEGO-Anything: Coding Agents for 3D Scene Reconstruction Coding agents rebuild a scene from one image as an editable Blender program. On LEGO-Bench, GPT-6-astra scores 53.4% indoors and 39.6% outdoors, and a training-free plugin improves all six models by up to 62.7% relative.
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Sampling and verifying candidate terminal commands before execution, with the generator and harness unchanged, raises TMAX-9B Pass@1 on TerminalBench-Lite from 50.00% to 68.03% using a GPT-5.6 Sol verifier over 8 candidates.
Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence HexaAnything represents robot task state and procedure as code and calls a VLA as a tool, raising RoboCasa365 Composite-Unseen success from 34.3% to 38.3% over native XR-1.
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Co-evolving search queries with solutions raises OpenEvolve’s normalized discovery gain across 21 tasks from 61.3% to 82.3% with Gemini-3.8-Flash and from 74.1% to 78.0% with GPT-5.6-Luna. Qwen3.5-9B does not benefit.
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents UC Santa Barbara and MIT release 213 anomaly-finding tasks across 13 3D environments. VLM agents succeed on 28.2% to 42.3% of tasks, VLA-then-VLM pipelines on 6.6% to 17.4%, and humans on 83.4%.

