Most agent memory systems curate at write time. When a task finishes, the system reads the trajectory and distills it into a reflection, a workflow, a skill or a reasoning strategy, then stores that artifact for similarity retrieval later. The authors of Just-in-Time Memory name two costs of this design. Anything the distiller discards is gone for every future task that needed it. A single fixed summary also has to serve every future query, although one trajectory can teach several lessons. Their example is a household interaction that might teach one later task how to heat or cool an object and teach another an object-placement strategy. Which lesson matters depends on a task that does not exist yet at write time.
There is a training problem as well. If you want to learn the distiller, the reward for a storage decision arrives only when some later query retrieves the artifact, possibly many tasks later. The closest prior work, SkillOS, trains a skill curator with GRPO and has to group related tasks to create that delayed signal. According to the JitMem paper, SkillOS’s own ablations identify the grouping as a major contributor to its performance.
Below, we’ll take a look at both similar and opposite approaches - and when to us what.
JitMem stores raw traces and writes a briefing for each new task
JitMem has four parts. The memory bank stores complete raw trajectories: the task description plus every observation and action, with no summarization. A trajectory enters the bank only if the executor model, acting as an LLM judge, decides the task succeeded. The retriever is plain BM25 over task descriptions only, not over trajectory content, and returns the top 3 traces. The curator is Qwen3-8B with thinking disabled. It reads the new task together with the retrieved raw traces and writes a short natural-language payload: which past experiences are relevant, which strategies worked, and specific guidance for this task. The executor is a frozen model that receives the payload prepended to its prompt and never sees the raw traces. The payload is discarded after use. Only the new trajectory is considered for storage.
Because the payload is written for the task that consumes it, the same stored trace produces different payloads for different tasks. The paper shows one trace producing heat-before-placement guidance for “put a hot potato in fridge” and location-verification guidance for “put a newspaper in sofa” (JitMem).

The design also changes the credit-assignment problem. The curator’s output is used immediately on the same task, so the task’s own success score is its reward. Training is GRPO with group size 8, batch size 32, learning rate 1e-6 and 100 steps, with no value network and no task grouping. The executor stays frozen. The training bank is fixed, built once from the base executor’s successful runs on the training set using ground-truth labels, so that reward reflects payload quality and does not depend on which traces happen to be available.
Most of the gain appears before any training
The headline results compare RL-trained curators on the same Qwen3-8B base with a Qwen3-8B executor. JitMem reaches 77.4 success rate on ALFWorld against 61.2 for SkillOS, and 32.8 on WebShop against 16.5 (JitMem). On τ²-bench, which has no standard training split, only untrained variants were evaluated. There, JitMem with a prompted GPT-5.4 curator reaches 73.4 macro-average success against 70.4 for ReasoningBank with the same GPT-5.4 curator.
The comparison that isolates the idea uses no training at all. With Gemini-2.5-Pro as both curator and executor on WebShop, untrained JitMem reaches 61.0 success rate against 41.0 for SkillOS under the same configuration. With GPT-5.4 as executor on ALFWorld, untrained JitMem with a Qwen3-8B curator scores 79.3 and beats ReasoningBank (77.9) and SkillOS (70.0), both of which used GPT-5.4 as curator. In that comparison, a smaller curator working at read time beat larger curators working at write time.
The ablations test each design choice separately. Removing the current task from the curator’s input, which turns it into a generic summarizer, costs the trained version up to 11.4 points on ALFWorld and 10.4 on WebShop. Distilling traces ReasoningBank-style before storage costs the untrained version 6.8 to 8.2 points on WebShop. Storing failed trajectories with correctness labels, the practice ReasoningBank and SkillOS use, costs the untrained version 2.3 to 3.4 points on WebShop compared with storing only successes. Forcing the trained curator to work with no retrieved traces drops it to the untrained level or below, by up to 15.2 points on WebShop, which the authors read as evidence that RL taught it to distill the traces rather than to produce hints from its own weights.
Two results bear directly on cost. The curator trained with a Qwen3-8B executor transfers to GPT-5.4 at 86.7 ALFWorld success, within 1.4 points of a curator trained directly against GPT-5.4 (88.1). On ALFWorld with a GPT-5.4 executor, untrained JitMem adds 1.9K input tokens over no memory, against 10.7K for ReasoningBank and 13.4K for untrained SkillOS, and cuts executor steps from 17.8 to 13.2.
The paper also reports where the method stops helping. On τ²-bench, the gains come from Telecom (72.6 against 61.6 for ReasoningBank with GPT-5.4). On Airline and Retail, no memory method beats the no-memory agent beyond variance. The authors read this as read-time curation helping most when tasks require synthesizing procedural guidance rather than simple fact retrieval.
JAM agrees on read time and spends far more compute per query
Just-In-Time Agent Memory, or JAM, from the Beijing Academy of Artificial Intelligence with Peking University and Hong Kong Polytechnic University, makes the same diagnosis in different terms. It calls write-time systems Ahead-of-Time (AOT) and says their request-agnostic compression can discard fine-grained details and cross-session dependencies that later become important. Its fix matches JitMem in outline: keep the raw history and build context at query time.

The two designs differ in how much work happens at read time. JitMem makes one curator call per task. JAM runs an agent. Its Memorizer stores raw sessions as files in a directory hierarchy, with a short memo per session and a README per directory built from those memos. Its Researcher, a Qwen3.5-4B model, loops over three tools (open a directory, search with hybrid BM25 plus BGE-M3 retrieval, browse a file), decides when it has enough evidence, and writes a context with source provenance. Training is supervised fine-tuning on filtered teacher trajectories followed by GRPO, rewarded on recall of the gold source sessions. JitMem rewards final task success instead.
On LoCoMo, JAM reaches 52.09 F1 against 46.13 for MemAgent-14B and 32.10 for Mem0 running on the same Qwen3.5-4B backbone (JAM). It answers a LoCoMo query in 13.81 seconds against 58.08 for MemAgent, with a median of 3 research rounds. Its query latency is higher than that of the one-shot AOT systems it beats, and the authors state this trade openly. Its one-time workspace build takes 79.53 seconds on LoCoMo, against 6.82 for LightMem.
JAM also qualifies the pure raw-store position. Its write-time layer holds navigational summaries rather than lessons, and that layer matters. Replacing the hierarchy with a flat raw store lowers LoCoMo F1 from 52.09 to 49.06 and raises latency from 13.81 to 17.11 seconds per query. Removing either BM25 or dense retrieval also hurts. JitMem lists its BM25-over-task-descriptions retriever as a likely bottleneck at scale, but neither paper tests whether JAM’s hybrid retrieval would help JitMem. JitMem shows raw storage working with no write-time summary at all. JAM shows that, on LoCoMo, a navigational index over raw storage helps.
Adobe’s skill bank shows write-time memory can work when every edit is replayed
The Designer-RSI paper from Adobe and Brown University is the counterexample. Its agent drives equivalents of Photoshop, Illustrator and InDesign through more than 230 tools, and its memory consists of natural-language SKILL.md playbooks, which are write-time artifacts. The frozen model never changes. The bank widens by minting a skill once an uncovered subtask recurs 3 times, and deepens by rewriting skills that fail repeatedly, contrasting the failed runs with the same skill’s successful ones. Every change goes through a replay gate. The gate freezes the upstream state, including the retrieved assets, replays the candidate in the same batch against the incumbent skill (for a rewrite) or the no-skill agent (for a mint), judges each pair with randomized order, and admits the change only if no prompt loses and at least one wins. Across five rounds over 1,406 briefs, the gate rejected 100 of 231 rewrites and 67 of 136 mints.

The results favor write-time skills. With Claude-Sonnet-4, GenEval2 execution success rises from 72.7% to 99.3%. Against the no-skill agent across four design benchmarks, the evolved bank wins 61.8% of pairwise comparisons on Claude-Sonnet-4 and 67.6% on Claude-Opus-4.6. The paper also reports mean latency overhead of 3.4% to 6.2% (Designer-RSI).
One ablation is consistent with JitMem’s warning. The 76 skills distilled from product documentation, written before any task was seen, did not beat having no skills: 68.62 completeness against 69.08, with a 46.4% win rate. Minting alone reached a 49.4% win rate and rewriting alone 48.6%. Only the combination reached 58.5% (p = 0.025) on 200 held-out briefs. Round 4 regressed at low completeness thresholds, and the authors attribute this to newly minted skills that had not yet been revised and misfired on requests outside the cluster they were minted from. Their summary is that minting expands coverage while rewriting converts that coverage into reliability. JitMem instead writes a new payload for each task, at the cost of one extra curator call per task.
This difference matters: Designer-RSI stores procedures, such as the steps of a double-exposure effect, that recur across many briefs. JitMem and JAM store episodes and facts whose relevance depends on the query. Procedures that recur may be worth distilling once and gating carefully. The gate still covers only the replayed cases, as the authors note, and gains concentrate in completeness while aesthetics barely move (65.92 to 66.53).
Curated context still has to be weighted correctly by the model that reads it
All three systems end by placing memory into a model’s context and assuming the model uses it appropriately. MemCalib, from the University of Science and Technology of China, Alibaba’s Qwen Business Unit and Fudan University, tests that assumption. It splits memory blocks into atomic propositions and labels each one Ignore (leave no answer-specific trace), Bound (provide local support) or Control (decide a conclusion). A DeepSeek-V4-Pro judge then scores how each atom was actually used. The judge agreed with a human annotator on 96.7% of 150 sampled judgments.
No evaluated model does this well. The best, GPT-5.6-SOL, reaches a Sample Calibration Score of 46.25 and gets every atom right on 28.40% of test examples. Larger models are not necessarily better: Qwen3-8B scores 31.17 against 26.54 for Qwen3.5-35B-A3B. Most models over-use memory, and Qwen3-8B under-uses it, which the authors read as models relying on a broad prior about whether memory should be trusted rather than judging each proposition. Relative to the shared cold-start checkpoint, every other post-training method tested, including GRPO, reduced one error direction while increasing the other on at least one model. Their MemCalib-RL, which assigns credit to response tokens by removing the atoms in each reward channel and measuring the change in token likelihood, raises Qwen3-8B to 79.54 SCS against 72.25 for the strongest baseline, GDPO, and is the only method that reduces both error directions on all three models tested.
Two findings in the other papers resemble these errors, though MemCalib measured neither. JitMem’s finding that failure-labeled traces hurt, because “the curator cannot fully suppress the noise,” resembles over-use, but it concerns the curator rather than the answering model. Designer-RSI’s first stated limitation is that a retrieved skill may not make the model follow it when the model has a strong default strategy, which resembles under-use. A JitMem-style curator might reduce misuse by giving the executor less irrelevant material to weigh, but none of these papers tests that.
What to build from this
For an agent that repeats similar tasks, the JitMem setup is inexpensive to try. Log full raw trajectories. Admit only runs that a judge marks as successful. Retrieve three with BM25 on the task description, and add one curator call that receives the current task plus those traces and writes a short briefing. The untrained curator carried most of the gain in the paper. GRPO training on task success adds more if you have a training split, and a curator trained with a Qwen3-8B executor transferred to GPT-5.4 with a 1.4-point loss on ALFWorld.
If your queries run over long, cross-session histories, JAM is the design tested on that setting, and its Researcher took about 14 seconds per LoCoMo query. The papers do not compare JAM and JitMem directly. If your agent repeats recognizable procedures across many clients, distilled skills with a replay gate are defensible, but expect documentation-derived skills to help little until they are revised against real failures.
For enterprise clients, JitMem and JAM both keep raw interaction records rather than summaries. That shifts the governance question from what the system chose to remember to how long it retains what it saw. This follows from the designs, and neither paper addresses it.
Nobody has tested how the executor uses the curated payload
The open question sits between JitMem and MemCalib. JitMem measures whether a curated payload raises task success. It does not measure whether the frozen executor gives each line of that payload the right weight, and on τ²-bench Airline and Retail no memory method beats the no-memory agent beyond variance. MemCalib measures that weighting on benchmark-built memory blocks, not on payloads written by a curator. An experiment that scores curator-written payloads with MemCalib’s Ignore, Bound and Control rubric would show whether read-time curation reduces the calibration problem or leaves it with the executor. The work we discussed today supports curating at read time. It doesn’t show whether the executor then uses what it receives correctly.



