JitMem: Just-in-Time Memory for agents

Almost all agent memory systems decide what to remember when a task ends. JitMem keeps complete runs and prepares a summary only when needed, for the task at hand. How it works, how much it improves results, how it compares with JAM, Designer-RSI and MemCalib and what it takes to try it.

AIR&DAIAI agentsMemoryLLMReinforcement learningRetrievalData governance
Contents
  1. The limit of summaries written in advance
  2. How JitMem works
  3. What it measures
  4. Three related papers
  5. What we think
  6. Sources
Four figures on JitMem and agent memory curated at read time
Figures from the papers cited. Sources at the end.

An agent that repeats similar tasks needs memory. Today the most common solution works like this: when a task ends, the system rereads what the agent did and summarises it into a text to keep, for example a lesson learned, a procedure or a skill. In later tasks it retrieves the most similar summaries. The decision about what to remember is taken when memory is written.

Just-in-Time Memory, or JitMem, posted on arXiv on 23 September, moves that decision to when memory is read. It is one of four papers from the last two weeks on the same problem, compared side by side by the LLM Watch newsletter.

The limit of summaries written in advance

The authors point to two problems. The first is that whatever the summary drops is lost for good, including for future tasks that would have needed it. The second is that a fixed summary has to serve every purpose, while the same experience can teach different things. In the paper’s example, the same sequence of actions in a simulated house can help one task remember to heat an object before putting it in the fridge, and another remember to check where an object is before moving it. Which lesson matters depends on a task that does not exist yet when memory is written.

There is also a training problem. If you want to teach a model to write good summaries, the reward arrives only when a future task uses them, possibly much later. SkillOS, the closest prior work, has to group similar tasks to create that signal.

How JitMem works

The system has four parts.

  1. Memory keeps complete runs, with no summaries: the task description and every step the agent took. A run enters only if the model judges it successful.
  2. Search uses BM25, a classic keyword search algorithm, over task descriptions only, and returns the three closest runs.
  3. The curator, a Qwen3-8B, reads the new task and the three runs and writes a short text: which experiences matter, what worked, what to do in this case.
  4. The executor, a model that is not modified, receives that text at the top of its prompt and carries out the task. It never sees the original runs, and the text is discarded after use.

For training the advantage is direct: the curator’s text is used straight away, on the same task, so the reward is the success of that task. The curator is trained with GRPO, a reinforcement learning method, in 100 steps and without grouping tasks. The executor stays unchanged.

What it measures

With the same base model, Qwen3-8B, trained JitMem solves 77.4% of tasks on ALFWorld against 61.2% for SkillOS and 32.8% on WebShop against 16.5%. ALFWorld and WebShop are standard test environments: a simulated house and an online shop.

The clearest result comes without training. With GPT-5.4 as executor on ALFWorld, JitMem with an untrained Qwen3-8B curator reaches 79.3%. It beats ReasoningBank (77.9%) and SkillOS (70.0%), which use GPT-5.4, a much larger model, to write memory. A small curator working at read time does better than large curators working at write time.

The authors also removed one choice at a time to see how much it weighs:

  • if the curator does not see the current task and writes a generic summary, up to 11.4 points are lost on ALFWorld;
  • if runs are summarised before being stored, the untrained curator loses 6.8 to 8.2 points on WebShop;
  • also storing failed runs, labelled as failures, gives worse results than storing only successes;
  • if the trained curator works with no runs to read, it falls back to the untrained level or below, by up to 15.2 points. For the authors this shows that training taught it to use memory, not to give generic advice.

Two results concern costs. A curator trained together with Qwen3-8B also works with GPT-5.4: 86.7% on ALFWorld, 1.4 points from one trained directly with GPT-5.4. And with GPT-5.4 on ALFWorld JitMem adds only 1,900 tokens to the prompt, against 10,700 for ReasoningBank and 13,400 for SkillOS. The agent completes the task in 13.2 steps instead of 17.8.

The method does not always help. On τ²-bench, which simulates customer service, it improves results in the telecommunications domain (72.6% against 61.6% for ReasoningBank, both with GPT-5.4). In the airline and retail domains no memory method does better than the agent with no memory. According to the authors, memory curated at read time helps most when guidance on how to proceed has to be worked out, less when retrieving a fact is enough.

JAM (Just-In-Time Agent Memory, Beijing Academy of Artificial Intelligence with Peking University and Hong Kong Polytechnic University) makes the same diagnosis. JAM also keeps the complete history, but at query time it runs a research agent, a Qwen3.5-4B. The archive is organised in folders, with a short memo for each session and a README for each folder. The agent opens folders, searches and reads files until it has enough evidence, then writes a context that cites its sources. On LoCoMo, a benchmark of long conversations, it reaches 52.09 F1 against 46.13 for MemAgent-14B and 32.10 for Mem0, in 13.81 seconds per query against 58.08 for MemAgent. One result corrects the most radical position: with a flat archive, without folders and memos, F1 drops to 49.06 and time rises to 17.11 seconds. It pays to keep raw data together with an index that helps find the way around it.

Designer-RSI (arXiv 2609.22086, Adobe and Brown University) goes the opposite way, with good results. It is a graphic design agent that uses equivalents of Photoshop, Illustrator and InDesign through more than 230 tools. Its memory is made of skills written in advance, meaning procedures in SKILL.md files. Every new or rewritten skill has to pass a test: it is replayed on the same cases and compared with the version in use, or with the agent without skills. It gets in only if it makes no case worse and at least one better. Over five rounds and 1,406 requests the test rejected 100 of 231 rewrites and 67 of 136 new skills. With Claude-Sonnet-4, success on GenEval2 rises from 72.7% to 99.3%. The 76 skills derived only from product documentation, written before seeing real tasks, did not improve anything (68.62 against 69.08 completeness). The difference with JitMem lies in what is remembered: Designer-RSI keeps procedures that recur, JitMem and JAM keep experiences and facts whose usefulness depends on the query.

MemCalib (arXiv 2609.24259, University of Science and Technology of China, Alibaba’s Qwen and Fudan University) checks an assumption the other three take for granted: that the model uses the memory it receives well. It splits memory into single statements and marks each as to be ignored, to be used as support or to be used to decide the answer. Then a judge model checks how each was actually used; the judge agrees with a human annotator 96.7% of the time. The best model tested, GPT-5.6 Sol, stops at 46.25 Sample Calibration Score and uses every statement correctly in only 28.40% of cases. Most models trust memory too much and Qwen3-8B too little. A larger model is not necessarily better. The proposed training method, MemCalib-RL, brings Qwen3-8B to 79.54.

What we think

Trying it is cheap. For an agent that repeats similar tasks four steps are enough: store complete runs, keep only successful ones, retrieve three with a keyword search on the task description and add a call to a model that reads task and runs and writes a short briefing. In the paper much of the gain already comes without training. Training adds more if you have tasks with a verifiable outcome.

It depends on what you want to remember. For procedures that recur the same way across many clients, skills written once and admitted only if they make nothing worse are a sensible choice, as long as they are rewritten on real failures and not only on documentation. For long histories across many sessions JAM is the reference, though it takes a few seconds per query. For experiences whose usefulness depends on the task, in the comparisons JitMem measured, memory curated at read time gives the best results.

Keeping everything raises a governance question. JitMem and JAM keep complete interactions rather than summaries. The question moves from what the system decided to remember to how long it keeps what it saw. In a company those runs contain customer data, documents and sometimes credentials passed through tools. Where personal data is present, the GDPR rules on storage limitation and minimisation apply: a defined retention period and a filter on what enters memory are needed. The two papers do not address this, which we covered in the second part of our agentic loop series.

One question remains open. JitMem measures whether the curator’s text makes more tasks succeed, not whether the executor gives each line the right weight. MemCalib measures exactly that weight, but on memories prepared for the test, not on texts written by a curator. Scoring JitMem’s briefings with MemCalib’s method would show whether memory curated at read time also helps the model use it well.

Sources

Need support?Under attack?Service Status
Need support?Under attack?Service Status