MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks

2 hours ago 3



Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have introduced JAZ, a stripped-down agent framework. In benchmark tests it beat Letta, formerly known as MemGPT, on recall-heavy tasks. It also topped ACE, a specialized self-improvement harness, while spending less money to do it. Instead of giving a language model a filing cabinet, give it the ability to write code that reaches into its own past. One primitive to rule them all The paper is titled “Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity.” It was posted to arXiv under the identifier arXiv:2609.26891. At its core, JAZ relies on a single LLM-based primitive called invoke. That’s the whole toolkit, more or less. JAZ takes a different route. The agent’s history and its prompt are exposed as variables inside a code environment. The model can then write executable code to inspect, slice, or manipulate that history directly. The benchmark numbers The researchers tested JAZ on two fronts: long-term recall and self-improvement. For recall, they used the StuLife benchmark and ran JAZ on the GPT-5.4 nano model. JAZ scored 70%. Letta scored 62%. That eight-point gap is n...

Read Entire Article