REMORY Develops Residual Memory Method to Compact Context
agents
| Source: HF Papers | Original article
Researchers unveil REMORY, a neural memory network that augments textual summaries with a bounded set of soft memory tokens to help long‑horizon agents retain context within limited windows.
Long‑horizon AI agents have long struggled with the trade‑off between retaining enough context to make informed decisions and staying within the finite token windows of large language models. A new paper and accompanying code release introduce **REMORY**, a neural memory network that augments a compact textual summary with a bounded sequence of “soft” memory tokens. The tokens are generated from the full history and the summary, then appended after the summary, forming a residual‑style connection that lets a frozen LLM recover information that would otherwise be lost.
Early experiments show the approach delivers measurable gains on several state‑of‑the‑art models. On Qwen‑3.8‑27B and GLM‑5.3‑Flash, REMORY improves performance on long‑horizon benchmarks while cutting down repeated tool outputs and tool‑related errors. On the SummHay evaluation suite, the method boosts source‑attribution scores and approaches full‑context joint performance while using only about 5 % of the original input positions.
The development matters because it offers a scalable way to keep agents’ reasoning grounded in earlier events without inflating context size. By preserving nuanced details in soft tokens, agents can maintain higher fidelity in tool use, reduce hallucinations, and make more consistent decisions over extended interactions—key hurdles for autonomous assistants, research bots, and AI‑driven workflows.
The next steps will likely focus on broader validation across diverse tasks and integration into existing agent frameworks such as those explored in recent reinforcement‑learning and tool‑augmented projects. Watch for follow‑up studies that test REMORY with larger models, real‑world deployments, and potential refinements to the token generation process that could further shrink the memory footprint while preserving decision quality.
Sources
Back to AIPULSEN