APM-Bench tests cross-session persistent memory for egocentric streaming video assistants
benchmarks
| Source: HF Papers | Original article
APM‑Bench, a new benchmark, measures cross‑session persistent memory in egocentric streaming video assistants, addressing gaps in existing short‑clip evaluations.
A team of researchers from Shanghai Jiao‑Tong University and the Eastern Institute of Technology Ningbo has unveiled **APM‑Bench**, a new benchmark designed to test cross‑session persistent memory in egocentric streaming‑video assistants. The dataset reframes egocentric video streams as “multi‑session life trajectories,” comprising 549 individual sessions grouped into 104 longer trajectories and evaluated against 2,719 candidate items. By structuring the data this way, APM‑Bench measures whether an always‑on assistant can retrieve information it observed in earlier, interrupted sessions—a capability that current streaming benchmarks largely ignore.
The release matters because real‑world personal assistants must operate across fragmented interactions, remembering past events to provide context‑aware help. Existing benchmarks tend to focus on single, continuous clips, leaving a gap in evaluating memory that persists through pauses and device switches. Early results show that naïve raw‑video replay still delivers the highest recall, but at a prohibitive storage cost of roughly 3 GiB per hour, and no current system simultaneously optimises utility, latency and storage. By quantifying this trade‑off, APM‑Bench gives researchers a concrete target for building more efficient memory mechanisms.
The benchmark follows a series of recent efforts to assess multimodal memory, such as the VoxMem and Beyond Dyadic Memory studies we covered earlier this month. Its introduction is likely to spur a wave of model‑level innovations aimed at balancing recall with practical constraints, and could influence the design of consumer‑facing AI features—like Instagram’s newly announced conversational video‑editing assistant.
Going forward, the community will watch for follow‑up papers that propose memory architectures capable of meeting APM‑Bench’s demands, as well as any adoption of the benchmark by major AI labs. Success in this arena could bring truly persistent, context‑rich video assistants from research labs into everyday devices.
Sources
Back to AIPULSEN