AI memory systems are getting a more standardized comparison. The Agent Memory Leaderboard (AML) has released its first public evaluation results, focusing on “Text Memory” using a shared benchmark framework designed to improve consistency across submissions.

AML separates memory-system performance from downstream components by using a common interface for memory (Add → Search) and then having the benchmark platform handle answering and judging (Answer → Eval) under the same evaluation process. The effort is meant to make scores more comparable and reproducible, addressing a common issue in agent evaluations where differences in datasets, prompts, models, or judges can obscure what the memory system itself contributes.

The initial release includes two tracks: Open-source Methods (Text Memory) and Commercial Products (Text Memory). According to the organizers, 136 teams registered and about two-thirds of submissions completed the first evaluation. The first live commercial-product rankings list MemoraX first (58.02), followed by MemOS (45.89) and NTES-MEMORY-SMART (44.21). Additional details and full rankings are published on the AML leaderboard, with further analyses planned for later rounds.