Show HN: Agent Memory Leaderboard – first public results for AI memory systems (agentmemoryleaderboard.ai)

3 points by IreneAI ↗ HN
We just released the first public results from the Agent Memory Leaderboard (AML).

The first evaluation focuses on Text Memory across two tracks:

- Open-source Methods - Commercial Products

136 teams registered, and 69 representative memory frameworks completed the first evaluation.

Commercial Products — Text Memory:

1. MemoraX — 58.02 2. MemOS — 45.89 3. NTES-MEMORY-SMART — 44.21

The benchmark uses a common evaluation framework and a clearer system boundary:

Memory system: Add → Search Benchmark: Answer → Eval

The goal is to make memory-system results more comparable and reproducible.

3 comments

[ 15.8 ms ] story [ 487 ms ] thread
Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of the table.