Benchmarks
Benchmarks: what total-agent-memory measures and what it does not claim
Every number on this page comes from a file in the repository that you can open, and the runner that produced it is public. Where a number has a limit, the limit is written next to it. Where another product publishes a different kind of number, we name it and do not put it in the same chart.
How to read the numbers
- Retrieval recall, R@5
- The right evidence came back in the top five results. It says nothing about the answer an agent then writes. This is the number we publish for LongMemEval.
- Answer accuracy
- A generator writes an answer and a judge model scores it. Most memory products publish this. It depends on the generator and the judge as much as on the memory, so it cannot be compared with a retrieval number.
- Development run
- A measurement taken while the code was being changed to improve it, on the same questions. It is not an independent evaluation and we label it as such.
Benchmarks
Numbers we can actually reproduce.
Every value below links to a result file; runners live in benchmarks/. Retrieval numbers were measured on 13.x; the 14.x releases make no new retrieval claim. Figures that competitors publish on a different metric are named, not plotted. LongMemEval is a 500-question long-term memory benchmark (Wu et al., 2024); we score its 470 non-abstention questions.
evals/results-2026-04-17.json,
benchmarks/results/v13-locomo-retrieval.json.
Zero LLM calls and zero network requests on the default save/search path — enforced by
tests/test_no_llm_hot_path_v11.py.
- No top-10 ranking and no new default answer-quality improvement is claimed.
- LoCoMo 66.49% and LongMemEval 73.20% are controlled development QA measurements, not independent evaluations.
- Against Mem0 OSS on LoCoMo with the same prompt and judge, the paired difference was +0.84 pp, 95% CI [−3.03; +4.96]: no quality advantage is established.
- The optional BGE reranker at one thread measured p95 ≈ 2,179 ms on Linux ARM64 — above the 200 ms target.
- An earlier comparison of our retrieval recall with a competitor's answer accuracy was invalid and has been withdrawn.
Sources: docs/RELEASE_V14.md · docs/LOCOMO_V14_RESULTS.md · release verification report
Found a wrong number? Open an issue with a link to the source and we will re-run our eval and correct this page. file an issue →
The like-for-like run against Mem0 OSS
During development of 14.0 we ran LoCoMo QA (1,540 questions, one judge model, the same answer prompt) for both systems. total-agent-memory scored 72.21% with its context mode; Mem0 OSS with the same prompt scored 71.36%. The paired difference is +0.84 points with a 95% interval from −3.03 to +4.96, so no quality advantage is established. On the same 200 questions, local retrieval medians were 28.36 ms for total-agent-memory and 71.83 ms for Mem0 OSS. Mem0 OSS ran with a local Qdrant and FastEmbed setup, which is not the hosted platform Mem0 reports its own numbers on.
The changes measured in that run were developed on the same questions, so it is a development run, not a held-out evaluation.
Sources
- LongMemEval per-question results, 13.0 store mode
- Third-party recomputation of the LongMemEval result
- LoCoMo 14.0 run, including the Mem0 OSS comparison
- 14.0 release validation report
- All benchmark runners and raw outputs
Comparing with a specific tool? See the comparison pages.