Skip to content

Benchmarks

Benchmarks: what total-agent-memory measures and what it does not claim

Every number on this page comes from a file in the repository that you can open, and the runner that produced it is public. Where a number has a limit, the limit is written next to it. Where another product publishes a different kind of number, we name it and do not put it in the same chart.

How to read the numbers

Retrieval recall, R@5
The right evidence came back in the top five results. It says nothing about the answer an agent then writes. This is the number we publish for LongMemEval.
Answer accuracy
A generator writes an answer and a judge model scores it. Most memory products publish this. It depends on the generator and the judge as much as on the memory, so it cannot be compared with a retrieval number.
Development run
A measurement taken while the code was being changed to improve it, on the same questions. It is not an independent evaluation and we label it as such.

Benchmarks

Numbers we can actually reproduce.

Every value below links to a result file; runners live in benchmarks/. Retrieval numbers were measured on 13.x; the 14.x releases make no new retrieval claim. Figures that competitors publish on a different metric are named, not plotted. LongMemEval is a 500-question long-term memory benchmark (Wu et al., 2024); we score its 470 non-abstention questions.

LongMemEval R@5 · 13.0 store mode
recall_any@5 by question type · retrieval, not answer accuracy
raw json →
all 470 questions 95.1% knowledge-update 100.0% multi-session 98.3% single-session-user 95.3% single-session-assistant 94.6% temporal-reasoning 92.9% single-session-preference 80.0%
Not plotted — published on a different metric:
mem0 · answer accuracySupermemory · answer accuracyZep · DMR benchmark
Search latency
warm = caches hot · cold = first query after start
warm p50 · eval 2026-04-17 0.065 ms
warm p95 · eval 2026-04-17 2.97 ms
LoCoMo query p50 · 13.x 18.2 ms
cold p50 · eval 2026-04-17 1,333 ms
Warm rows are cached steady-state queries; the LoCoMo row is full hybrid retrieval. Sources: evals/results-2026-04-17.json, benchmarks/results/v13-locomo-retrieval.json. Zero LLM calls and zero network requests on the default save/search path — enforced by tests/test_no_llm_hot_path_v11.py.
Limits we state for 14.x
  • No top-10 ranking and no new default answer-quality improvement is claimed.
  • LoCoMo 66.49% and LongMemEval 73.20% are controlled development QA measurements, not independent evaluations.
  • Against Mem0 OSS on LoCoMo with the same prompt and judge, the paired difference was +0.84 pp, 95% CI [−3.03; +4.96]: no quality advantage is established.
  • The optional BGE reranker at one thread measured p95 ≈ 2,179 ms on Linux ARM64 — above the 200 ms target.
  • An earlier comparison of our retrieval recall with a competitor's answer accuracy was invalid and has been withdrawn.

Sources: docs/RELEASE_V14.md · docs/LOCOMO_V14_RESULTS.md · release verification report

Found a wrong number? Open an issue with a link to the source and we will re-run our eval and correct this page. file an issue →

The like-for-like run against Mem0 OSS

During development of 14.0 we ran LoCoMo QA (1,540 questions, one judge model, the same answer prompt) for both systems. total-agent-memory scored 72.21% with its context mode; Mem0 OSS with the same prompt scored 71.36%. The paired difference is +0.84 points with a 95% interval from −3.03 to +4.96, so no quality advantage is established. On the same 200 questions, local retrieval medians were 28.36 ms for total-agent-memory and 71.83 ms for Mem0 OSS. Mem0 OSS ran with a local Qdrant and FastEmbed setup, which is not the hosted platform Mem0 reports its own numbers on.

The changes measured in that run were developed on the same questions, so it is a development run, not a held-out evaluation.

Sources

Comparing with a specific tool? See the comparison pages.