Tool reference · Evaluation
memory_eval_locomo
Runs the LongMemEval-style scenario suite against the live store in a chosen mode and reports recall@5/10, latency and any LLM or network calls made.
local stdio server read-only
Caution. Resets the in-process performance counters. balanced and deep modes may call an LLM and take much longer.
When to use
- You want one summary number set for recall quality in fast mode.
- You want to prove the fast mode makes zero LLM or network calls.
- You want to compare fast, balanced and deep modes on the same scenarios.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
top_k | integer | 5 | — |
limit | integer | — | Cap how many scenarios to run. |
mode | "fast" | "balanced" | "deep" | "fast" | — |
scenarios_path | string | — | Optional override path. |
Example
Arguments
{
"top_k": 5,
"limit": 20,
"mode": "fast"
} Result shape
{
"scenarios_total": 20,
"scenarios_passed": 18,
"recall_at_5": 0.9,
"recall_at_10": 0.95,
"latency_ms": 74.3,
"mode": "fast",
"llm_calls_during_eval": 0,
"network_calls_during_eval": 0,
"details": {
"recall": {
"total": 20,
"passed": 18,
"r_at_1": 0.7,
"r_at_5": 0.9,
"r_at_10": 0.95
},
"prevention": {
"total": 0,
"passed": 0,
"rate": 0
},
"latency": {
"mean_ms": 40.1,
"p50_ms": 35,
"p95_ms": 74.3,
"max_ms": 90.2
}
}
} latency_ms is the 95th percentile. On a crash the result is { status: "error", reason, mode }.
Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.
Server description
The description the server sends to your agent in tools/list, captured from the v14.7.0 source:
v11.0 Phase 8: run the LongMemEval-style recall+prevention scenario suite (loaded from evals/scenarios/) against the live store. Forces MEMORY_MODE=fast by default. Returns {scenarios_total, scenarios_passed, recall_at_5, recall_at_10, latency_ms, mode, llm_calls_during_eval, network_calls_during_eval}.