Tool reference · Evaluation
memory_eval_recall
Measures recall on your own scenario dataset, or on a tiny built-in check that saves two records and searches for them.
local stdio server read-only
Caution. Without dataset_path it saves two test records into the real store under project "memory_eval_recall" and does not remove them.
When to use
- You have a dataset file of queries with expected results.
- You want a quick smoke test that save and search work end to end.
- You want the same result format as memory_eval_locomo for a custom dataset.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
dataset_path | string | — | — |
top_k | integer | 5 | — |
limit | integer | — | — |
mode | "fast" | "balanced" | "deep" | "fast" | — |
Example
Arguments
{
"top_k": 5,
"mode": "fast"
} Result shape
{
"scenarios_total": 2,
"scenarios_passed": 2,
"recall_at_5": 1,
"recall_at_10": 1,
"latency_ms": 21.6,
"mode": "fast",
"llm_calls_during_eval": 0,
"network_calls_during_eval": 0,
"dataset": "builtin"
} dataset is "builtin" or the dataset_path you passed.
Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.
Server description
The description the server sends to your agent in tools/list, captured from the v14.7.0 source:
v11.0 Phase 8: generic recall benchmark on a dataset path or a small built-in fixture. Same payload shape as memory_eval_locomo.