Skip to content
Docs menu

Tool reference · Evaluation

memory_eval_recall

Measures recall on your own scenario dataset, or on a tiny built-in check that saves two records and searches for them.

local stdio server read-only
Caution. Without dataset_path it saves two test records into the real store under project "memory_eval_recall" and does not remove them.

When to use

  • You have a dataset file of queries with expected results.
  • You want a quick smoke test that save and search work end to end.
  • You want the same result format as memory_eval_locomo for a custom dataset.

Parameters

NameTypeDefaultDescription
dataset_path string — —
top_k integer 5 —
limit integer — —
mode "fast" | "balanced" | "deep" "fast" —

Example

Arguments

{
  "top_k": 5,
  "mode": "fast"
}

Result shape

{
  "scenarios_total": 2,
  "scenarios_passed": 2,
  "recall_at_5": 1,
  "recall_at_10": 1,
  "latency_ms": 21.6,
  "mode": "fast",
  "llm_calls_during_eval": 0,
  "network_calls_during_eval": 0,
  "dataset": "builtin"
}

dataset is "builtin" or the dataset_path you passed.

Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.

Server description

The description the server sends to your agent in tools/list, captured from the v14.7.0 source:

v11.0 Phase 8: generic recall benchmark on a dataset path or a small built-in fixture. Same payload shape as memory_eval_locomo.

Found a mistake? Open an issue on GitHub.

Search