Tool reference · Evaluation
memory_eval_long_context
Saves many filler records plus one unique "needle" record and checks whether search still finds the needle.
local stdio server read-only
Caution. Writes n_records + 1 real records into your store under project "memory_eval_long_context" on every run and does not remove them. Large values are slow.
When to use
- You want to know if recall holds up as the store grows.
- After changing ranking or fusion settings.
- You want latency for a search over a crowded project.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
n_records | integer | 200 | — |
top_k | integer | 5 | — |
mode | "fast" | "balanced" | "deep" | "fast" | — |
Example
Arguments
{
"n_records": 200,
"top_k": 5,
"mode": "fast"
} Result shape
{
"scenarios_total": 1,
"scenarios_passed": 1,
"recall_at_5": 1,
"recall_at_10": 1,
"latency_ms": 46.8,
"mode": "fast",
"llm_calls_during_eval": 0,
"network_calls_during_eval": 0,
"n_records": 200
} n_records has a minimum of 10.
Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.
Server description
The description the server sends to your agent in tools/list, captured from the v14.7.0 source:
v11.0 Phase 8: large-context recall scenario. Saves N records and queries them at the tail. Reuses eval_harness scenarios tagged 'long_context' if present.