Skip to content
Docs menu

Tool reference · Evaluation

memory_eval_long_context

Saves many filler records plus one unique "needle" record and checks whether search still finds the needle.

local stdio server read-only
Caution. Writes n_records + 1 real records into your store under project "memory_eval_long_context" on every run and does not remove them. Large values are slow.

When to use

  • You want to know if recall holds up as the store grows.
  • After changing ranking or fusion settings.
  • You want latency for a search over a crowded project.

Parameters

NameTypeDefaultDescription
n_records integer 200 —
top_k integer 5 —
mode "fast" | "balanced" | "deep" "fast" —

Example

Arguments

{
  "n_records": 200,
  "top_k": 5,
  "mode": "fast"
}

Result shape

{
  "scenarios_total": 1,
  "scenarios_passed": 1,
  "recall_at_5": 1,
  "recall_at_10": 1,
  "latency_ms": 46.8,
  "mode": "fast",
  "llm_calls_during_eval": 0,
  "network_calls_during_eval": 0,
  "n_records": 200
}

n_records has a minimum of 10.

Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.

Server description

The description the server sends to your agent in tools/list, captured from the v14.7.0 source:

v11.0 Phase 8: large-context recall scenario. Saves N records and queries them at the tail. Reuses eval_harness scenarios tagged 'long_context' if present.

Found a mistake? Open an issue on GitHub.

Search