Skip to content
Docs menu

Tool reference · Evaluation

memory_eval_locomo

Runs the LongMemEval-style scenario suite against the live store in a chosen mode and reports recall@5/10, latency and any LLM or network calls made.

local stdio server read-only
Caution. Resets the in-process performance counters. balanced and deep modes may call an LLM and take much longer.

When to use

  • You want one summary number set for recall quality in fast mode.
  • You want to prove the fast mode makes zero LLM or network calls.
  • You want to compare fast, balanced and deep modes on the same scenarios.

Parameters

NameTypeDefaultDescription
top_k integer 5 —
limit integer — Cap how many scenarios to run.
mode "fast" | "balanced" | "deep" "fast" —
scenarios_path string — Optional override path.

Example

Arguments

{
  "top_k": 5,
  "limit": 20,
  "mode": "fast"
}

Result shape

{
  "scenarios_total": 20,
  "scenarios_passed": 18,
  "recall_at_5": 0.9,
  "recall_at_10": 0.95,
  "latency_ms": 74.3,
  "mode": "fast",
  "llm_calls_during_eval": 0,
  "network_calls_during_eval": 0,
  "details": {
    "recall": {
      "total": 20,
      "passed": 18,
      "r_at_1": 0.7,
      "r_at_5": 0.9,
      "r_at_10": 0.95
    },
    "prevention": {
      "total": 0,
      "passed": 0,
      "rate": 0
    },
    "latency": {
      "mean_ms": 40.1,
      "p50_ms": 35,
      "p95_ms": 74.3,
      "max_ms": 90.2
    }
  }
}

latency_ms is the 95th percentile. On a crash the result is { status: "error", reason, mode }.

Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.

Server description

The description the server sends to your agent in tools/list, captured from the v14.7.0 source:

v11.0 Phase 8: run the LongMemEval-style recall+prevention scenario suite (loaded from evals/scenarios/) against the live store. Forces MEMORY_MODE=fast by default. Returns {scenarios_total, scenarios_passed, recall_at_5, recall_at_10, latency_ms, mode, llm_calls_during_eval, network_calls_during_eval}.

Found a mistake? Open an issue on GitHub.

Search