Skip to content
Docs menu

Tool reference · Maintenance & performance

benchmark

Runs the built-in test scenarios against your memory and reports recall at k, how often past mistakes are flagged, and latency.

local stdio server read-only
Caution. Scores depend on what is in your store: scenarios whose expected records do not exist will fail.

When to use

  • After changing search settings, to check recall did not get worse.
  • You wrote your own scenario files and want to score them.
  • You want latency percentiles for recall on your real data.

Parameters

NameTypeDefaultDescription
scenarios_path string — Custom scenarios dir or file

Example

Arguments

{
  "scenarios_path": "evals/scenarios/recall_core.json"
}

Result shape

{
  "total": 12,
  "passed": 11,
  "failed": 1,
  "recall": {
    "total": 12,
    "passed": 11,
    "r_at_1": 0.75,
    "r_at_5": 0.9167,
    "r_at_10": 0.9167
  },
  "prevention": {
    "total": 0,
    "passed": 0,
    "rate": 0
  },
  "latency": {
    "mean_ms": 38.2,
    "p50_ms": 31.5,
    "p95_ms": 72.9,
    "max_ms": 88.4
  },
  "scenarios": [
    {
      "name": "recall_postgres_pool",
      "type": "recall",
      "passed": true,
      "latency_ms": 29.1,
      "details": {
        "first_match_rank": 1
      }
    }
  ]
}

Without scenarios_path it loads every JSON file in evals/scenarios/.

Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.

Server description

The description the server sends to your agent in tools/list, captured from the v14.7.0 source:

Run the eval harness: recall_at_k, prevention_rate, latency percentiles. Loads scenarios from evals/scenarios/*.json by default.

Found a mistake? Open an issue on GitHub.

Search