Skip to content
Docs menu

Tool reference · Save & recall

memory_recall

Searches everything saved in memory across past sessions and returns the best-matching records grouped by type.

local stdio server read-only

When to use

  • Before starting a task, to check whether it was already solved or decided.
  • You hit an error and want to know if a fix was recorded earlier.
  • You need a cheap list of candidate IDs first (mode: 'index') and will fetch full text later.
  • You want source excerpts or indexed passages to answer from (mode: 'context' or 'evidence').
  • You answer from context mode and your records are short (pass fill_budget: true). The plain context mode excerpts the top limit hits, so most of context_max_chars can stay empty; fill_budget searches up to 100 hits deep and keeps whole records, in rank order, until the budget is used. Session neighbours default to 0 here; neighbors: N adds them.

Parameters

NameTypeDefaultDescription
queryrequired string — What to search for
project string — Filter by project name
type "decision" | "fact" | "solution" | "lesson" | "convention" | "all" "all" —
limit integer 10 —
mode "search" | "index" | "timeline" | "context" | "evidence" "search" Progressive-disclosure mode: 'search' (default) = normal results, 'index' = ultra-compact metadata only (id+title+score+type+project+created_at, ~40-60 tok/hit, no cognitive expansion, use memory_get(ids=...) to fetch full content), 'timeline' = chronological compact view; 'context' = source excerpts; 'evidence' = indexed passage search with bounded follow-up retrieval
context_max_chars integer 24000 Context character budget; evidence mode uses the same value as a stricter UTF-8 byte budget. (min 512)
evidence_followup boolean true Evidence mode: allow one search for an explicitly supplied missing_relation.
missing_relation { subject, relation, time } — Evidence mode: missing relation; subject must occur in the original question.
neighbors integer 2 Timeline/context modes: records before/after each hit; context accepts 0–3.
fill_budgetNew in 14.7.0 boolean false Context mode: search deeper (up to 100 hits) and keep whole hits in rank order until context_max_chars is used, instead of excerpting the top `limit` hits to fit. Neighbours default to 0 here; pass `neighbors` to keep each hit with the records around it.
detail "compact" | "summary" | "full" | "auto" "full" Level of detail: 'compact' ~50 tokens/result (id+title+score), 'summary' truncates content to 150 chars, 'full' returns everything, 'auto' picks based on query complexity (paths/urls/code → full, short → compact). Ignored when mode!='search'.
branch string — Filter by git branch (also includes branch-agnostic records)
fusion "rrf" | "legacy" "rrf" Score fusion method: 'rrf' = Reciprocal Rank Fusion (better multi-tier ranking), 'legacy' = original additive scoring
rerank boolean false Enable CrossEncoder re-ranking for higher precision (adds ~30ms latency)
diverse boolean false Enable MMR diversity to reduce redundant results (useful for broad queries)
expand_context boolean false Add graph-related records (1-hop neighbors via knowledge graph) as 'expansion' results
expand_budget integer 5 Max number of additional records to include via graph expansion
topics string[] — Filter results to records tagged with any of these topics (from deep enrichment)
entities string[] — Filter by extracted entity names (technology/person/project, case-insensitive)
intent string — Filter by classified intent (question|procedural|fact|decision|problem|solution|incident|plan)
decisions_only boolean false Return only structured decisions (v8.0): type=decision AND tags contain 'structured'. Results include parsed schema payload under 'decision'.

Example

Arguments

{
  "query": "billing job lock",
  "project": "my-api",
  "type": "all",
  "limit": 5,
  "detail": "summary"
}

Result shape

{
  "query": "billing job lock",
  "total": 1,
  "detail": "summary",
  "fusion": "rrf",
  "total_tokens": 96,
  "results": {
    "decision": [
      {
        "id": 1842,
        "content": "Use PostgreSQL advisory locks for the nightly billing job instead of a Redis lock.",
        "context": "",
        "project": "my-api",
        "tags": [
          "database",
          "billing"
        ],
        "confidence": 1,
        "importance": "high",
        "created_at": "2026-09-20T10:14:03Z",
        "session_id": "a1b2c3",
        "score": 0.912,
        "via": [
          "fts",
          "semantic"
        ],
        "recall_count": 3,
        "decay": 0.98,
        "_tokens": 96
      }
    ]
  },
  "tiers_used": [
    "fts",
    "semantic"
  ],
  "semantic_diagnostics": []
}

In search mode a cognitive block (related rules, past failures, extra solutions) may be added. Other modes return a different shape: index returns a flat results list with mode: 'index'.

Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.

Server description

The description the server sends to your agent in tools/list, captured from the v14.7.0 source:

Search ALL memory: decisions, solutions, facts, lessons from ALL past sessions. 6-stage pipeline: FTS5+BM25 → semantic → fuzzy → graph → (optional) CrossEncoder → (optional) MMR. Default: hybrid mode (BM25 + semantic + RRF). Use BEFORE starting any task. v11.0: routes to fast hot path when MEMORY_MODE=fast (default). Use memory_search_fast / memory_explain_search for explicit fast routing.

Found a mistake? Open an issue on GitHub.

Search