Tool reference · Save & recall
memory_recall
Searches everything saved in memory across past sessions and returns the best-matching records grouped by type.
local stdio server read-only
When to use
- Before starting a task, to check whether it was already solved or decided.
- You hit an error and want to know if a fix was recorded earlier.
- You need a cheap list of candidate IDs first (mode: 'index') and will fetch full text later.
- You want source excerpts or indexed passages to answer from (mode: 'context' or 'evidence').
- You answer from context mode and your records are short (pass fill_budget: true). The plain context mode excerpts the top limit hits, so most of context_max_chars can stay empty; fill_budget searches up to 100 hits deep and keeps whole records, in rank order, until the budget is used. Session neighbours default to 0 here; neighbors: N adds them.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
queryrequired | string | — | What to search for |
project | string | — | Filter by project name |
type | "decision" | "fact" | "solution" | "lesson" | "convention" | "all" | "all" | — |
limit | integer | 10 | — |
mode | "search" | "index" | "timeline" | "context" | "evidence" | "search" | Progressive-disclosure mode: 'search' (default) = normal results, 'index' = ultra-compact metadata only (id+title+score+type+project+created_at, ~40-60 tok/hit, no cognitive expansion, use memory_get(ids=...) to fetch full content), 'timeline' = chronological compact view; 'context' = source excerpts; 'evidence' = indexed passage search with bounded follow-up retrieval |
context_max_chars | integer | 24000 | Context character budget; evidence mode uses the same value as a stricter UTF-8 byte budget. (min 512) |
evidence_followup | boolean | true | Evidence mode: allow one search for an explicitly supplied missing_relation. |
missing_relation | { subject, relation, time } | — | Evidence mode: missing relation; subject must occur in the original question. |
neighbors | integer | 2 | Timeline/context modes: records before/after each hit; context accepts 0–3. |
fill_budgetNew in 14.7.0 | boolean | false | Context mode: search deeper (up to 100 hits) and keep whole hits in rank order until context_max_chars is used, instead of excerpting the top `limit` hits to fit. Neighbours default to 0 here; pass `neighbors` to keep each hit with the records around it. |
detail | "compact" | "summary" | "full" | "auto" | "full" | Level of detail: 'compact' ~50 tokens/result (id+title+score), 'summary' truncates content to 150 chars, 'full' returns everything, 'auto' picks based on query complexity (paths/urls/code → full, short → compact). Ignored when mode!='search'. |
branch | string | — | Filter by git branch (also includes branch-agnostic records) |
fusion | "rrf" | "legacy" | "rrf" | Score fusion method: 'rrf' = Reciprocal Rank Fusion (better multi-tier ranking), 'legacy' = original additive scoring |
rerank | boolean | false | Enable CrossEncoder re-ranking for higher precision (adds ~30ms latency) |
diverse | boolean | false | Enable MMR diversity to reduce redundant results (useful for broad queries) |
expand_context | boolean | false | Add graph-related records (1-hop neighbors via knowledge graph) as 'expansion' results |
expand_budget | integer | 5 | Max number of additional records to include via graph expansion |
topics | string[] | — | Filter results to records tagged with any of these topics (from deep enrichment) |
entities | string[] | — | Filter by extracted entity names (technology/person/project, case-insensitive) |
intent | string | — | Filter by classified intent (question|procedural|fact|decision|problem|solution|incident|plan) |
decisions_only | boolean | false | Return only structured decisions (v8.0): type=decision AND tags contain 'structured'. Results include parsed schema payload under 'decision'. |
Example
Arguments
{
"query": "billing job lock",
"project": "my-api",
"type": "all",
"limit": 5,
"detail": "summary"
} Result shape
{
"query": "billing job lock",
"total": 1,
"detail": "summary",
"fusion": "rrf",
"total_tokens": 96,
"results": {
"decision": [
{
"id": 1842,
"content": "Use PostgreSQL advisory locks for the nightly billing job instead of a Redis lock.",
"context": "",
"project": "my-api",
"tags": [
"database",
"billing"
],
"confidence": 1,
"importance": "high",
"created_at": "2026-09-20T10:14:03Z",
"session_id": "a1b2c3",
"score": 0.912,
"via": [
"fts",
"semantic"
],
"recall_count": 3,
"decay": 0.98,
"_tokens": 96
}
]
},
"tiers_used": [
"fts",
"semantic"
],
"semantic_diagnostics": []
} In search mode a cognitive block (related rules, past failures, extra solutions) may be added. Other modes return a different shape: index returns a flat results list with mode: 'index'.
Values are illustrative; the keys follow the server's handler. MCP clients receive the result as JSON text content.
Server description
The description the server sends to your agent in tools/list, captured from the v14.7.0 source:
Search ALL memory: decisions, solutions, facts, lessons from ALL past sessions. 6-stage pipeline: FTS5+BM25 → semantic → fuzzy → graph → (optional) CrossEncoder → (optional) MMR. Default: hybrid mode (BM25 + semantic + RRF). Use BEFORE starting any task. v11.0: routes to fast hot path when MEMORY_MODE=fast (default). Use memory_search_fast / memory_explain_search for explicit fast routing.