Documentation
Configuration
Every setting is an environment variable with a safe default. You do not need to set anything for a normal local install.
How to set a variable
Variables are read when the server starts, so restart your IDE (or reconnect the MCP server) after a change. Where you put them depends on how you run TAM:
- IDE-managed server (the usual case): add them to the
envblock of the MCP entry in your client's config, next tocommand. - Docker Compose: put them in a
.envfile next to the compose file. The repository ships.env.examplewith the common ones. - Shell:
export NAME=valuebefore startingtamortam-team serve.
{
"mcpServers": {
"total-agent-memory": {
"command": "/Users/you/.tam/.venv/bin/total-agent-memory",
"env": { "MEMORY_MODE": "fast", "MEMORY_LLM_ENABLED": "false" }
}
}
} Remote LLM providers receive the text used in those tasks. Keep Ollama local or set
MEMORY_LLM_ENABLED=falseif that content must stay on your machine. See Privacy & local-first.
Mode
One switch picks a profile; every flag it sets can still be overridden individually.
| Variable | Default | What it does |
|---|---|---|
MEMORY_MODE | fast | ultrafast (full-text only, no embedding on save unless cached), fast (embeddings + full-text + vectors, no LLM on save or search), balanced (fast plus the background enrichment worker), deep (older synchronous behaviour: quality gate, contradiction check and reranker in the request path; slower). |
MEMORY_USE_LLM_IN_HOT_PATH | false | Allow LLM calls while saving and searching. deep sets it to true. |
MEMORY_ALLOW_OLLAMA_IN_HOT_PATH | false | Let search fall back to Ollama embeddings when FastEmbed is unavailable. balanced and deep set it to true. |
MEMORY_ENRICHMENT_ENABLED | false | Run the background enrichment worker (summaries, keywords, entities). On in balanced and deep. |
MEMORY_ASYNC_ENRICHMENT | true (set by mode) | Do enrichment in the background instead of during the save call. deep sets it to false. |
MEMORY_QUALITY_GATE_ENABLED | false (auto in deep) | Score a record before saving and reject low-quality ones. |
MEMORY_CONTRADICTION_DETECT_ENABLED | false (auto in deep) | Check new records against existing ones for contradictions at save time. |
MEMORY_ENTITY_DEDUP_ENABLED | false (auto in deep) | Merge duplicate entity names in the knowledge graph at save time. |
MEMORY_COREF_ENABLED | false | Rewrite pronouns ("after this it broke") into self-contained text using recent session history. Opt-in in every mode. |
MEMORY_QUERY_REWRITE | 0 | Rewrite search queries with an LLM before searching. Opt-in; costs an API call per query. |
MEMORY_RERANK_ENABLED | false | Honour rerank=true in tool calls (CrossEncoder reranker; needs the [rerank] extra). deep sets it to true. |
MEMORY_CROSS_ENCODER_ENABLED | false | Enable the torch CrossEncoder path used by deep. |
Paths, ports & transport
Where data lives and how the server is reached.
| Variable | Default | What it does |
|---|---|---|
TAM_MEMORY_DIR | ~/.tam | Data directory (database, logs, caches). CLAUDE_MEMORY_DIR is the older name and still works with a deprecation warning. A legacy ~/.claude-memory/ is moved to ~/.tam/ on first run and a symlink is left behind. |
DASHBOARD_PORT | 37737 | Port of the local web dashboard. |
DASHBOARD_BIND | 127.0.0.1 | Address the local dashboard listens on. A non-wildcard address is also accepted as a Host name. |
DASHBOARD_ALLOWED_HOSTSNew in 14.6.0 | empty | Comma-separated extra host names the local dashboard answers to. Requests with any other Host than loopback, DASHBOARD_BIND or these names get 421; a foreign Origin gets 403. Set it when you open the dashboard by a LAN name or from Docker. |
MCP_TRANSPORT | stdio | stdio for IDEs that start the server themselves; http / streamable-http to serve MCP over HTTP (the Docker path). |
MCP_HTTP_HOST | 127.0.0.1 | Bind address for the HTTP transport and for tam-team serve. |
MCP_HTTP_PORT | 3737 | Port for the HTTP transport and for tam-team serve. |
MCP_HTTP_ALLOWED_HOSTSNew in 14.6.0 | empty | Comma-separated extra host names the MCP HTTP transport accepts (DNS-rebinding protection). Host must be loopback, a non-wildcard MCP_HTTP_HOST or one of these (otherwise 421); a present Origin must match one of them (otherwise 403). Needed when clients reach /mcp by a LAN name or a Docker service name. |
MCP_HTTP_WORKERS | 1 | Number of HTTP server processes sharing the MCP port. |
TAM_NO_SETUPNew in 14.6.0 | unset | Set to 1 so a plain tam at a terminal never starts the setup wizard. |
TAM_SETUP_FILENew in 14.6.0 | <memory dir>/setup.json | Path of the setup record written by tam setup. |
MEMORY_REPORT_TZNew in 14.6.0 | TZ or the system zone | IANA time zone for the day, week and month boundaries of memory_report. |
FASTEMBED_CACHE_PATH | FastEmbed default | Where embedding models are cached. Docker images point it at a volume. |
LLM provider
Optional. Fast mode works without any LLM. These settings are used by background enrichment, memory_answer and other LLM-backed tools.
| Variable | Default | What it does |
|---|---|---|
MEMORY_LLM_ENABLED | auto | auto probes for a usable model, true/force always uses it, false disables every LLM call. |
MEMORY_LLM_PROVIDER | ollama | ollama, openai, openai-compatible, anthropic or auto (picks OpenAI, then Anthropic, from whichever API key is set). |
MEMORY_LLM_MODEL | qwen2.5-coder:7b (Ollama), gpt-4o-mini (OpenAI), claude-haiku-4-5 (Anthropic) | Model name. Required for openai-compatible. |
MEMORY_LLM_API_BASE | provider default | Base URL override. Required for openai-compatible (Chat Completions with JSON Schema support). |
MEMORY_LLM_API_KEY | unset | Key for the LLM provider. Takes priority over OPENAI_API_KEY / ANTHROPIC_API_KEY. Not inherited by openai-compatible. |
OPENAI_API_KEY | unset | Used by the openai LLM and embedding providers when no MEMORY_*_API_KEY is set. |
ANTHROPIC_API_KEY | unset | Used by the anthropic provider when MEMORY_LLM_API_KEY is not set. |
COHERE_API_KEY | unset | Used by the cohere embedding provider. |
DASHSCOPE_API_KEYNew in 14.6.0 | unset | Used by the dashscope embedding provider (Alibaba Cloud Model Studio) when MEMORY_EMBED_API_KEY is not set. Keys are bound to a region; see MEMORY_EMBED_API_BASE. |
OLLAMA_URL | http://localhost:11434 | Ollama address. In Docker, use an address the container can reach, e.g. http://host.docker.internal:11434. |
MEMORY_VISION_MODEL | unset | Model for describing images. Without it, no vision model is called automatically. |
MEMORY_LLM_PROBE_TTL_SEC | 60 | How long the "is the LLM reachable" probe result is cached. |
MEMORY_LLM_TIMEOUT_SEC | 60 | Fallback timeout for LLM requests, in seconds. |
MEMORY_TRIPLE_TIMEOUT_SEC | 30 | Timeout for knowledge-graph triple extraction. |
MEMORY_ENRICH_TIMEOUT_SEC | 45 | Timeout for deep enrichment. |
MEMORY_REPR_TIMEOUT_SEC | 60 | Timeout for generating record representations (summary, keywords, questions). |
MEMORY_TRIPLE_MAX_PREDICT | 2048 | Token cap for triple extraction. On CPU-only hosts lower this before raising timeouts. |
MEMORY_TRIPLE_PROVIDER | MEMORY_LLM_PROVIDER | Per-phase provider override for triple extraction. |
MEMORY_TRIPLE_MODEL | MEMORY_LLM_MODEL | Per-phase model override for triple extraction. |
MEMORY_ENRICH_PROVIDER | MEMORY_LLM_PROVIDER | Per-phase provider override for enrichment. |
MEMORY_ENRICH_MODEL | MEMORY_LLM_MODEL | Per-phase model override for enrichment. |
MEMORY_REPR_PROVIDER | MEMORY_LLM_PROVIDER | Per-phase provider override for representations. |
MEMORY_REPR_MODEL | MEMORY_LLM_MODEL | Per-phase model override for representations. |
MEMORY_REASON_PROVIDER | MEMORY_LLM_PROVIDER | Provider for the reasoning steps of memory_answer (reader, verifier, contradiction scoring). |
MEMORY_REASON_MODEL | MEMORY_LLM_MODEL | Model for those reasoning steps. |
Embeddings
Local FastEmbed (ONNX, no torch) by default. Changing the model of an existing store requires re-embedding.
| Variable | Default | What it does |
|---|---|---|
MEMORY_EMBED_PROVIDER | fastembed | fastembed (local), openai, cohere, or dashscope (since 14.6.0: Alibaba Cloud Model Studio, model text-embedding-v4, key in DASHSCOPE_API_KEY). Vector dimension must match the existing database or you must re-embed. |
MEMORY_EMBED_MODEL | provider default | Embedding model for cloud providers (e.g. text-embedding-3-small). Also an older alias that overrides the text model. |
MEMORY_EMBED_API_BASE | provider default | Base URL override for the embedding provider. For dashscope the default is the international (Singapore) endpoint https://dashscope-intl.aliyuncs.com/compatible-mode/v1; use https://dashscope.aliyuncs.com/compatible-mode/v1 for a Beijing-region key. |
MEMORY_EMBED_API_KEY | unset | Key for the embedding provider. |
MEMORY_EMBED_DIMENSIONSNew in 14.6.0 | 1024 (dashscope); model's native size otherwise | Requested vector size for openai and dashscope. For text-embedding-v4 allowed values are 2048, 1536, 1024, 768, 512, 256, 128 and 64. An invalid value stops the server instead of silently changing size; changing it on an existing store requires re-embedding. |
MEMORY_EMBED_BATCH_SIZENew in 14.6.0 | 10 (dashscope), 64 otherwise | Texts per embedding request for dashscope. The DashScope API accepts at most 10; larger values are refused. |
MEMORY_EMBED_MAX_RETRIESNew in 14.6.0 | 6 | Retries for the dashscope provider on HTTP 408, 425, 429 and 5xx and on network errors. Other 4xx errors (auth, quota) fail immediately. |
MEMORY_EMBED_TIMEOUT_SECNew in 14.6.0 | 60 | Per-request timeout for the dashscope provider, in seconds. |
MEMORY_EMBED_MAX_BACKOFF_SECNew in 14.6.0 | 30 | Upper limit for one retry delay, in seconds. A Retry-After header from the server is honoured up to this limit; otherwise the delay grows exponentially with jitter. |
MEMORY_TEXT_EMBED_MODEL | sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 | Model for ordinary records. For an English-only store BAAI/bge-base-en-v1.5 retrieves better but handles other languages poorly. Re-embed after changing it: python src/reembed.py --fastembed. |
MEMORY_CODE_EMBED_MODEL | jinaai/jina-embeddings-v2-base-code | Model for records classified as code. Empty falls back to the text model. |
MEMORY_LOG_EMBED_MODEL | empty (text model) | Model for records classified as logs. |
MEMORY_CONFIG_EMBED_MODEL | empty (text model) | Model for records classified as configuration. |
MEMORY_DEFAULT_EMBEDDING_SPACE | text | Space for content that is not classified. |
MEMORY_EMBED_THREADS | 1 | Compute threads per FastEmbed model. Raise only after measuring CPU and latency. |
MEMORY_TORCH_THREADS | 1 | PyTorch threads for the optional reranker and NLI models. Restart after changing. |
Recall & ranking
Tune how search results are ordered and filtered.
| Variable | Default | What it does |
|---|---|---|
MEMORY_CROSS_RERANK | auto | Local cross-encoder over the fused candidates of memory_recall. auto uses it once loaded in the background, on waits for it, off keeps the fused order. |
MEMORY_CROSS_RERANK_MODEL | Xenova/ms-marco-MiniLM-L-6-v2 | English-only models are skipped for queries mostly outside the Latin script; jinaai/jina-reranker-v2-base-multilingual covers other languages. |
MEMORY_CROSS_RERANK_WINDOW | 50 | How many fused candidates the cross-encoder scores (minimum 2). |
MEMORY_CROSS_RERANK_WEIGHT | 2.0 | Weight of the cross-encoder rank when it is fused with the other ranks. |
MEMORY_CROSS_RERANK_CONTEXT | 400 | Characters of the neighbouring turns read with each candidate. 0 scores each record alone. |
MEMORY_CONTEXT_RESOLVE_DATES | on | In memory_recall(mode="context"), annotate relative dates ("last Thursday [Thu 14 December 2023]"). off returns records as stored. |
MEMORY_RECALL_EXCLUDE | 1 | Filter operational records out of recall. 0 disables the filter. |
MEMORY_RECALL_EXCLUDED_TAGS | recovery,auto-extract | Comma-separated tags hidden from recall unless you ask for them. |
MEMORY_DECAY_PARENT_DAYS | 90 | Freshness half-life, in days, for results matched by full-text, fuzzy or graph search. |
MEMORY_DECAY_RAW_DAYS | 180 | Half-life for matches on the raw text. |
MEMORY_DECAY_KEYWORDS_DAYS | 90 | Half-life for matches on generated keywords. |
MEMORY_DECAY_COMPRESSED_DAYS | 60 | Half-life for matches on the compressed view. |
MEMORY_DECAY_QUESTIONS_DAYS | 45 | Half-life for matches on generated questions. |
MEMORY_DECAY_SUMMARY_DAYS | 30 | Half-life for matches on generated summaries. |
MEMORY_EDGE_HALF_LIFE_DAYS | 60 | Half-life of knowledge-graph edges when expanding context. |
Answers & contradictions
Behaviour of memory_answer, the optional grounded reader. It needs an LLM provider.
| Variable | Default | What it does |
|---|---|---|
MEMORY_NEGATIVE_RETRIEVAL | true | Run a second, contradiction-seeking search before answering. false skips it. |
MEMORY_CONTRADICTION_POLICY | resolve | resolve passes both sides with their dates to the reader, which answers with the latest value and names the one it replaced. abstain answers "Not enough information". |
MEMORY_CONTRADICTION_SCORER | llm | llm uses the reasoning provider; jev uses TypeSafe's Jev model and needs TYPESAFE_API_KEY. |
TYPESAFE_API_KEY | unset | Key for the jev scorer. |
TYPESAFE_BASE_URL | https://api.typesafe.ai | Endpoint for the jev scorer. |
TYPESAFE_DEFAULT_MODEL | jev-latest | Model for the jev scorer. |
Obsidian projection
TAM can write a human-readable activeContext.md for each project into an Obsidian vault.
| Variable | Default | What it does |
|---|---|---|
MEMORY_ACTIVECONTEXT_VAULT | ~/Documents/project/Projects | Vault folder the projection is written to. |
MEMORY_ACTIVECONTEXT_DISABLE | unset | Set to 1 to stop writing the projection. |
Team server & remote client
Read by tam-team, docker-compose.team.yml and the tam-remote bridge.
| Variable | Default | What it does |
|---|---|---|
TAM_TEAM_DIR | ~/.tam-server | Team server data directory (same as tam-team --root). |
TAM_TEAM_MAX_WORKERS | 3 | Workspace processes kept warm (personal, team, shared). More uses more RAM; fewer means more cold loads. |
TAM_TEAM_OPERATION_TIMEOUT | 120 | Seconds one memory operation may take. The server never retries a write automatically after a timeout. |
TAM_TEAM_PORT | 3738 | Docker Compose: host port for the web interface and /mcp/. |
TAM_TEAM_PUBLISH_HOST | 127.0.0.1 | Docker Compose: host address the port is published on. Keep it on loopback behind a reverse proxy. |
TAM_TEAM_HOST_GATEWAY | host.docker.internal:host-gateway | Docker Compose: host alias so the container can reach a model server on the Docker host. |
TAM_TEAM_MASTER_KEYNew in 14.6.0 | <data dir>/master.key | Fernet key that encrypts provider API keys saved in the dashboard. Back it up separately from tam-team backup. |
TAM_TEAM_INVITE_TTL_HOURSNew in 14.6.0 | 72 | Lifetime of dashboard invite codes. |
TAM_TEAM_SESSION_IDLE_MINUTESNew in 14.6.0 | 30 | Dashboard session idle timeout. |
TAM_TEAM_SESSION_MAX_HOURSNew in 14.6.0 | 12 | Maximum dashboard session length. |
TAM_TEAM_LOGIN_MAX_FAILURESNew in 14.6.0 | 5 | Failed sign-ins for one user from one IP within 15 minutes before a lockout. |
TAM_TEAM_LOGIN_LOCK_MINUTESNew in 14.6.0 | 15 | Length of a sign-in lockout. |
TAM_TEAM_COOKIE_SECURENew in 14.6.0 | auto | auto marks the session cookie Secure on HTTPS requests; always or never force it. |
TAM_TEAM_TRUST_PROXYNew in 14.6.0 | false | Honour X-Forwarded-Proto and X-Forwarded-For from a reverse proxy you control. |
TAM_TEAM_SETUP_TOKEN_TTL_MINUTESNew in 14.6.0 | 60 | Lifetime of the one-time code for the web setup wizard. |
TAM_REMOTE_URL | required | Client bridge: the HTTPS /mcp/ endpoint. Plain HTTP is accepted only for loopback addresses. |
TAM_REMOTE_TOKEN_FILE | required | Client bridge: path to the personal token file. Re-read on every request. |
TAM_REMOTE_TIMEOUT | 180 | Client bridge: request timeout in seconds. |
Experimental (v9) switches
Older experimental knobs that are still read. Leave them at their defaults unless you are benchmarking.
| Variable | Default | What it does |
|---|---|---|
V9_PARALLEL_RETRIEVAL | false | Run retrieval tiers in parallel. |
V9_CACHE_L1_ENABLED | false | In-process result cache. |
V9_CACHE_L1_SIZE | 1000 | Entries in that cache. |
V9_CACHE_L1_TTL_SEC | 300 | Lifetime of cache entries, in seconds. |
V9_CACHE_L2_ENABLED | false | Second-level cache. |
V9_EMBED_BACKEND | fastembed | Embedding backend selector used by the v9 benchmark wiring. |
V9_LOCOMO_TUNED_PATH | ./models/locomo-tuned-minilm | Path to a locally fine-tuned embedding model for experiments. |
V9_RERANKER_BACKEND | ce-marco | Reranker backend selector for the v9 path. |
V9_RERANKER_MODEL | unset | Reranker model override for the v9 path. |
V9_RERANKER_FP16 | 1 | Use half precision for the v9 reranker when supported. |