⚠️ Disclaimer: This project is not production-ready. It is currently an educational project intended for learning, experimentation, and research purposes only. Do not use it in production environments or for critical workloads.
memory_mcp is a Rust-based Model Context Protocol (MCP) server that gives AI agents a structured long-term memory layer backed by SurrealDB.
It is designed for workflows where agents need more than short-lived chat context: episodic memory, extracted entities and facts, bi-temporal validity, ranked context assembly, and graph-style relationships between people, companies, tasks, and decisions.
- Overview
- What it provides
- Architecture
- Quick start
- Streamable HTTP SaaS profile
- Configuration
- MCP tools
- Development
- Testing
- Project layout
- Documentation
- Contributing
- License
Memory MCP implements a memory system for AI agents with core goals:
- preserve important source material as episodes
- extract entities, facts, and links in a deterministic way
- track knowledge over both valid time and transaction time
- assemble compact, relevant context for downstream reasoning
- support policy-tag-aware retrieval and access filtering within one Active Namespace
In practice, an agent can ingest emails, notes, or working documents, resolve entities consistently, store facts with provenance, and later ask for ranked context instead of replaying entire histories.
- Bi-temporal knowledge model for valid time and ingestion time
- Episode ingestion for storing raw source material
- Entity resolution with alias handling and deterministic IDs
- Fact extraction for metrics, promises, and other structured knowledge
- Context assembly for ranked retrieval by query, policy tags, and time cutoff
- Graph relationships between episodes, entities, and facts
- Optional semantic retrieval providers including in-process
local-candle - Pluggable NER backends for entity extraction:
anno,regex, explicit Anno NuNER ONNX, and two native Candle zero-shot GLiNER backends (selectable viaNER_EXTRACTOR) - SurrealDB support for embedded and remote deployments
- Optional filesystem ingestion inside
servefor filesystem-backed auto-ingest workflows (activated byMEMORY_INGESTION_INBOX) - MCP-native interface for tool-driven agent workflows
- Structured logging with predictable operational behavior
Memory MCP is a layered system with a narrow service seam between protocol adapters and domain logic. The MCP and CLI interfaces share the same protocol-agnostic capabilities, so behavior does not diverge between an agent calling a tool and an operator running a command locally.
flowchart TD
Agent["AI agent / MCP client"]
Hooks["Agent host hooks\nstop / precompact"]
Operator["Operator / CI"]
Agent -->|"stdio JSON-RPC"| MCP["MCP protocol layer\nhandlers, params, parsers, sessions"]
Hooks -->|"hidden lifecycle CLI\nsubcommands"| CLI["CLI layer\nserve, reembed, init"]
Operator --> CLI
MCP --> Tools["Protocol-agnostic tools\ningest, extract, resolve, retrieve, explain, invalidate"]
CLI --> Tools
Tools --> Capabilities["Capabilities\nsmall use-case adapters"]
Capabilities --> Context["ServiceContext\nnarrow dependency seam\nrate limiting + stores + providers"]
Context --> Domain["Domain services\ningestion, facts, entities, claims,\nembeddings, lifecycle, procedures"]
Context --> Retrieval["Context pipeline\nlexical, semantic, graph, community,\ntemporal filtering and ranking"]
Domain --> Storage["Storage abstraction\nnarrow stores + append-only migrations"]
Retrieval --> Storage
Storage --> DB[("SurrealDB\nActive Namespace")]
Important boundaries
main.rsis intentionally thin: argument parsing and dispatch only.mcp/is a protocol adapter; business logic stays inservice/.tools/andservice/capabilities/are reusable from both MCP and CLI.- Storage is selected once at startup. Requests do not choose a namespace.
- Facts and claims are never deleted. They are invalidated while preserving historical traceability.
The write path turns source material into durable, structured memory. Ingestion is deterministic and idempotent: sending the same source again returns the existing episode instead of creating a duplicate.
flowchart TD
Source["Raw source\nemail, note, document, file"] --> Ingest["ingest"]
Ingest --> Limit["RateLimiter.check_access\nper-caller token bucket"]
Limit --> Parse["Prepare content\nPDF / HTML / plaintext"]
Parse --> EpisodeID["Deterministic episode ID v2"]
EpisodeID --> Existing{"Episode already exists?"}
Existing -->|"yes"| Episode["Return episode:<id>\nidempotent dedupe"]
Existing -->|"no"| PersistEpisode["Persist episode\nt_ref + t_ingested"]
PersistEpisode --> Episode
Episode --> Extract["extract(episode_id)"]
Extract --> NER["Entity extraction\nanno, regex, anno-onnx,\nclassic GLiNER, LFM2 GLiNER"]
NER --> Resolve["Entity resolution\naliases -> canonical IDs"]
Extract --> Facts["Fact extraction\nstructured facts + provenance"]
Facts --> Triples["Triple extraction\nsubject / predicate / object"]
Facts --> Embeddings["Embedding generation\ncached + background retry"]
Facts --> Claims["Claim projection\nvalue, qualifiers, validity, source_span"]
Claims --> Reconcile["Claim reconciliation\nmatch, upsert, retract, backfill"]
Reconcile --> Store[("Durable memory\nSurrealDB")]
Resolve --> Store
Triples --> Store
Embeddings --> Store
Facts --> Store
The claim pipeline preserves provenance: source_span points back to the
source range that produced a claim, while remaining outside deterministic
claim identity. This lets metadata improve traceability without changing
whether two claims are considered the same claim.
The read path fuses several retrieval strategies, then applies policy, temporal, provenance, and budget constraints before returning a compact context pack.
flowchart TD
Query["assemble_context\nquery + budget + flags"] --> Limit["RateLimiter.check_access"]
Limit --> Cache{"Context cache hit?"}
Cache -->|"yes"| Cached["Return cached context"]
Cache -->|"no"| Prepare["Normalize parameters\nexpand aliases, resolve mode"]
Prepare --> Lexical["Lexical retrieval\nterm and field matches"]
Prepare --> Semantic["Semantic retrieval\nquery embeddings + similarity"]
Prepare --> Graph["Graph retrieval\nentity links, triples, bounded hops"]
Prepare --> Community["Community retrieval\nentity/community summaries"]
Prepare --> Experience["Experience retrieval\nrepeated topics and preferences"]
Lexical --> Fuse["Fuse candidates"]
Semantic --> Fuse
Graph --> Fuse
Community --> Fuse
Experience --> Fuse
Fuse --> Filter["Filter\nvalid time, policy tags, access"]
Filter --> Rank["Rank\nrelevance, decay, source priority,\nsemantic score, temporal focus"]
Rank --> Select["Select\nbudget, per-source caps, grounding"]
Select --> Shape["Shape response\nranked or timeline view"]
Shape --> Track["Record fact access\nrecency feedback"]
Track --> Result["Context items\ncontent + rationale + provenance"]
Result --> Explain["explain\ncitation-ready source snippets"]
A query can therefore succeed even when one retrieval signal is weak: lexical matches, semantic similarity, graph expansion, community summaries, and experience candidates are fused before ranking. Each result includes enough rationale and provenance for an agent to decide whether to use it.
Memory distinguishes when something was true from when the system learned it. This is essential for correcting stale knowledge without erasing the historical record.
erDiagram
EPISODE ||--o{ FACT : yields
FACT }o--o{ ENTITY : links
FACT ||--o{ CLAIM : projects
FACT ||--o{ TRIPLE : produces
ENTITY }o--o{ COMMUNITY : belongs_to
EPISODE {
string episode_id PK
string source_type
datetime t_ref
datetime t_ingested
}
FACT {
string fact_id PK
string content
array entity_links
array embedding
datetime t_valid
datetime t_invalid
}
CLAIM {
string claim_id PK
string schema_ref
string value
array source_span
datetime t_valid
datetime t_invalid
}
ENTITY {
string entity_id PK
string canonical_name
array aliases
}
TRIPLE {
string triple_id PK
string subject
string predicate
string object
}
COMMUNITY {
string community_id PK
string summary
}
t_ref/t_valid: when the source says the information is true.t_ingested: when Memory MCP recorded the source.t_invalid: when a fact or claim was superseded; the row remains available for audit and historical queries.
| Module | Purpose |
|---|---|
mcp |
MCP handlers, params, parsers, and tool-facing types (stdio profile) |
service |
Core business logic for ingest, extract, retrieval, graph operations, and validation |
storage |
Database integration and persistence helpers (one Active Namespace) |
models |
Shared domain models and request/response types |
config |
Environment-driven configuration loading |
logging |
Logging setup and log-level utilities |
observability |
Optional Prometheus installation and bounded runtime metrics |
tools |
Protocol-agnostic tool implementations shared by MCP and CLI |
cli |
CLI subcommand adapters and lifecycle hooks |
http |
Streamable HTTP composition root (SaaS profile, feature-gated) |
control |
OIDC, browser sessions, account/operator API, CSRF, and deletion flow (feature-gated) |
- Rust 1.97.1+ only when compiling from source (matches workspace
rust-version) - No external SurrealDB service is required for the default embedded mode
- Download the release asset for your platform and verify its accompanying SHA-256 checksum. Rename it to
memory_mcp(or create a symlink with that name); the Windows asset already includes.exe. - Put the renamed executable on
PATH. - Run
memory_mcp initfor the default VS Code snippet, or pass one of the exact targetsclaude-desktop,codex,zed, orenv. - Copy the printed host-native snippet into the indicated configuration file.
- Ingest one source, run
extract --episode-id <episode-id>, then runassemble-contextto verify a real fact is recalled.
The default path needs no environment variables, configuration file, external database, API key, network request, or model download. It uses a user-owned embedded database, Anno extraction, and lexical/graph retrieval immediately. memory_mcp init prints configuration only: it does not edit host files, change environment variables, start a database, download models, or access the network.
A Rust toolchain is needed only when a prebuilt release is unavailable:
cargo install --path crates/memory-mcp --lockedThis builds the same full-capability application as the release binary; it is not a reduced onboarding build.
The clean-machine harness measures the path from the selected persona's start to a
real fact recalled by assemble-context; it does not measure GUI host startup or
memory quality. It uses isolated HOME, XDG_DATA_HOME, CARGO_HOME, and working
directories and prints machine-readable timings with median and p90 aggregates:
scripts/measure_ttv.sh --binary ./target/release/memory_mcp --persona release-binary --repeat 5
scripts/measure_ttv.sh --binary ./target/release/memory_mcp --persona host-config-user --repeat 5
scripts/measure_ttv.sh --cargo-install --source . --persona rust-user --repeat 5The fixture is a summary-like requirement episode because the existing extractor
intentionally limits note fallback facts to summary-capable source types. The
validator rejects malformed responses, empty fact arrays, and episode-only fallback
items, so a run is successful only when a persisted fact—not an episode fallback—is
recalled. Installation, host-snippet preparation, storage initialization, episode
write, extraction, and fact recall are reported separately. A median total of
<= 300 seconds is the measured target, not a guarantee; the first clean rust-user
run on this macOS workspace took 544.098 seconds, with 542.634 seconds spent in
isolated compilation/install and approximately 1.46 seconds in the application
path.
cargo run --release -- serve
# or
make serve-releaseFor local NER workloads, run the MCP server from a release build. The development
profile leaves the memory_mcp crate at opt-level = 0; dependency code is optimized,
but GLiNER window orchestration and span enumeration are not. Performance claims and
timeout investigations are valid only for release builds. Use cargo run only for
development and functional debugging.
The binary uses stdio transport, which makes it suitable for local MCP client integration.
The default embedded mode needs no SURREALDB_* variables. To select a remote
SurrealDB explicitly, use one of the supported remote schemes (ws, wss,
http, or https) and provide non-empty credentials:
SURREALDB_URL=ws://127.0.0.1:8000/rpc \
SURREALDB_EMBEDDED=false \
SURREALDB_DB_NAME=memory \
SURREALDB_NAMESPACE=org \
SURREALDB_USERNAME=<your-remote-username> \
SURREALDB_PASSWORD=<your-remote-password> \
RUST_LOG=info \
cargo run --quiet --bin memory_mcpmem:// and rocksdb:// are not remote URL schemes. For an explicit local
RocksDB location, set SURREALDB_DATA_DIR; otherwise the server uses a
user-owned data directory by default.
Filesystem ingestion turns a directory into a passive memory intake pipe:
drop or save files into the configured inbox and the stdio MCP server ingests
them through the full ingest → extract pipeline without manual tool calls.
Activation
Set MEMORY_INGESTION_INBOX to an existing absolute directory when starting
serve (the variable is optional; when absent, startup behavior is unchanged).
The binary must be built with the fs-watch feature (official release binaries
include it). A binary compiled without the feature rejects a configured inbox
with an actionable startup error.
# Single terminal — the MCP server owns filesystem ingestion
RUST_LOG=info \
MEMORY_INGESTION_INBOX=$HOME/projects/atlas/inbox \
SURREALDB_DATA_DIR=$HOME/.memory-mcp/atlas \
memory_mcp serveWhat it does
- Startup validates the inbox, attaches the OS watcher, then scans existing supported files in the background (files cannot fall into a scan-to-watch gap)
- Watches the inbox recursively for file create and modify events
- Processes files only after size and modification time stabilize
- Skips symlinks and unsupported file types silently
- Tracks durable revisions: each distinct set of bytes at a path is one immutable revision; renaming a file starts a new lineage; deleting a file never invalidates memory
- One failed file never stops ingestion or MCP; the watcher backend is recreated with bounded backoff and then enters a logged degraded state
Supported file types
| Extension | Format | Extracted content |
|---|---|---|
.pdf |
Text content (pages, paragraphs) | |
.docx |
Word document | Body text, headings, tables |
.xlsx |
Spreadsheet | Cell values, sheet structure |
.pptx |
Presentation | Slide text, speaker notes |
.md, .markdown |
Markdown | Headings, lists, code blocks |
.txt |
Plain text | Raw text content |
.eml |
Email message | Subject, sender, recipients, body, date |
Files with other extensions (.json, .png, .zip, etc.) are silently skipped.
MCP host example (Zed)
{
"context_servers": {
"memory_mcp": {
"command": "memory_mcp",
"args": [],
"env": {
"MEMORY_INGESTION_INBOX": "/absolute/path/to/inbox",
"SURREALDB_DATA_DIR": "/absolute/path/to/atlas-data"
}
}
}
}MCP host example (Claude Desktop)
{
"mcpServers": {
"memory_mcp": {
"command": "memory_mcp",
"args": [],
"env": {
"MEMORY_INGESTION_INBOX": "/absolute/path/to/inbox",
"SURREALDB_DATA_DIR": "/absolute/path/to/atlas-data"
}
}
}
}Each stdio client process needs its own SURREALDB_DATA_DIR; changing only the
database name or namespace does not avoid the embedded directory lock.
The repository also contains an optional app-oriented MCP surface for reviewer and inspector workflows. It is intentionally feature-gated so the eight canonical memory tools stay available without exposing extra session/resource endpoints by default.
Build or run with apps enabled:
cargo run --features mcp-apps -- serveRecommended verification for this surface:
cargo check --all-targets --features mcp-apps
cargo clippy --all-targets --features mcp-appsHow it works internally
Architecture flow
serve (stdio MCP) with MEMORY_INGESTION_INBOX set
│
▼
FsWatchRuntime::start(service, config)
│
├─ Validate: inbox must be absolute, readable, not a symlink
├─ Attach OS watcher (watcher-first) — before the scan starts
├─ Requeue: expired leases + one retry cycle for failed revisions
│
├─ Spawn event bridge: forwards Create/Modify events
├─ Spawn startup scan: enqueues existing supported files recursively
└─ Spawn sequential processor: drains the durable inbox revision store
SHARED DISCOVERY (both event bridge and scan)
│
├─ prepare_candidate: reject symlinks + unsupported extensions,
│ wait for size + mtime stability
├─ Hash raw bytes (SHA-256) → immutable revision identity
├─ discover_prepared: persist durable prepared-content snapshot
│
└─ Processor (sequential, lease-based):
├─ ingest → extract (from the durable snapshot, never the path)
├─ retry transient failures (bounded, exponential)
└─ mark processed; failed revisions requeue once per startup
Revision and deduplication behavior
How rapid saves are handled
When you save a file, editors often fire multiple filesystem events in quick succession (write + metadata + timestamp). Files are processed only after size and modification time stabilize (two consecutive matching samples), and each distinct set of raw bytes becomes exactly one immutable inbox revision:
- Revision identity is SHA-256 over the raw bytes plus the normalized lineage (path relative to the inbox)
- Re-scanning or re-observing identical bytes returns the existing revision — no duplicate episode or facts
- A file that changes creates a new revision (new episode, same
source_lineage) - Renaming a file starts a new lineage (new episode source lineage); deleting a file never invalidates memory
Command-line reference
Activation
MEMORY_INGESTION_INBOX=/absolute/path/to/inbox memory_mcp serve
The variable must be a non-empty absolute path to an existing readable
directory that is not a symlink. Omit it to keep filesystem ingestion
disabled. The binary must include the `fs-watch` feature (official release
binaries do); a binary compiled without it rejects a configured inbox with
an actionable startup error.
Important notes:
- One process watches exactly one inbox recursively; symlinks and unsupported files are skipped.
- Files are processed only after size and modification time stabilize.
- Each distinct set of bytes is one immutable revision; renaming starts a new lineage; deleting a file never invalidates memory.
- One failed file or a degraded watcher backend never stops MCP or queued work.
Logging during filesystem ingestion
What to expect at each log level
| Level | Events you'll see |
|---|---|
info |
fs_watch.ready (startup), fs_watch.revision (per-revision outcome: relative path + short revision prefix only) |
warn |
fs_watch.degraded (watcher backend exhausted after bounded backoff) |
debug |
fs_watch.shutdown outcome on clean exit |
Revision events contain relative paths and short revision prefixes only; file contents and absolute inbox roots never appear in logs except startup diagnostics.
Run the renderer for the current VS Code schema and copy its JSON into
.vscode/mcp.json:
memory_mcp init --target vscodeThe generated snippet uses servers.memory_mcp with a stdio command of
memory_mcp and no environment variables. After cargo build --release or
cargo install --path crates/memory-mcp --locked, the installed binary can be
used directly by the host.
memory_mcp_http is the multi-user, remote deployment of Memory MCP. Each
authenticated request maps to one Account and one Tenant, then runs against
that Tenant's dedicated storage. Account, tenant, and namespace identities
stay independent; the namespace that holds a Tenant's data is server-generated
and never selectable from a request.
This section is the operator's quick reference. Full environment contract, request validation, and release gates live in the Streamable HTTP SaaS specification, ADR-0052, and the operations runbooks.
Build with the streamable-http feature. Add control-plane for OIDC and
account management, and control-plane-ui to embed the Dioxus SPA.
cargo build --release --locked --features streamable-http,control-plane
# HTTP boundary
MEMORY_MCP_HTTP_PUBLIC_BASE_URL=https://mcp.example.com \
ALLOWED_HOSTS=mcp.example.com \
ALLOWED_ORIGINS=https://mcp.example.com \
MEMORY_MCP_HTTP_TRUSTED_PROXY_CIDRS=10.0.0.0/8,172.16.0.0/12 \
# SurrealDB control (Account/plan/credential Registry) and tenant engine
SURREALDB_CONTROL_URL=wss://surreal.example.com/rpc \
SURREALDB_CONTROL_USERNAME=... \
SURREALDB_CONTROL_PASSWORD=... \
SURREALDB_CONTROL_NAMESPACE=control \
SURREALDB_CONTROL_DB=registry \
SURREALDB_TENANT_URL=wss://surreal.example.com/rpc \
SURREALDB_TENANT_USERNAME=... \
SURREALDB_TENANT_PASSWORD=... \
SURREALDB_TENANT_NAMESPACE=tenant \
SURREALDB_TENANT_DB=tenant \
# Keyed verifiers (32-byte hex; raw secrets are never persisted or logged)
MEMORY_MCP_API_KEY_PEPPER=... \
MEMORY_MCP_HTTP_IDENTITY_INDEX_KEY=... \
MEMORY_MCP_HTTP_SESSION_KEY=... \
MEMORY_MCP_HTTP_OIDC_STATE_KEY=... \
MEMORY_MCP_HTTP_OIDC_NONCE_KEY=... \
MEMORY_MCP_HTTP_CSRF_KEY=... \
# Signup policy and OIDC
MEMORY_MCP_HTTP_SIGNUP_MODE=invite_only \
MEMORY_MCP_HTTP_OIDC_ISSUER=https://issuer.example.com \
MEMORY_MCP_HTTP_OIDC_CLIENT_ID=memory_mcp \
MEMORY_MCP_HTTP_OIDC_AUDIENCE=https://mcp.example.com \
MEMORY_MCP_HTTP_OIDC_REDIRECT_URI=https://mcp.example.com/auth/oidc/callback \
# Stable replica identity (required for multi-replica deployments)
MEMORY_MCP_HTTP_REPLICA_ID=node-a \
./target/release/memory_mcp_http| Route | Auth | Use |
|---|---|---|
POST /mcp |
Bearer API key | Modern MCP Streamable HTTP (2026-07-28). Only POST is accepted; GET/DELETE return 405. |
/api/v1/account/* |
Browser session + CSRF | Self-service: API keys, profile, account deletion |
/api/v1/operator/* |
OIDC operator + CSRF + recent-auth | Operator-only: provisioning retry, suspend, purge, recovery |
/auth/oidc/* |
OIDC flow | Login, callback, logout (only when the control plane is enabled) |
/health/live, /health/ready |
Public | Process liveness and admission readiness |
/metrics |
Public (no app auth) | Prometheus scrape. Restrict at the reverse proxy or network layer. |
/ and SPA fallback |
Public | Dioxus control-plane UI (only with the control-plane-ui feature) |
MCP requests authenticate with a Bearer API key issued per Account. The
server-generated key has the shape mem_sk_<key_id>_<256-bit-secret>; the
secret is shown once and never recoverable. Each Account can hold up to ten
active keys. Revocation is durable immediately and externally effective
within the documented bound. Account APIs cannot grant operator status.
Browser control-plane endpoints authenticate with a secure server-side session cookie (HttpOnly, SameSite). The session has both an idle and an absolute expiry and rotates after login. Destructive actions require recent authentication, typically within ten minutes.
OIDC is the operator path. When the control plane is enabled, the binary
requires MEMORY_MCP_HTTP_OIDC_* configuration. Enabling the control plane
without it is a startup error. OIDC login uses Authorization Code with PKCE,
exact issuer/audience validation, encrypted state and nonce, and an algorithm
allowlist. The browser session and MCP API keys are independent: a browser
session never authenticates POST /mcp, and an API key never authenticates the
control plane.
Operator access is granted only through MEMORY_MCP_HTTP_OPERATOR_IDENTITIES,
an immutable allowlist of issuer|hex(subject_verifier) entries. Operators
audit, retry, suspend, resume, purge, and inspect recovery status; nothing
else.
The backend listens on 0.0.0.0:8080 by default. Production terminates TLS at
a reverse proxy. The proxy must:
- Disable response buffering for
POST /mcpand set the proxy's read timeout above the ordinary request deadline - Apply a separate, longer idle policy to long-lived
subscriptions/listenstreams; the ordinary deadline does not apply to them - Enforce host and origin from the configured allowlists
- Restrict
/metricsto your metrics network - Not rewrite MCP mirrored headers
Wildcard origins are rejected at startup. Missing Origin is accepted only
for non-browser MCP clients.
The HTTP profile enforces admission limits at the edge (per-tenant and
global), separate subscription and ordinary-request budgets, and a tunable
runtime pool. Quota overage returns HTTP 429 with a Retry-After header
and a guidance field. Public launch requires remote SurrealDB and explicit
quota values for signup_mode=open:
MEMORY_MCP_HTTP_MAX_INGESTED_BYTES
MEMORY_MCP_HTTP_MAX_EPISODE_COUNT
MEMORY_MCP_HTTP_INGEST_PER_MINUTE
MEMORY_MCP_HTTP_MAX_OPEN_APP_SESSIONS
MEMORY_MCP_HTTP_MAX_ACTIVE_API_KEYS
MEMORY_MCP_HTTP_PER_TENANT_REQUEST_CONCURRENCY
MEMORY_MCP_HTTP_EXTRACTION_CONCURRENCY
These seed Registry plan version 1 only when no plan exists. An existing durable plan is never overwritten.
Two features run long-lived response streams:
subscriptions/listen: filtered, ordered change events for Apps and resources. The server reauthorizes at least every 30 seconds, bounded to 60 seconds. Revocation terminates the stream within that bound.extractwith theio.modelcontextprotocol/tasksextension:extractis the only tool that may return a Task. Clients that do not advertise the extension receive a bounded synchronous result, or a preflight rejection when the work exceeds the configured size limit.
App Sessions, durable Tasks, and the change-event outbox have separate,
configurable retention and queue limits. The full list of MEMORY_MCP_HTTP_*
tuning variables lives in the specification.
The user-driven deletion flow is irreversible. After recent OIDC re-authentication, the user confirms with a typed phrase and a one-use token. The server revokes live API keys and browser sessions, transitions the Account and Tenant to terminal deletion states, and preserves a durable tombstone so the namespace binding is never reused. Memory records are invalidated; only ephemeral Task and App Session rows are physically removed after their normal retention/TTL window. There is no cancellation window.
Set MEMORY_MCP_HTTP_REPLICA_ID to a stable deployment identity. Without it,
the process falls back to a PID-based identity, which is safe only for a
single process. Workers across replicas coordinate durable work through
fenced leases; a stable replica ID makes lease ownership observable.
mem:// is test-only. rocksdb:// works for development, demos, and
single-process tests. Public production requires remote SurrealDB. The
embedded profile emits a startup warning and is not HA, not rolling, and
not horizontally scalable.
Confirm the reverse proxy enforces TLS, host, and origin. Confirm
/metrics is restricted. Run the conformance suite. Complete the
operations runbooks (restore drill, credential rotation)
and the §20.5 release evidence for open signup.
Configuration is loaded from environment variables.
| Variable | Type | Default | Required | Description |
|---|---|---|---|---|
SURREALDB_DB_NAME |
string | memory |
No | Database name |
SURREALDB_NAMESPACE |
string | main |
No | One namespace; changing it takes effect after restart and never moves data |
SURREALDB_USERNAME |
string | root (embedded); explicit value required (remote) |
Remote only | Database username |
SURREALDB_PASSWORD |
string | root (embedded); explicit value required (remote) |
Remote only | Database password |
SURREALDB_URL |
URL | unset (embedded) | Remote only | Remote connection URL using ws, wss, http, or https |
SURREALDB_EMBEDDED |
boolean | inferred from SURREALDB_URL |
No | Explicit true/false; remote URLs select remote mode and all other URLs select embedded mode when unset |
SURREALDB_DATA_DIR |
path | $XDG_DATA_HOME/memory_mcp; else $HOME/.local/share/memory_mcp; else ./.memory_mcp (embedded); unset (remote config) |
No | Custom embedded data directory; an existing executable-relative data/surrealdb directory may be reused for compatibility, and the effective default root also backs local model caches |
SURREALDB_EMBEDDING_DIMENSION |
unsigned integer | unset | No | Existing vector dimension override; the provider fallback is 384 for local-candle and 1536 for other embedding providers |
The default local path is embedded RocksDB with no external service, credential,
or model download required to start storage. Remote mode requires a valid URL and
non-empty explicit username and password. NER defaults to the in-process Anno
backend; zero-config does not mean the binary has no dependencies.
The following settings are optional for power users. They are read by the same executable used by the no-configuration quick start.
| Variable | Type | Default | Description |
|---|---|---|---|
RUST_LOG |
string | info |
Logging level; canonical values are trace, debug, info, warn, and error; warning aliases warn, and unknown values fall back to info |
MEMORY_LOG_FILE |
path | unset | Write structured log events to this file instead of stderr; the file is created if missing (parent directory must exist), opened in append mode, and flushed after every line; on open failure the process falls back to stderr with a warning |
MEMORY_PROMETHEUS_LISTEN_ADDR |
socket address (IP:port) |
unset | Prometheus HTTP listener address; active only when the prometheus feature is compiled and this variable is set |
QUERY_LOGGING_ENABLED |
boolean | false |
Persist assemble_context analytics rows into query_log when true |
QUERY_LOG_RETENTION_DAYS |
unsigned integer | 90 |
Days to retain persisted query_log analytics before best-effort pruning |
LIFECYCLE_ENABLED |
boolean | false |
Enable background lifecycle jobs |
LIFECYCLE_DECAY_INTERVAL_SECS |
unsigned integer | 3600 |
Decay worker interval in seconds |
LIFECYCLE_ARCHIVAL_INTERVAL_SECS |
unsigned integer | 86400 |
Archival worker interval in seconds |
LIFECYCLE_DECAY_THRESHOLD |
floating-point number | 0.3 |
Confidence threshold for fact invalidation |
LIFECYCLE_ARCHIVAL_AGE_DAYS |
unsigned integer | 90 |
Days before archiving episodes |
LIFECYCLE_DECAY_HALF_LIFE_DAYS |
floating-point number | 365 |
Half-life in days for decay computation |
EMBEDDINGS_ENABLED |
boolean | false when unset and no provider is set; true when a provider is set and this variable is unset |
Enable semantic retrieval; explicit false takes precedence over provider selection |
EMBEDDINGS_PROVIDER |
string enum | disabled when both variables are unset; local-candle when EMBEDDINGS_ENABLED=true without a provider |
Embedding backend: local-candle, openai-compatible, or ollama; when EMBEDDINGS_ENABLED is unset, setting a provider enables embeddings, while explicit false disables them and explicit true enables the selected/default provider |
EMBEDDINGS_MODEL |
string | intfloat/multilingual-e5-small for local-candle; required for external providers when enabled |
Model identifier for the selected embedding provider |
EMBEDDINGS_MODEL_DIR |
path | unset (derived under the effective data/cache root for local-candle) |
Optional local cache directory for local-candle |
EMBEDDINGS_BASE_URL |
URL | unset for local-candle; https://api.openai.com/v1 for openai-compatible; http://127.0.0.1:11434 for ollama |
Base URL for remote embedding providers |
EMBEDDINGS_MAX_TOKENS |
unsigned integer | 384 |
Max token budget before local-candle chunks long inputs |
EMBEDDINGS_TIMEOUT_SECS |
unsigned integer | 15 |
Timeout for remote embedding calls |
EMBEDDINGS_RECOVERY_INTERVAL_SECS |
positive unsigned integer | 60 |
Initial delay before the in-process recovery worker probes a remote provider after degraded startup; failed probes use exponential backoff |
EMBEDDINGS_AUTO_RECOVERY |
boolean | true |
Enable automatic in-process recovery after a failed remote startup preflight; set false for explicit opt-out |
EMBEDDINGS_SIMILARITY_THRESHOLD |
floating-point number | 0.7 |
Minimum cosine similarity for semantic matches |
EMBEDDINGS_API_KEY |
string | unset | Optional bearer token for OpenAI-compatible providers |
NER_EXTRACTOR |
string enum | anno (unset) |
Entity extraction backend selector. Closed catalog: anno (lightweight, download-free), regex (project-owned deterministic), anno-onnx (Anno NuNER ONNX, local-path only), urchade/gliner_multi-v2.1 (classic Candle GLiNER), VAGOsolutions/SauerkrautLM-LFM2.5-GLiNER (native Candle LFM2 GLiNER). Unknown values and arbitrary repository IDs are rejected. The removed NER_PROVIDER and NER_MODEL variables fail with migration guidance if present |
NER_CACHE_DIR |
path | <data>/models/ner |
Artifact store root for model-backed extractors (Anno ONNX, classic GLiNER, VAGO LFM2) |
NER_LABELS |
comma-separated list | person, company, location, product, event, technology |
Runtime labels for model-backed extractors; trimmed, lowercased, deduplicated in first-declared order |
NER_THRESHOLD |
floating-point number | 0.5 |
Confidence threshold for model-backed extractors (each backend owns an evaluated default; explicit in-range values override it) |
NER_MAX_CONCURRENCY |
positive integer | 1 |
Concurrent local NER inference limit |
NER_IDLE_UNLOAD_SECS |
unsigned integer | 0 |
Seconds of inactivity before any model-backed extractor unloads its model; 0 keeps it loaded for the process lifetime |
GLINER_BATCH_SIZE |
positive integer | 1 |
Max windows per transformer forward pass; increase only after workload-specific benchmarking |
GLINER_MAX_BATCH_TOKENS |
positive integer | 1536 |
Max padded tokens per batch |
GLINER_DEVICE |
string enum | cpu |
Device for the native Candle GLiNER backends: cpu, metal, or auto; metal requires --features metal, while auto uses Metal when available and otherwise falls back to CPU (with an event) |
MEMORY_CLAIM_ROLLOUT_STAGE |
string enum | shadow |
Claim reconciliation rollout stage: disabled, shadow, relations, or evidence |
MEMORY_CLAIM_CANDIDATE_PAGE_SIZE |
unsigned integer | 256 |
Candidate page size for claim reconciliation |
MEMORY_CLAIM_INLINE_CANDIDATE_LIMIT |
unsigned integer | 1024 |
Inline claim candidate limit |
MEMORY_CLAIM_INLINE_BUDGET_MS |
unsigned integer | 50 |
Inline claim reconciliation budget in milliseconds |
ENTITY_FUZZY_THRESHOLD |
floating-point number | 0.85 |
Entity fuzzy-match threshold |
Advanced provider selection may cause network access or model downloads. Keep these variables unset for the local-first quick start.
Read only by the memory_mcp_http binary built with the streamable-http feature. Set these to deploy the SaaS profile; the Streamable HTTP SaaS profile section above explains the operational behavior and reverse-proxy contract, and the Streamable HTTP SaaS specification is the contract of record.
HTTP boundary
| Variable | Type | Default | Description |
|---|---|---|---|
MEMORY_MCP_HTTP_BIND |
socket address (IP:port) |
0.0.0.0:8080 |
Listen address |
MEMORY_MCP_HTTP_PUBLIC_BASE_URL |
URL | unset | Required. Public base URL used for OIDC redirects and absolute links |
ALLOWED_HOSTS |
comma-separated list | unset | Required for production. Wildcard and unset values are rejected at startup; missing Host returns 403 |
ALLOWED_ORIGINS |
comma-separated list | unset | Required for production. Wildcard values are rejected; missing Origin is allowed only for non-browser MCP clients, present Origin must match |
MEMORY_MCP_HTTP_TRUSTED_PROXY_CIDRS |
comma-separated CIDR list |
unset | Trusted reverse-proxy CIDRs for forwarded Host/Origin; if unset, the values are ignored entirely |
MEMORY_MCP_HTTP_BODY_LIMIT |
bytes | 8388608 (8 MiB) |
Maximum request body size; oversized bodies return 413 |
MEMORY_MCP_HTTP_REQUEST_DEADLINE_SECS |
seconds | 120 |
Ordinary request handler deadline; does not apply to subscriptions/listen |
MEMORY_MCP_HTTP_SHUTDOWN_GRACE_SECS |
seconds | 30 |
Time the server waits for in-flight requests and SSE streams during shutdown |
SurrealDB (control Registry and tenant engine) — required
| Variable | Type | Description |
|---|---|---|
SURREALDB_CONTROL_URL |
URL | Remote ws, wss, http, or https connection URL for the control Registry |
SURREALDB_CONTROL_USERNAME |
string | Registry username; non-empty for remote |
SURREALDB_CONTROL_PASSWORD |
string | Registry password; non-empty for remote |
SURREALDB_CONTROL_NAMESPACE |
string | Registry namespace (separate from tenant namespaces) |
SURREALDB_CONTROL_DB |
string | Registry database name |
SURREALDB_TENANT_URL |
URL | Remote URL for the tenant engine that creates and binds immutable tenant namespaces |
SURREALDB_TENANT_USERNAME |
string | Tenant engine username |
SURREALDB_TENANT_PASSWORD |
string | Tenant engine password |
SURREALDB_TENANT_NAMESPACE |
string | Tenant engine namespace |
SURREALDB_TENANT_DB |
string | Tenant engine database name |
Keyed verifiers and secrets — required, 32-byte hex each (raw secrets are never persisted or logged)
| Variable | Description |
|---|---|
MEMORY_MCP_API_KEY_PEPPER |
Pepper for the keyed HMAC verifier of Account API keys; rotating it invalidates every existing key. Must be ≥ 32 bytes of secret material; the server does not require hex encoding for this field |
MEMORY_MCP_HTTP_IDENTITY_INDEX_KEY |
Blind index key for OIDC subject verifiers; rotating it requires every OIDC identity to relink |
MEMORY_MCP_HTTP_SESSION_KEY |
HMAC key for browser-session cookie verifiers; rotating it invalidates every browser session |
MEMORY_MCP_HTTP_OIDC_STATE_KEY |
AEAD key for OIDC state nonces; rotating it invalidates in-flight login flows |
MEMORY_MCP_HTTP_OIDC_NONCE_KEY |
AEAD key for OIDC ID-token nonces; rotating it invalidates in-flight login flows |
MEMORY_MCP_HTTP_CSRF_KEY |
HMAC key for CSRF tokens; rotating it invalidates every active browser session |
Signup policy, control plane, and OIDC
| Variable | Type | Default | Description |
|---|---|---|---|
MEMORY_MCP_HTTP_SIGNUP_MODE |
enum: invite_only | open |
unset | Required. invite_only rejects self-service sign-up; open requires the seven plan seed variables below |
MEMORY_MCP_HTTP_ENABLE_CONTROL_PLANE |
boolean | false |
Enable OIDC, browser sessions, and /api/v1 endpoints |
MEMORY_MCP_HTTP_ENABLE_CONTROL_PLANE_UI |
boolean | false |
Serve the embedded Dioxus SPA from / (requires control-plane-ui build) |
MEMORY_MCP_HTTP_OIDC_ISSUER |
URL | unset | Required when the control plane is enabled. Exact issuer match is enforced on every login |
MEMORY_MCP_HTTP_OIDC_CLIENT_ID |
string | unset | Required when the control plane is enabled |
MEMORY_MCP_HTTP_OIDC_AUDIENCE |
URL string | unset | Required when the control plane is enabled. Exact audience match is enforced against the ID token's aud claim; supply a single audience identifier (the server does not currently parse a list) |
MEMORY_MCP_HTTP_OIDC_REDIRECT_URI |
URL | unset | Required when the control plane is enabled. Must match the registered redirect URI exactly |
MEMORY_MCP_HTTP_OIDC_ALLOWED_ALG |
enum | RS256 |
JWT algorithm allowlist; accepted values are RS256, ES256, and EdDSA. Tokens signed with any other algorithm are rejected. Mismatched values fail startup with ConfigInvalid |
MEMORY_MCP_HTTP_OPERATOR_IDENTITIES |
comma-separated issuer|hex(subject_verifier) list |
unset | Immutable operator allowlist. Account APIs cannot grant operator status |
Plan seed (required for signup_mode=open) — if any one of these is set, all seven must parse as u64/usize. The values seed Registry plan version 1 only when no plan exists; an existing durable plan is never overwritten.
| Variable | Type | Description |
|---|---|---|
MEMORY_MCP_HTTP_MAX_INGESTED_BYTES |
u64 |
Cumulative ingested source bytes per tenant |
MEMORY_MCP_HTTP_MAX_EPISODE_COUNT |
u64 |
Total episode count per tenant |
MEMORY_MCP_HTTP_INGEST_PER_MINUTE |
u64 |
Token-bucket ingest rate per tenant |
MEMORY_MCP_HTTP_MAX_OPEN_APP_SESSIONS |
usize |
Concurrent App Sessions per tenant |
MEMORY_MCP_HTTP_MAX_ACTIVE_API_KEYS |
usize |
Active API keys per account |
MEMORY_MCP_HTTP_PER_TENANT_REQUEST_CONCURRENCY |
usize |
Concurrent ordinary requests per tenant |
MEMORY_MCP_HTTP_EXTRACTION_CONCURRENCY |
usize |
Concurrent extract operations per tenant |
Runtime pool, subscriptions, tasks, replica identity
| Variable | Type | Default | Description |
|---|---|---|---|
MEMORY_MCP_HTTP_POOL_CAP |
usize |
32 |
Maximum number of tenant runtimes kept warm |
MEMORY_MCP_HTTP_RUNTIME_IDLE_TTL_SECS |
seconds | 900 (15 min) |
Idle eviction window for unpinned runtimes |
MEMORY_MCP_HTTP_RUNTIME_CAPACITY_WAIT_MS |
milliseconds | 2000 |
Maximum time a request waits for runtime capacity before returning 503 |
MEMORY_MCP_HTTP_RUNTIME_ACTIVATION_TIMEOUT_SECS |
seconds | 30 |
Maximum time the activator waits for a tenant runtime to become ready |
MEMORY_MCP_HTTP_GLOBAL_REQUEST_LIMIT |
u32 |
256 |
Global concurrent ordinary-request admission budget |
MEMORY_MCP_HTTP_SUBSCRIPTION_LIMIT |
u32 |
32 |
Global concurrent subscriptions/listen admission budget (separate from ordinary requests) |
MEMORY_MCP_HTTP_MAINTENANCE_PARALLELISM |
usize |
4 |
Scheduler maintenance-job concurrency |
MEMORY_MCP_HTTP_SUBSCRIPTION_QUEUE_CAPACITY |
usize |
64 |
Bounded per-listener event queue; slow consumers are disconnected |
MEMORY_MCP_HTTP_SUBSCRIPTION_AUTH_RECHECK_SECS |
seconds | 30 |
Maximum interval between authorization re-checks for an open stream |
MEMORY_MCP_HTTP_TASK_RETENTION_SECS |
seconds | 604800 (7 days) |
How long completed durable Tasks are kept before the scheduler removes them |
MEMORY_MCP_HTTP_TASK_QUEUE_CAPACITY |
usize |
256 |
Bounded durable Task queue capacity |
MEMORY_MCP_HTTP_TASK_SYNC_MAX_BYTES |
usize |
1048576 (1 MiB) |
Preflight size limit: extract work above this returns a preflight rejection for clients that did not advertise Tasks |
MEMORY_MCP_HTTP_REPLICA_ID |
string | unset (falls back to process PID) | Stable replica identity. Set in multi-replica deployments; the PID fallback is safe only for a single process |
MEMORY_INGESTION_INBOX and any other stdio-only filesystem variable are rejected as a fatal startup error in the HTTP profile.
Additional startup-failure rules (not optional):
ALLOWED_HOSTSandALLOWED_ORIGINSare required and must be non-empty in any HTTP build. Missing or wildcard origins are rejected at startup (ConfigInvalid), and the server does not fall back to permissive defaults.- The control and tenant SurrealDB targets (
SURREALDB_CONTROL_*vsSURREALDB_TENANT_*) must differ in at least one ofurl,namespace, ordatabase. The server rejects identical bindings at startup to prevent the control Registry from writing into a tenant namespace. mem://is rejected at startup in any non-test HTTP build. Production HTTP must use a remotews/wss/http/httpsSurrealDB target or a documented embeddedrocksdb://profile.SURREALDB_FS_WATCH_INBOXis rejected as a fatal startup error if set. Filesystem ingestion is the stdio-only ingestion path; setting it in the HTTP profile means the deployment is misconfigured.
Build with the optional prometheus feature and set
MEMORY_PROMETHEUS_LISTEN_ADDR to expose the Prometheus endpoint. The runtime
exports three generic bounded metric families:
| Metric | Labels | Meaning |
|---|---|---|
memory_operation_calls_total |
operation, outcome |
Logical operation volume; outcome is success or error |
memory_operation_duration_seconds |
operation, outcome |
Operation latency histogram |
memory_operation_results_total |
operation, result |
Counts of bounded domain outputs such as facts, entities, or retrieved items |
Claim reconciliation exports additional memory_claim_* families under the
cardinality rules in ADR-0005.
The evaluation harness does not publish ephemeral Prometheus series: its
versioned JSON artifacts are the source of truth for batch latency, capacity,
retrieval quality, gates, and case outcomes. Individual record identifiers are
never metric labels; use structured logs for per-request diagnosis. See
ADR-0048.
The binary supports a few opt-in Cargo features:
| Feature | Effect |
|---|---|
mimalloc |
Use the mimalloc global allocator instead of the system allocator. This remains an explicit experiment: the fresh macOS matrix reduced physical footprint after GLiNER unload but increased observed RSS to about 2.56 GB, so it is not the server default. Build: cargo build --release --features mimalloc. |
accelerate |
Enable Candle's Apple Accelerate CPU backend. This is an explicit Apple-specific feature, not a portable package default; the current A/B did not pass the no-degradation gate, so do not present it as a production speedup. Build: cargo build --release --features accelerate. |
metal |
Enable Candle's Metal backend for explicit macOS GPU experiments. It is not a production default. Build: cargo build --release --features metal. |
mcp-apps |
Enable the optional interactive MCP app-session surface. It is not required for the eight core tools or the zero-config first-value path. Build: cargo build --release --features mcp-apps. |
control-plane-ui |
Compile the optional Dioxus control-plane SPA. It requires a prebuilt web bundle; see Control-plane UI asset packaging. |
prometheus |
Compile the optional Prometheus recorder/listener. Set MEMORY_PROMETHEUS_LISTEN_ADDR at runtime to expose /metrics. |
The allocator evidence is recorded in docs/performance/MEMORY_PROFILE.md, the CPU-backend result in docs/performance/NER_PERFORMANCE.md, and the policy in ADR-0034. For infrequent local GLiNER extraction, NER_IDLE_UNLOAD_SECS=30 is the measured workload-specific memory recommendation; the runtime compatibility default remains 0.
The control-plane-ui feature embeds the separately built Dioxus 0.7 web bundle
into the memory_mcp binary at compile time. The runtime does not read a
filesystem asset directory, and the build never fetches UI assets from the
network.
Build the UI with the Dioxus CLI matching the crate's 0.7 dependency, then pass an absolute bundle directory to the backend build:
cd crates/control-plane-ui
dx bundle --platform web --release --out-dir "$PWD/../../target/control-plane-ui-dist"
cd ../..
MEMORY_MCP_CONTROL_PLANE_UI_DIST="$PWD/target/control-plane-ui-dist" \
cargo build --release --features control-plane-uiThe bundle must contain a non-empty index.html. All regular files are copied
in deterministic path order into Cargo's OUT_DIR and embedded with
include_bytes!; symlinks, non-UTF-8 paths, and invalid bundle entries are
rejected. If control-plane-ui is enabled without the environment variable or
without a complete bundle, compilation fails with an actionable error instead
of producing a placeholder page. Builds without that feature do not require UI
assets.
Each server process selects exactly one SurrealDB namespace at startup. The
default is main; SURREALDB_NAMESPACE may select one other namespace. All
ordinary MCP, CLI, lifecycle, app, and worker operations use that namespace
implicitly — tool calls do not carry scope, project, or a request-level
namespace.
SURREALDB_NAMESPACE accepts exactly one name. The removed plural variable
SURREALDB_NAMESPACES is a hard configuration error; if your environment still
sets it, choose one name and replace it:
# old and unsupported
SURREALDB_NAMESPACES=kaspersky,org,personal,private-domain
# choose exactly one
SURREALDB_NAMESPACE=kasperskySwitching and restarting accesses another namespace without moving data. One Active Namespace does not isolate personal/corporate/family/project memories internally; operators who need separate authorization domains must run separate process configurations. Namespace transfer/export/import is not automatic.
SURREALDB_DB_NAME=memory
SURREALDB_NAMESPACE=org
SURREALDB_USERNAME=root
SURREALDB_PASSWORD=root
SURREALDB_URL=ws://127.0.0.1:8000/rpc
SURREALDB_EMBEDDED=false
RUST_LOG=info
QUERY_LOGGING_ENABLED=false
QUERY_LOG_RETENTION_DAYS=90
# Lifecycle background jobs (optional)
LIFECYCLE_ENABLED=true
LIFECYCLE_DECAY_INTERVAL_SECS=3600
LIFECYCLE_ARCHIVAL_INTERVAL_SECS=86400
LIFECYCLE_DECAY_THRESHOLD=0.3
LIFECYCLE_ARCHIVAL_AGE_DAYS=90
# LIFECYCLE_DECAY_HALF_LIFE_DAYS=365
# Optional local model configuration
# EMBEDDINGS_ENABLED=true
# EMBEDDINGS_PROVIDER=local-candle
# EMBEDDINGS_MODEL=intfloat/multilingual-e5-small
# EMBEDDINGS_MODEL_DIR=./data/models/intfloat/multilingual-e5-small
# NER_EXTRACTOR=urchade/gliner_multi-v2.1
# NER_CACHE_DIR=./data/models/ner
# GLINER_DEVICE=cpuNER_EXTRACTOR=urchade/gliner_multi-v2.1 is the only model-backed
selector that performs remote acquisition. The classic GLiNER backend
follows these rules so MCP readiness never waits on the network:
- First install with no local cache: the MCP
initializerequest succeeds immediately. Extraction is unavailable in this process; the server returns a structuredmodel_not_readyerror withretryable=false,restart_required=true, andactivation=next_restart. A background refresh task starts after MCP readiness and downloads the 1+ GB checkpoint with cancel-safe staging. - Operator guidance: the structured event
ner.artifact_refresh.candidate_ready(withactivation=next_restart) on stderr/logs means a new revision was staged locally. Restart Memory MCP (and therefore Zed) to activate it. Retryingextractin the same process cannot activate the model — the active extractor and fingerprint are immutable for the process lifetime. - Download-free alternatives:
NER_EXTRACTOR=anno(the default),NER_EXTRACTOR=regex, andNER_EXTRACTOR=anno-onnx(CPU only, manual checkpoint) do not perform network acquisition. Anno and Regex are the recommended choices for offline installs. - Operational failures (inaccessible
NER_CACHE_DIR, missing read permissions) remain explicit startup errors and never silently downgrade extraction.
See ADR-0051 for the state machine, cancellation guarantees, and rejected alternatives.
The server supports three embedding backends, controlled by EMBEDDINGS_PROVIDER:
| Provider | What it is | Default dimension | Requires network? |
|---|---|---|---|
local-candle |
In-process BERT model via Candle (Rust ML) | 384 | Only for first download |
openai-compatible |
External OpenAI-compatible HTTP API | 1536 (configurable) | Yes, every call |
ollama |
External Ollama HTTP API | 1536 (configurable) | Yes, every call |
At startup the server resolves a target embedding identity from the configured provider, model, base URL, and effective dimension.
For remote providers (openai-compatible, ollama) the effective dimension is normally detected with a single short dimension probe request to the provider.
Two startup behaviors keep this from blocking serve:
- If
SURREALDB_EMBEDDING_DIMENSIONis set, the probe is skipped entirely — the override is authoritative at startup, so the server resolves its embedding identity without any network access. A wrong override then surfaces as a dimension-validation error on embed; usereembedas the recovery path. - If no override is set and the provider is unreachable, the probe fails fast (single attempt, bounded by a short probe timeout) and the server degrades to lexical/graph-only retrieval instead of stalling startup.
That identity is persisted per namespace in embedding_state:fact as an active_signature once the namespace is known to be compatible.
In normal serve startup, the Active Namespace is checked before semantic retrieval is enabled.
If it is already marked ready for the same signature and has no missing vectors, semantic retrieval starts normally. If it is marked backfill_pending, startup resumes the recovery worker; a matching ready state with embedding IS NONE facts is also treated as resumable for compatibility with states written before the durable marker existed.
If its state is missing but it is clearly compatible (empty namespace or sampled legacy vectors all match the current dimension), the service bootstraps a ready state automatically.
If it is marked rebuilding, failed, or has embeddings that do not match the configured target, the service degrades to lexical/graph-only retrieval instead of mixing incompatible vectors. When a signature differs but missing vectors exist, the service starts degraded and schedules safe backfill of only those missing vectors; the old persisted signature remains until reembed completes.
That is the safety rail: after a provider switch, normal MCP traffic keeps working, but semantic retrieval is intentionally disabled until embeddings are rebuilt.
To switch, change the environment variables and restart. The server does not silently rewrite old vectors during normal startup.
The runtime now separates two modes:
Normal mode (memory_mcp or memory_mcp serve) — safe startup checks run first. If stored embeddings are incompatible with the configured target, semantic retrieval is disabled and the process logs embedding.rebuild_required.
Maintenance mode (memory_mcp reembed) — a dedicated one-shot command that forces the configured embedding provider on, rewrites every fact embedding, persists progress, and exits when complete.
This keeps the public MCP tool surface unchanged while giving operators a deterministic recovery path after provider changes.
A remote embedding provider is an external dependency, so the server separates startup availability from later semantic recovery:
- Startup performs one bounded dimension preflight. If the remote endpoint is unavailable, the server logs
embedding.preflight_failedandembedding.startup_decision, installs the disabled provider, and continues with lexical/graph-only retrieval. MCP operations do not wait for the provider's runtime retry loop. Remote HTTP errors include the sanitized endpoint, model, and boundedresponse_payload; JSON fields such asapi_key,authorization,token,input, andpromptare redacted. - When
EMBEDDINGS_AUTO_RECOVERYis enabled, a background worker waitsEMBEDDINGS_RECOVERY_INTERVAL_SECS(default60) and probes again. Probe failures use15s → 30s → 60sexponential backoff capped at300s; transport failures and HTTP errors, including404, remain retryable. After three consecutive failures the repetitive probe event is logged at debug level. - If the probe dimension matches the active index and the persisted signature is equal or absent, the worker first persists
embedding_state:fact.status = "backfill_pending", then swaps in the provider, invalidates the context cache, and backfills facts created while degraded. Only after backfill completes does it persiststatus = "ready". Backfill selects onlyembedding IS NONE, usesfact_idorder and batches of100, and never drops the HNSW index. - If the dimension matches but the signature differs, the worker enables the provider for new writes, logs
embedding.reembed_required, and backfills only facts without vectors. It preserves the old persisted signature and never rewrites existing vectors, so a restart remains degraded untilreembed. If the dimension differs, semantic mode stays disabled andembedding.reembed_requiredpoints toreembed. - After compatible recovery and an empty applicable backfill set, the worker exits. Server shutdown cancels and joins it through the existing lifecycle shutdown path. A crash during backfill is restart-safe:
backfill_pendingor the missing-vector count causes the next startup to resume.
The observed startup 404 is not proof of an air gap: it usually means that the configured URL, route, model, or provider API shape is wrong. The recovery worker keeps probing so a transient endpoint failure does not require a restart, but a persistent 404 still requires correcting the endpoint configuration.
Use the maintenance command after changing any embedding target that should become authoritative for stored facts:
EMBEDDINGS_PROVIDEREMBEDDINGS_MODELEMBEDDINGS_BASE_URL- effective embedding dimension (including override/probe changes)
Example:
memory_mcp reembedFrom the workspace during development:
cargo run --quiet --bin memory_mcp -- reembedFlags:
| Flag | Default | Description |
|---|---|---|
--max-failures N |
10% of total (min 10) | Maximum failed facts before aborting. Use 0 for fail-fast behavior. |
--retry-failed |
off | Retry only facts that failed in a previous run. |
What the command does:
- Resolves the configured target signature and dimension.
- Loads or creates a persisted control-plane job record at
embedding_job:fact_reembed. - Marks the Active Namespace as
rebuildinginembedding_state. - Rewrites all fact embeddings in the Active Namespace, including invalidated / historical facts.
- Stores fresh metadata on each fact (
embedding_provider,embedding_model,embedding_dimension,embedding_signature,embedding_updated_at). - Marks the Active Namespace
readyon success, orfailedif the failure quota is exceeded.
The job is restart-safe for the same target signature: if the process stops mid-run, invoking memory_mcp reembed again resumes from the persisted Active-Namespace cursor instead of starting from scratch. Legacy aggregate job records are read only for the currently selected namespace; other entries are left untouched.
Ctrl+C handling: Pressing Ctrl+C interrupts the run gracefully. The current fact finishes, job state is persisted with status interrupted, and the exit code is 130. Resume with memory_mcp reembed.
Continue-on-error: By default, the command continues processing after a fact failure. If the number of failures stays within the quota (10% of total, minimum 10), the run completes with status completed_with_errors. Use --max-failures 0 to restore fail-fast behavior. After a run with errors, use --retry-failed to retry only the failed facts.
See ADR-0018 for the full architectural rationale.
The maintenance flow supports two modes:
TTY mode (interactive terminal): When stderr is a TTY, a live progress bar shows:
Reembedding [org] ██████████░░░░░░░░ 1240/3000 (41%) eta 2m 15s | 38/s ✓1230 ✗10
- Percentage, processed/total, ETA in human-readable format, facts/sec
- Success/failure counters (
✓1230 ✗10) - Namespace label in the bar prefix
- Spinner during service initialization
- Redraw throttled to 10 Hz
After completion, a compact summary is printed to stdout:
✓ Reembed completed (with errors)
Total: 3000 facts
Processed: 3000 (2990 succeeded, 10 failed)
Duration: 135.2s
Speed: 22 facts/sec
10 facts failed. Re-run with --retry-failed to retry only failures.
Non-TTY mode (pipes, CI, scripts): When stderr is not a TTY, the command falls back to structured log events:
reembed.init_completed— service initialized, ready to processreembed.namespace_started/reembed.namespace_completedreembed.index_recreating/reembed.index_recreatedreembed.progress— batch-level progress (every 100 facts)reembed.fact_failed— a fact failed to re-embed (with error reason)reembed.job_interrupted— Ctrl+C receivedreembed.job_completed— final outcome withoutcomefieldreembed.job_failed— quota exceeded or unrecoverable errormain.reembed_completed— compact summary with totals and elapsed time
Job statuses persisted in the control-plane record: running, completed, completed_with_errors, failed, interrupted.
To restore semantic retrieval safely after a provider change:
Change the embedding environment variables.
Run memory_mcp reembed (or cargo run --quiet --bin memory_mcp -- reembed).
Wait for the maintenance run to complete successfully.
Start the normal MCP server again.
Until step 3 completes, the server may intentionally run with semantic retrieval disabled while lexical and graph-based retrieval continue to work.
For external embedding backends (openai-compatible and ollama), the server now treats transient provider issues differently from hard configuration errors.
Bounded retries with backoff are applied automatically for:
- request timeouts / connect failures
- HTTP
429rate limits - retryable upstream statuses such as
408,425,500,502,503, and504
If those retries still do not recover the provider:
write-paths keep the fact write and schedule an in-memory background retry to fill in the missing embedding later;
query-time semantic retrieval falls back to lexical / graph-only results for the current request and schedules a background warm-up of a short-lived query embedding cache for repeated identical queries;
memory_mcp reembed still stops after bounded retries and keeps the maintenance job in a failed state so operators can fix the provider and rerun it explicitly.
Important limitation: the deferred background path is intentionally in-memory only. If the process restarts before a background retry succeeds, those deferred retries are lost and will be attempted again only when a new request hits the same path.
The EMBEDDINGS_SIMILARITY_THRESHOLD (default 0.7) filters semantic search results: only facts with cosine similarity ≥ threshold are returned. After a provider switch, this threshold effectively filters out all old facts because cross-provider similarity scores are meaningless.
If you only use lexical (BM25/FTS) retrieval and graph-expanded context assembly, the provider switch has no impact on those retrieval tiers — they do not use embeddings.
assemble_context remains lexical/BM25-first, but now applies deterministic query-mode routing before ranking results:
- explicit
view_modestill wins; - temporal-history queries such as
timeline of Atlas changes in Q1 2026automatically resolve to timeline ordering whenview_modeis omitted; - named entity anchors can expand into 1-hop graph context (2 hops for explicit connection/path questions) without requiring semantic retrieval.
Persisted query analytics are optional and disabled by default.
When QUERY_LOGGING_ENABLED=true, successful assemble_context calls write a row to the query_log table with:
queryview_moderesolved_view_modequery_flagsretrieval_tiersresult_countlatency_msretrieval_tiercache_hitlogged_at
Old query_log rows are pruned with a best-effort retention pass after successful writes. By default, rows older than 90 days are deleted; override this with QUERY_LOG_RETENTION_DAYS=<days>.
This switch only controls database-backed query analytics. Regular runtime logs still follow RUST_LOG.
memory_mcp emits structured logs across the plan-added functionality using the standard levels below:
info— lifecycle milestones and successful high-level operations such asingest,extract,assemble_context, filesystem-ingestion readiness and per-revision outcomes, and community rebuild passesdebug— feature-path decisions such as document ingest transport detection (file/directory/url/inline), view-mode selection, graph insight assembly, hub/community map building, and successfulquery_logwrites when enabledtrace— fine-grained diagnostics such as cache misses/sets,query_logskips when disabled, retrieval-tier summaries, appendedexperiencefacts, and Active-Namespace community rebuild detailswarn— recoverable issues such as unknownview_modefallback, access-heat tracking failures, query analytics write failures, degraded worker passes, andfs_watch.degradederror— terminal failures such as process-level startup/serve failures
Recommended presets:
RUST_LOG=infofor normal local/server usageRUST_LOG=debugwhen validating new ingest/view-mode/graph behaviorRUST_LOG=tracewhen debugging retrieval tiers, cache behavior, or filesystem-ingestion revision decisions
An .env file already exists in the repository root, so you can keep local values there if your MCP host or shell loads it.
When MEMORY_LOG_FILE is set to a non-empty path, all structured log events are written to that file instead of stderr. This is useful for MCP hosts that do not expose the server's stderr. The file is opened in append mode (no rotation); the parent directory must already exist. If the file cannot be opened, a warning is emitted to stderr and logging continues there.
The public MCP surface is centered on a small set of high-value operations rather than endpoint-by-endpoint plumbing.
| Tool | Purpose |
|---|---|
ingest |
Store an episode with source metadata and timestamps |
extract |
Extract entities, facts, and links from an episode or raw content |
resolve |
Canonicalize an entity name and aliases into a stable entity record |
assemble_context |
Return ranked memory context for a query |
explain |
Expand context items with source citations and multi-source provenance |
invalidate |
Mark a fact as no longer valid as of a given time |
open_app |
Launch an optional MCP app session and return a session-backed resource URI |
app_command |
Execute coarse-grained actions against an open MCP app session |
When the MCP host supports resources, the server also exposes app discovery and session resources such as ui://memory/apps and ui://memory/app/{app}/{session_id} for inspector, diff, ingestion review, lifecycle, and graph views.
The explain() operation returns complete provenance lineage for each fact:
- Direct sources — episodes that directly generated the fact
- Linked sources — episodes connected via shared entities
Returns:
all_sources: Array of provenance sources including:episode_id: Source episode identifierepisode_content: Excerpt from the episodeepisode_t_ref: Episode timestamprelationship: "direct" (created fact) or "linked" (via entity)entity_path: Path from fact to episode via entity (if linked)
This enables full audit trails, understanding of information propagation, and building trust through transparency.
This design lines up with the intent-driven MCP guidance reflected in the docs: fewer tools, clearer semantics, better outcomes.
As of 2026-03-27, memory_mcp implements adaptive memory alignment with SOTA research:
-
Fact-augmented index keys: Entity names, aliases, and temporal markers (month-year, ISO dates) indexed at ingest for enriched BM25 retrieval. FTS matches on both
contentandindex_keys. -
Heat-aware lifecycle: Recently-accessed facts protected from decay/archival via
access_countandlast_accessedfields. Retrieval increments by 1, explain increments by 3 (stronger signal). -
Timeline retrieval:
assemble_contextsupportsview_mode=timelinewith optionalwindow_start/window_endfor chronological queries. Results sorted byt_valid(oldest first). -
LongMemEval-style acceptance tests: Coverage for multi-session reasoning, temporal reasoning, knowledge update, abstention, and direct fact lookup.
See the Architecture Decision Records (ADR-0008 source continuity, ADR-0022 compact responses, ADR-0040 narrow retrieval, ADR-0049 claim evidence fidelity) and the compatibility contract for the current runtime contract.
cargo check
cargo fmt
cargo clippy -- -D warnings
cargo doc --no-depsPerformance measurements live under crates/eval-harness/benches/ and use
Criterion. They are not part of cargo test.
# Pipeline stages (ingest, extraction, claims, retrieval, end-to-end)
cargo bench -p eval-harness --bench pipeline -- --noplot
# NER on CPU
cargo bench -p eval-harness --bench ner_cpu -- --noplot
# NER on Metal (macOS only; feature belongs to memory_mcp)
cargo bench -p eval-harness --features memory_mcp/metal --bench ner_metal -- --noplot
# NER CPU with Candle Accelerate (macOS only; currently experimental)
cargo bench -p eval-harness --features memory_mcp/accelerate --bench ner_cpu -- --noplot
# Contention
cargo bench -p eval-harness --bench contention -- --noplotSee docs/performance/NER_PERFORMANCE.md for raw samples, contention results,
and the Criterion reproduction contract.
The server advertises the official io.modelcontextprotocol/tasks extension.
extract is the only task-capable tool. A client that advertises the extension
calls extract through ordinary tools/call and receives a task handle with
taskId, status, timestamps, TTL, and a suggested polling interval at the
result level. Poll tasks/get until the task is terminal; completed payloads are
embedded in the detailed task’s result field and failed payloads in its error
field. tasks/update is available for input responses and tasks/cancel requests
cooperative cancellation. Task listing and a separate terminal-result request are
not part of this extension contract.
Clients that do not advertise io.modelcontextprotocol/tasks continue to receive
synchronous extract results. rmcp’s TaskManager supplies the default five-minute
TTL, polling metadata, lifecycle, and retention behavior.
CPU is the production default. Metal remains an experimental, explicit opt-in until its candidate parity, latency, contention, and memory gates are recorded for the deployment hardware.
# Run with Metal GPU (requires --features metal)
GLINER_DEVICE=metal cargo run --release --features metal -- serve
# Auto tries Metal and falls back to CPU; do not make it a deployment default before gating
GLINER_DEVICE=auto cargo run --release --features metal -- servecrates/memory-mcp/src/main.rs— main MCP server binary (memory_mcp); stdio profile, CLI, and lifecycle hookscrates/memory-mcp/src/bin/memory_mcp_http.rs— Streamable HTTP SaaS binary (memory_mcp_http); built when thestreamable-httpfeature is enabledcrates/eval-harness/src/main.rs— evaluation harness binary (memory-eval); never linked into the production binarycrates/control-plane-ui/src/main.rs— Dioxus web SPA build target (control-plane-ui); built with the Dioxus CLI and embedded by the backend when thecontrol-plane-uifeature is enabled
MCP input/output schemas are exposed by the server itself through the protocol's
tool metadata and remain regression-covered by the schema tests under
crates/memory-mcp/src/mcp/.
Run the production crate's test suite:
cargo testRun every workspace member, including the private evaluation harness:
cargo test --workspaceUseful narrower runs:
cargo test --test service_integration
cargo test --test service_acceptance
cargo test --test tools_e2eCoverage output is stored under coverage/ when generated with Tarpaulin.
.
├── AGENTS.md
├── Cargo.toml # workspace root
├── Makefile # thin eval profile adapters
├── crates/
│ ├── memory-mcp/ # production package (std and HTTP binaries)
│ │ ├── migrations/
│ │ ├── src/ # library, two binary entry points, MCP/HTTP, control plane, and domain services
│ │ └── tests/ # production integration and release-gate tests
│ ├── control-plane-ui/ # Dioxus 0.7 web SPA (control-plane-ui feature)
│ │ └── src/ # router, API client, login/keys/delete/status pages
│ └── eval-harness/ # private evaluation package
│ ├── benches/ # Criterion benchmark families
│ ├── src/ # domain, artifact, metrics, gate, suites, runner, CLI
│ └── tests/ # harness integration tests and fixtures
├── evals/
│ ├── corpora/ # immutable corpus manifests + NER corpora
│ ├── longmemeval_v2/ # prepared LoCoMo/LongMemEval corpora
│ ├── performance/ # pinned-runner config
│ ├── profiles/ # pr.json, release.json, nightly.json, ner_quality.json
│ ├── results/ # recorded comparison results (e.g. NER)
│ └── schema/ # eval-artifact-v1.json
├── hooks/ # lifecycle capture scripts (stop, precompact, profile)
├── scripts/ # time-to-value harness, parity checks
└── docs/
The eval-harness crate (memory-eval binary) provides a profile-driven,
truthful evaluation system. It is never linked into the production binary.
# PR profile (target 10 min)
make eval-pr
# Release profile (target 20 min)
make eval-release
# Nightly profile (full end-to-end)
make eval-nightly
# Corpus preparation (one-time, requires network)
cargo run -p eval-harness --bin memory-eval -- prepare-corpus \
--manifest evals/corpora/longmemeval.json \
--output-root data/corporaProfiles, modes, outcome semantics, artifact schema, and baseline governance
are documented in the design spec at
docs/superpowers/specs/2026-07-28-truthful-evaluation-system-design.md
and the supporting ADRs under docs/adr/.
docs/superpowers/specs/2026-08-27-streamable-http-saas.md— Streamable HTTP SaaS design specificationdocs/superpowers/specs/2026-07-28-truthful-evaluation-system-design.md— evaluation architecture and designdocs/superpowers/specs/2026-07-30-token-efficient-responses-design.md— compact tool responses designdocs/adr/— Architecture Decision Records (52 ADRs, including ADR-0038 one Active Namespace and ADR-0052 Streamable HTTP SaaS profile)docs/compatibility/one-active-namespace-identities.md— scope/namespace compatibility contractdocs/operations/— operator runbooks (protocol conformance, credential rotation, known limitations, SurrealDB restore drill)docs/performance/— memory profile and NER performance measurementsdocs/evals/— evaluation results, benchmark reports, claim reconciliation baselines, and procedural memory evidencedocs/BACKLOG.md— open engineering backloghooks/README.md— lifecycle hooks contract and editor-by-editor configuration
This repository follows the conventions in AGENTS.md.
In particular:
- keep public APIs stable unless a change is explicitly requested
- avoid introducing dependencies without approval
- prefer typed errors and deterministic behavior
- run formatting, clippy, and tests before considering work done
Every memory tool can be invoked directly from the command line. The CLI shares the same implementation as the MCP protocol — zero code duplication.
| Command | Description |
|---|---|
serve (default) |
Run the stdio MCP server. Set MEMORY_INGESTION_INBOX to enable filesystem ingestion (fs-watch feature) |
reembed |
Rebuild all fact embeddings after a provider switch. Flags: --max-failures N, --retry-failed |
lifecycle |
Inspect or run lifecycle maintenance: dashboard, archive-candidates, restore-archived, recompute-decay, rebuild-communities |
ingest |
Store raw source material as an episode |
extract |
Extract entities, facts, and relationships |
resolve |
Resolve entity aliases to a canonical entity id |
invalidate |
Invalidate a fact while preserving history |
explain |
Get citation-ready source snippets |
assemble-context |
Assemble ranked, relevant context for a query |
init [--target TARGET] |
Print deterministic, output-only host setup for vscode, claude-desktop, codex, zed, or env |
init is the one authorized output-only onboarding exception to the ordinary
CLI surface. It does not build a service, touch storage, edit files, or change
environment variables.
# Ingest an episode
memory_mcp ingest \
--source-type email \
--source-id msg-001 \
--content "I will finish the API by Friday." \
--t-ref 2026-06-30T10:00:00Z
# Extract entities and facts
memory_mcp extract --episode-id episode:abc123
# Extract from inline content
memory_mcp extract \
--content "Alice works at Acme Corp." \
--source-type ad-hoc \
--t-ref 2026-06-30T10:00:00Z
# Resolve an entity
memory_mcp resolve \
--entity-type person \
--canonical-name "Alice Smith" \
--aliases Alice --aliases "A. Smith"
# Query assembled context
memory_mcp assemble-context \
--query "What did Alice promise?" \
--budget 10
# Inspect lifecycle state
memory_mcp lifecycle dashboard
# Run a dry-run archival selection without changing storage
memory_mcp lifecycle archive-candidates episode:old-1 \
--dry-run
# Recompute confidence decay (requires --confirmed to mutate)
memory_mcp lifecycle recompute-decay --confirmed
# Invalidate a fact
memory_mcp invalidate \
--fact-id fact:xyz \
--reason "Decision reversed" \
--t-invalid 2026-06-30T00:00:00Z
# Get provenance citations
memory_mcp explain \
--context-items '[{"content":"API delivery","source_episode":"episode:abc"}]'Memory-operation CLI subcommands print the ToolResponse<T> as pretty JSON to stdout. The lifecycle command prints an operation/result JSON envelope, and the output-only init command prints its documented result object to stdout:
{
"status": "success",
"result": "episode:abc123",
"guidance": "Call extract next to derive entities and facts.",
"has_more": false,
"total_count": 1,
"next_offset": null
}Structured log events go to stderr (controlled by RUST_LOG), or to the
file named by MEMORY_LOG_FILE when that variable is set. Successful
CLI results, including memory_mcp init, go to stdout as JSON. Configuration
failures and other error responses go to stderr as JSON:
{
"error": "Invalid `t_ref` value: bad-date. ...",
"kind": "Validation",
"exit_code": 2
}| Code | Meaning |
|---|---|
| 0 | Success |
| 1 | Internal / storage / config error |
| 2 | Validation error or not found |
This project is licensed under the MIT license. See LICENSE for details.
memory_mcp supports agent-host lifecycle integration through an internal
control plane that does not add new public tools. The eight-tool MCP surface
remains exactly eight tools. The ordinary CLI surface has one separate,
output-only onboarding exception: memory_mcp init.
- Architecture: A versioned host lifecycle bridge invokes internal
LifecycleRecallandLifecycleCapturecapabilities, which reuse the existingassemble_contextand inlineextractpaths. - Trust: Derived from the invocation channel, never from public arguments. External content cannot become privileged instruction, preference, policy, retraction, or procedure.
- Growth control: Ignored and duplicate events create zero durable rows. Accepted content is stored once. Quotas prevent unbounded ingestion.
- Procedural memory: Separately gated and projected through the existing
FactType::Experienceseam. Currently shadow-only.
flowchart LR
Session["Agent session"] --> Recall["RECALL\nassemble_context"]
Recall --> Work["Agent work\nreason, edit, call APIs"]
Work --> Outcome{"Significant outcome?"}
Outcome -->|"no"| Work
Outcome -->|"yes"| Capture["CAPTURE\ningest + extract"]
Capture --> Store[("Bi-temporal memory")]
Store --> Recall
Stop["memory_stop_hook.sh"] --> Capture
Precompact["memory_precompact_hook.sh"] --> Capture
The lifecycle integration deliberately reuses the normal retrieval and extraction paths. Hooks do not create a second memory implementation, and external content is treated as data rather than privileged instruction.
See:
- ADR 0016
- Hook scripts contract — transport, environment variables, editor-by-editor configuration, and hidden lifecycle CLI subcommands
- Evaluation Results
- Procedural Memory