Skip to content

Latest commit

 

History

198 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Engram

Trustable memory infrastructure for AI agents, assistants, and teams.

Engram is a standalone memory service that gives AI systems structured, durable, and trustable memory. It works for a single long-running assistant, and it is designed for the harder case: multiple agents sharing memory without corrupting, overwriting, or blindly trusting each other.

It's not a flat key-value memory store. Engram is a memory system of record with taxonomy, relationships, temporal validity, review states, provenance, conflict detection, and explainable recall.

An engram is the physical trace a memory leaves in brain tissue — the literal substrate of stored memory.

Why Engram?

Most AI memory layers answer the first-order question:

Can we store and recall facts?

Engram is built around the harder and more important question:

Can an AI system trust what it remembers, know where it came from, know whether it is still true, and safely act on it?

That matters for a single assistant remembering a user's preferences across sessions. It matters even more for coding agents, research agents, support agents, operations agents, and multi-agent teams that share institutional knowledge.

Engram is designed as trustable memory infrastructure, not just storage.

A lower-authority source can never silently replace a higher-authority memory.

Feature Flat memory stores Engram
Memory model Flat facts Structured taxonomy with wings and rooms
Trust model Minimal or none Review states: proposed → active → disputed → resolved
Relationships Usually none Knowledge graph with temporal validity
Single-agent use Basic recall Durable assistant memory with provenance, confidence, and lifecycle
Multi-agent use Per-agent silos Workspaces, visibility levels, and shared recall
Conflict handling Dedup only Write-time contradiction detection and resolution
Provenance Minimal Source trust, extraction model, subject, verification tracking
Audit trail Overwrites Append-first content plus audited metadata events
Classification Basic extraction LLM-backed and rule-based classification, tenant-configurable, no-LLM fallback
Recall quality Similarity only Scored ranking with "why recalled" explanations
Self-hostable Varies Docker Compose, one command

Who Engram Is For

Engram can be used anywhere an AI system needs durable, inspectable memory.

Single assistant / single user

Use Engram when one assistant needs to remember things safely across sessions:

  • user preferences
  • project context
  • standing instructions
  • important decisions
  • recurring workflows
  • constraints and invariants
  • long-term personal or professional context

A single-agent setup still benefits from Engram's trust model: source authority, confidence, review state, provenance, staleness, and explainable recall.

Coding agents

Use Engram to give coding agents persistent project memory:

  • architectural decisions
  • repo conventions
  • completed backlog items
  • known gotchas
  • deployment notes
  • test expectations
  • prior review decisions
  • "do not repeat this mistake" context

Multi-agent teams

Use Engram when multiple agents need to share memory without creating chaos:

  • workspace-level shared knowledge
  • private agent notes
  • tenant-wide organizational memory
  • conflict detection between competing claims
  • authority-aware supersession
  • memory review and promotion
  • audited updates instead of silent overwrites

This is where Engram's design is strongest: single-agent memory is the on-ramp; multi-agent institutional memory is the moat.

Quickstart

git clone https://github.com/Zutfen-LLC/engram.git
cd engram
cp .env.example .env  # set your passwords
docker compose up -d

This starts Postgres 16 with pgvector, the Engram API service, and a background worker that processes the job queue (embeddings, classification, conflict detection, recall telemetry). The schema migrates automatically on first boot.

Verify that the service is running:

curl http://localhost:8000/health
curl http://localhost:8000/ready

Create the first API key (auth is off by default for local dev; enable it for production). The printed key uses the eng_<key_id>_<secret> format and is shown only once:

docker compose exec engram-service engram bootstrap-key

For the full walkthrough — auth enablement, backup/restore, upgrades, embeddings, and troubleshooting — see docs/deployment.md. External control planes can use the optional, single-use read broker documented in docs/ops/service-delegation.md; its delegated credentials are never browser tokens or ordinary API keys.

See docs/design.md for the full architecture.

Local Python development setup

If you want the repo's .venv to be able to import the sibling SDK plus both adapters directly (so python -m engram_mcp and import engram_hooks work without setting PYTHONPATH), run:

bash scripts/setup-python-dev.sh
# or: make setup-python-dev

This bootstraps ./.venv, then installs these editable local packages into it:

  • sdk/engram-client
  • adapters/mcp-server
  • adapters/engram-hooks

After that, these commands work from the repo checkout:

.venv/bin/python -m engram_mcp
.venv/bin/engram-mcp
.venv/bin/python -c 'import engram_hooks'

Key Concepts

Memory Lifecycle

Every memory moves through a review pipeline:

written → proposed → active → disputed → resolved → superseded/archived
  • Proposed memories do not enter startup recall until reviewed or promoted.
  • Active memories are trusted enough for normal recall.
  • Disputed memories remain available with warnings where appropriate.
  • Archived and superseded memories are preserved for audit but excluded from default recall.

Auto-promotion has independent legacy-confidence and server-attested retention-evidence lanes. The evidence lane uses min(0.85, 0.20 * source_confidence_prior + 0.80 * retention_confidence) for governed kinds only; it changes only proposed → active, never provenance, authority, confidence, or human-verification fields. The promotion decision itself is a deterministic promotion assessment over the currently available evidence: recalls and useful feedback do not accumulate promotion evidence, and Path B (usage-validated quorum) is deferred and unimplemented — useful feedback changes importance (and may reset a startup-recall counter) only.

Per-item readiness is inspectable without re-running providers or the promotion-time conflict recheck via GET /v1/review/promotion-readiness/{item_id} (review/admin scope): it reports the bound receipt and its versions/model/provider, taxonomy and retention confidences, the computed evidence score and threshold, the minimum retention confidence required under the current formula and source prior (unreachable when a cap makes the threshold impossible), the selected lane or none, canonical blocker codes, cooling/eligibility clocks, the current promotion/classification job state, and whether the item can ever auto-promote under current policy without new evidence, configuration, or review. engram doctor additionally reports bounded, content-free readiness aggregates — counts by source type, kind, blocker code, evidence state (none / bound-qualified / bound-below-threshold / malformed/stale), job state (scheduled / missing / dead / overdue), terminal-vs-time-dependent status, and age buckets — and warns when the bounded startup-promotion window is dominated by terminal blockers (blocked backlog the lazy pass now rotates past rather than starves on).

POST /v1/recall with mode=startup retains its bounded, tenant-scoped lazy promotion pass only while ENGRAM_STARTUP_PROMOTION_MUTATION_ENABLED=true (the default, and the explicit rollback switch). It is capped at settings.startup_promotion_limit (default 20) and scans fairly through its persisted cursor. With the switch false, startup is lifecycle-read-only and sees only promotion state already committed by canonical worker/reconciliation processing; configuration fails closed unless both canonical rollout flags are enabled. For full sweeps of a large proposed backlog, wire the CLI to cron/systemd, or call the admin endpoint on demand:

# All tenants. Runs as the owner role (bypasses RLS) via ENGRAM_OWNER_DATABASE_URL:
engram promote-proposed

# Single tenant, capped at 1000 candidates:
engram promote-proposed --tenant <tenant-id> --limit 1000

# Exact evaluation with no writes, audit events, or actor creation:
engram promote-proposed --dry-run
POST /v1/admin/promote

The compatibility startup pass, CLI, and admin endpoint share one promotion service function and the same gates. Once the startup compatibility switch is off, canonical promotion.evaluate owns lifecycle mutation instead.

Promotion returns per-reason counts:

scanned
promoted
skipped_confidence
skipped_age
skipped_conflict
skipped_dispute            # blocked by another principal's dispute/negative feedback
skipped_conflict_recheck   # blocked by a promotion-time conflict recheck
skipped_disabled
skipped_kind_policy
skipped_evidence_disabled
skipped_no_retention_evidence
skipped_missing_source_prior
skipped_retention_disposition
skipped_taxonomy_confidence
skipped_evidence_score
skipped_evidence_version
skipped_evidence_inconsistent
skipped_review_policy

Canonical promotion.evaluate job (issue #155, ENG-PROMOTION-003B2/B3 — partial). Alongside the legacy targeted promotion.path_a job, Engram now has one canonical, item-scoped, current-state promotion-evaluation job contract: promotion.evaluate (engram.promotion.enqueue_promotion_evaluation, engram.promotion.evaluate_promotion_item_current_state, engram.worker.handle_promotion_evaluate). Its payload (promotion-evaluate-v1; promotion-evaluate-v2 adds exactly one optional field — see below) carries a stable memory_item_id, a closed trigger_type/trigger_id audit- provenance pair (item_created, classification_bound, classification_reassessed, feedback, conflict_changed, review_changed, provenance_changed, kind_changed, policy_changed, provider_recovery, reconcile, manual), and a requested_policy_version that is descriptive only. Each version's envelope is exact and closed: parse_promotion_evaluate_payload() rejects any field outside its version's set — v1 is exactly {contract_version, memory_item_id, trigger_type, trigger_id, requested_policy_version, ingest_id, correlation_id, dedupe_key}; v2 adds only execution_context_id, the durable execution-authority reference for non-ingest producers (mutually exclusive with ingest_id; no metadata/extra bag) — and independently recomputes dedupe_key from the parsed (memory_item_id, trigger_type, trigger_id) identity — a stored key that does not match the canonical one fails closed exactly like an unsupported contract version, rather than being trusted from the payload or the queue's unique-index behavior alone. enqueue_promotion_evaluation() builds payloads through the same build_promotion_evaluate_payload() rules the worker enforces at parse time, so enqueue-time construction and execution-time validation cannot drift apart; callers cannot supply their own dedupe_key. Unlike promotion.path_a's classification-run-targeted binding, the handler always evaluates whatever state is authoritative in the database at execution time — a trigger enqueued for a since-superseded observation is audit provenance only, never a filter and never forced back into the decision. Both job types run through the identical shared evaluator and mutation machinery (no second policy implementation), so a canonical-vs-legacy race for the same item yields at most one lifecycle transition and one audit event; the loser is an idempotent no-op. State-changing audit events additionally carry evaluation_id, job_id, job_contract_version, trigger_type, trigger_id, and requested_policy_version when triggered through promotion.evaluate — the legacy promotion.path_a audit-event JSON shape is unchanged.

This slice is intentionally partial, in two stages. The canonical job is behind ENGRAM_PROMOTION_EVALUATE_JOBS_ENABLED (default false). With the flag set, these producers now enqueue canonical evaluations (engram.promotion.maybe_enqueue_promotion_evaluation gates every one of them on the flag, on the item still being a live proposal, and on the tenant's auto_promote_enabled):

  • classification_boundclassification.refine's delayed evidence-promotion schedule (the original producer); can newly admit at the evidence cooling boundary.
  • item_created/v1/remember for every proposed item, scheduled at the exact legacy cooling boundary (created_at + auto_promote_min_age_hours); can newly admit via the legacy lane. This is the reassessment path for explicit-kind writes and below-threshold receipts, which previously had no future evaluation at all.
  • feedback — the /v1/feedback route after a committed verdict transition (recorded/updated; unchanged never enqueues). Can block (a new external noise verdict) or newly admit (a replacement lifting an existing noise block); a first-time useful verdict is consumed by no current gate and so merely refreshes diagnostics (enqueued anyway: cheap, flag-gated, forward-compatible with future usefulness lanes). Feedback effects (importance) remain promotion-blind — engram/feedback.py contains no promotion coupling by design; only the route-level orchestration schedules reevaluation.
  • conflict_changed/v1/items/{id}/resolve-conflict after a committed resolution on a live proposal; can newly admit (clears the conflict block). Conflict creation is deliberately not a producer: creation can only block, and any later evaluation re-derives the block from current state (disputes raised through the review route end the item's proposal, so they need no trigger either).
  • review_changed/v1/items/{id}/verify on a live proposal; refreshes diagnostics only (verification feeds no current promotion gate, and no committed review transition can land an item back on proposed — that is why the review-status route itself has no producer).
  • manualPOST /v1/admin/items/{item_id}/evaluate enqueues one immediate evaluation under admin authority; the item must be a live proposal (409 otherwise). Repeats enqueue independently; the evaluator stays idempotent. The admin scope is not a bypass of a bound memory profile: the target is resolved through the caller's existing-item write eligibility (out-of-boundary is the mutation API's non-disclosing 404), and a profile-bound caller's job carries a durable job_execution_contexts reference (promotion-evaluate-v2) so the worker reconstructs and applies the same pinned profile boundary at execution time — a missing/corrupt/unreconstructable authority reference fails closed through the ordinary retry/dead-letter path, never falling back to broader tenant-level authority. Unprofiled callers keep the v1 payload and the established compatibility behavior.

Still unwired in this slice: classification_reassessed (awaiting #157's versioned reassessment) and kind_changed/provenance_changed (no kind/source-provenance correction path exists in the API today). policy_changed/provider_recovery/reconcile are now wired through the bounded reconciliation backstop (see below). engram doctor and the per-item readiness endpoint recognize both job types. Startup recall's bounded rotating promotion pass (described above) is unchanged by this slice, still mutates directly rather than going through either job contract, and is not yet read-only. Durable, persisted promotion-assessment state remains future work (issue #159) — evaluation_id exists only within an audit event's reason JSON when a mutation or conflict block actually occurs, not as its own row. See the code comments in engram/promotion.py for the full contract, dedupe-key construction, and per-trigger decision-effect matrix.

Bounded promotion reconciliation backstop (ENG-PROMOTION-003B4)

Targeted triggers cover events the service observes itself; the backstop makes promotion reevaluation self-healing and independent of agent startup for the events it can miss — a lost targeted job, a dead cooling-boundary job, a policy change, a provider that recovered after classification failures.

targeted event -> promotion.evaluate
                     ^
                     |
bounded promotion reconciliation backstop

The backstop is orchestration, not a second promotion implementation (engram/promotion_reconciliation.py). It discovers live proposals (review_status='proposed' AND valid_to IS NULL AND superseded_by IS NULL, kind-policy-eligible) needing repair and enqueues canonical work; every lifecycle decision still flows through the shared evaluator behind promotion.evaluate. Its own job contract, promotion.reconcile (promotion-reconcile-v1), is a versioned, exact/closed, identifier-only envelope — {contract_version, reason, trigger_id, dedupe_key} with a closed reason vocabulary (backstop, policy_change, provider_recovery, operator_request) — that fails closed on unknown fields/reasons/versions through ordinary retry/dead-letter. No thresholds, scores, decision state, credentials, or content ever enter the payload, and there is no metadata bag.

Topology — no separate scheduler daemon; all chains use the existing PostgreSQL job queue. Each chain executes bounded passes: the periodic backstop has one active per-tenant chain, while finite requests retain independent durable progress and may safely overlap without sharing cursor coverage:

  • Periodic backstop chain (reason=backstop): each pass self-schedules the next at ENGRAM_PROMOTION_RECONCILIATION_INTERVAL_SECONDS (default 3600). The worker loop bootstraps/heals at most ENGRAM_PROMOTION_RECONCILIATION_TENANT_BATCH_LIMIT tenants per ensure pass (default 100), using an owner-only, content-free, persisted (created_at, id) cursor. It wraps fairly and resumes after restart, so no worker-loop iteration enumerates every tenant.
  • Request chains (policy_change / provider_recovery / operator_request): each finite request has its own durable (tenant, reason, trigger_id) cursor (coverage from the head), so overlapping reasons cannot advance one another's coverage. Explicit request ids remain idempotent after completion; a failed identity is terminal and requires a fresh request id. Each chain runs immediately-due continuation passes until its rotation reaches the tail, then stops — bounded per pass with durable continuation.

Candidate discovery is a keyset rotation over (created_at, id). Periodic progress is persisted in promotion_reconcile_state; finite request progress is persisted independently in promotion_reconcile_chains (migration 034). The periodic cursor is deliberately separate from #164's startup-rotation cursor and carries the tenant's kind_policy_revision plus content-free last-pass diagnostics. “Terminal under current policy” means the item was evaluated under a known reconciliation cursor epoch and is excluded from ordinary periodic selection until that epoch or relevant item/evidence state changes. The lightweight promotion_reconcile_terminal row stores only (tenant, item, epoch, observed_at)—no blocker, threshold, score, or decision. Relevant item, classification, feedback, and audit events invalidate it; policy, operator, and provider-recovery requests advance the epoch. Stable terminal rows are therefore assessed once, not once per periodic rotation, while later actionable rows continue through the bounded keyset window.

Each item pass inspects at most ENGRAM_PROMOTION_RECONCILIATION_PASS_LIMIT rows (default 20). The keyset scan is served directly by (tenant_id, created_at, id) partial indexes (idx_memitems_proposed_rotation, or the existing idx_memitems_backfill) — an index bound, not a LIMIT over a sort of the backlog.

What one pass repairs, per item, using the shared evaluator's own assessment:

  • Time-dependent (a lane is trust-qualified, cooling): if a healthy targeted job exists, leave it alone; otherwise enqueue a canonical promotion.evaluate (trigger_type='reconcile') with run_after equal to the exact authoritative eligibility boundary — never earlier. “Healthy” means pending/running and covering the current obligation, not merely having a live status. A cooling canonical job must be due at the exact evaluator-produced boundary; an already-due obligation accepts a due/overdue canonical job but not a future one. Legacy promotion.path_a must meet the same due-time rule and name the currently bound classification_run_id. Stale or incompatible legacy work remains untouched while canonical repair is enqueued. Indexed per-window EXISTS probes replace the former global 1,000-row history cap, so unrelated history cannot hide a healthy job.
  • Eligible now: same repair, due immediately.
  • Terminal under current policy (blocked without new evidence, review, or policy change): no repair — reconsideration arrives through that evidence's own producer, a policy-change chain (cursor reset), or provider recovery.
  • Provider recovery (reason=provider_recovery only): live proposals with no bound classification receipt and no healthy classification.refine job are repaired only when the immutable initial classification event proves the remember path selected source=auto_classified. Explicit-kind items (source=explicit_kind) and legacy rows with unknown intent are excluded. This restores previously intended async classification; it does not create new classification intent. Recovery is never a provider call inline, never a re-classification of already-bound evidence, and suppressed (recorded in diagnostics) while ENGRAM_CLASSIFICATION_PROVIDER=none. Request it only after the provider is available again; successful recovered classification follows the normal binding path and schedules promotion evaluation through the ordinary classification_bound producer. This is recovery of current failed classification work, not #157's future model/prompt/version reassessment.

What reconciliation deliberately does not repair: #156 structured extraction, #157 reassessment, #159 durable assessment history; bound-but-fallback classification receipts (that is reassessment, not recovery); direct-SQL tenant_config changes, which the service cannot observe.

Policy changes:

  • Memory-kind policy: a committed PATCH to /admin/memory-kinds/{name} that materially changes enabled or auto_promote_from_inferred (the two fields existing proposals' eligibility depends on) bumps the tenant's kind_policy_revision and schedules one bounded policy_change chain — never synchronous item fan-out. Item evaluations reached because of that revision carry trigger_type='policy_changed' with the revision in their trigger identity. Non-admission changes (description, display name, requires_review — which governs only new writes' initial review status) schedule nothing.
  • Tenant promotion configuration has no product mutation surface today (only direct SQL). After changing tenant_config promotion columns, explicitly request reconciliation — the honest path for unobservable changes: POST /v1/admin/promotion/reconcile (admin scope, tenant-scoped, content-free response, idempotent per (reason, request_id)) or engram reconcile-promotion [--tenant <id>] [--reason operator_request|provider_recovery] [--request-id <stable-id>]. A tenant-scoped request id is durable: its replay reports active while the chain runs or completed after it finishes without enqueuing duplicate work. A failed request id is terminal and must be retried with a fresh --request-id (the CLI exits 1 to require operator action). Without --tenant, the command advances one persisted tenant page (default 100) and prints the stable --request-id to repeat until complete=true; it never synchronously materializes every tenant.

Execution authority and safety: reconciliation runs under the worker's routed tenant app-role RLS context; its item evaluations are v1-unprofiled promotion.evaluate jobs under internal worker provenance (no fabricated CandidateIngest, no job_execution_contexts abuse, manual/profile-bound evaluation from #167 unchanged). Repairs, cursor advance, and chain continuation commit in one transaction — a crash before commit replays from the unchanged cursor, a crash after leaves durable work; repair identity is stable per observation so replays dedupe. Concurrent reconciliation/targeted/legacy/startup paths converge on one mutation and one authoritative review_change event (the shared evaluator's guarded update already guarantees this).

Rollout: ENGRAM_PROMOTION_RECONCILIATION_ENABLED (default false) is independent of ENGRAM_PROMOTION_EVALUATE_JOBS_ENABLED. Flag off: no backstop work is created, requests are fail-safe no-ops, and startup recall plus both targeted job types behave exactly as before. Reconciliation on with evaluate off: passes run and record (suppressed) discovered work rather than substituting any broader/legacy mutation mechanism — surfaced by engram doctor's promotion.reconciliation check (v1.2), which reports the rollout flags, cursor/epoch/due state, last-pass counts (evaluations enqueued, dead/missing found, provider-recovery work scheduled, terminal/healthy skipped, suppressed), and pending/dead reconciliation chains — content-free.

Thresholds come from tenant_config:

auto_promote_enabled
auto_promote_evidence_enabled       # existing tenants migrate as false
auto_promote_evidence_threshold     # default 0.70
auto_promote_confidence_threshold
auto_promote_min_age_hours

Trust Model

Trust is not binary. Every memory carries:

  • source_trust — trust in where the memory came from
  • source_confidence_prior — immutable write-time confidence selected from source policy; production classification never blends it afterwards
  • taxonomy_confidence — classifier confidence that kind/wing/room placement is correct
  • retention_confidence — the classifier's estimate that the content is durable and useful enough to keep; it is not epistemic or factual confidence (the retention receipt does not establish that the proposition is correct)
  • authority — immutable governance rank derived from provenance (inferred 10, untrusted agent 20, trusted agent 30, trusted import 40, explicit user 50)

Source trust remains tenant-configurable and affects recall ranking. Authority is fixed by code and controls whether one memory may supersede another; tuning recall scores cannot invert that hierarchy.

  • memory_confidence — legacy promotion-lane input; initialized from the source-policy prior at write time and left unchanged by later classification
  • extraction_confidence — confidence in the extraction process
  • human_verified — whether a human has confirmed it
  • authority level — explicit_user > trusted_import > trusted_agent > inferred

Authority hierarchy governs supersession. A lower-authority source can never silently replace a higher-authority memory.

Supersession is atomic: expiring the original and inserting the replacement happen in one transaction (the original is expired before the replacement is inserted, so the dedup unique index can never see both as active), and a failure between the two rolls back the original's expiration. The original points forward to its replacement (superseded_by), the replacement records what it replaced (item_events), and supersede behavior is covered by Postgres-backed tests that enforce the real unique index and RLS policies.

memory_confidence starts from source-type defaults (see docs/design.md §4). Taxonomy classification does not change it. New writes also preserve that selected default in source_confidence_prior; existing rows remain NULL rather than being historically reinterpreted.

Recall

Startup recall returns a deterministic, bounded working set of active memories, scored by:

score = importance × 0.30
      + source_trust × 0.25
      + memory_confidence × 0.20
      + recency × 0.15
      + human_verified × 0.10

The recency component is the larger of two decay signals: recall recency (decay from last_recalled_at, subject to the anti-feedback penalty) and freshness (decay from valid_from/created_at, half-weighted). This gives a fresh, never-recalled memory a modest recency contribution without letting freshness dominate trust/importance.

Pinned items bypass scoring and are included first, up to a ceiling.

Recall is bounded by default: omitted byte_budget / item_budget fall back to the configured defaults (recall_byte_budget, recall_item_budget) rather than leaving recall unbounded. Explicit caller budgets always override the defaults.

Every recalled item includes a reasons array explaining why it was included.

Anti-feedback-loop guardrails prevent the same memories from permanently dominating recall without useful feedback.

Feedback is canonical per principal and item: identical retries are unchanged, while changing useful/noise appends a replacement and preserves the prior row as history. Accepted first verdicts and changes are limited per principal per UTC day by tenant_config.feedback_daily_limit (default 500); exhaustion returns 429 and Retry-After. Optional recall_log_id attribution is accepted only when the caller owns that recall and the item was actually surfaced.

Semantic recall & search ranking

Semantic recall (mode=semantic) and semantic/hybrid search rank results by a deterministic trust-weighted score (scoring_version = semantic-v2), not pure cosine distance:

similarity   = clamp(1.0 − cosine_distance, 0.0, 1.0)
trust_score  = 0.30·source_trust + 0.30·memory_confidence + 0.25·importance
             + 0.10·human_verified + 0.05·review_status_factor   (clamped to [0.05, 1.0])
semantic_score = similarity × trust_score

Unresolved conflicts multiply trust_score by 0.75; proposed items multiply by 0.85. So a slightly-closer low-trust or unreviewed item cannot outrank a higher-trust memory. Semantic result rows preserve distance and additionally expose similarity_score, trust_score, and the final score for transparency.

Relationship-aware recall (graph + tunnel expansion)

POST /v1/recall with mode=semantic doesn't stop at the nearest matches — it also reconstructs the context around them. Semantic search finds relevant memories; relationship expansion finds the decisions they were derived from, what they contradict or support, and their neighbors in the same tunnel. That surrounding context is often more valuable than the single closest-matching memory in isolation.

query → semantic retrieval → graph expansion → tunnel expansion → merge → relationship-aware rescoring → budget packing → response
  • Graph expansion follows typed, depth-1 edges (memory_edges: derived_from, references, explains, contradicts, supports, depends_on, mentions) from the top semantic candidates — bounded (5 neighbors/item, 20 total by default), deterministic, no recursion.
  • Tunnel expansion pulls bounded neighbors from a memory's tunnel-linked (wing, room) (20 additions by default) — no full tunnel scan.
  • Every expanded candidate passes the same tenant/visibility/review-status trust gate as a direct semantic hit — expansion can only narrow, never bypass.
  • Final score blends semantic relevance (70%), relationship strength (15%), tunnel membership (10%), and importance (5%) — semantic relevance still dominates. Reported as scoring_version = semantic-v3 in /v1/recall (mode=semantic) responses and recall_logs.
  • Every expanded item explains itself: "linked via derived_from", 'same tunnel "Atlas"', alongside the existing reasons array — merged/tagged with an origin (semantic, graph, tunnel, or a combination like semantic+graph).
  • The existing byte/token/item budget packer is unchanged and authoritative — expanded memories compete for budget exactly like direct hits.

Startup recall is unaffected (no query to anchor expansion from). See docs/design.md §9 for the full pipeline, edge-weight mapping, and configuration knobs (ENGRAM_MAX_GRAPH_*, ENGRAM_MAX_TUNNEL_*, ENGRAM_RECALL_CANDIDATE_CEILING).

Search filters

/v1/search honors kind, wing, and room filters (AND semantics) across keyword, semantic, and hybrid modes. Filters apply before ranking/limit and alongside tenant/read-eligibility scoping, so ineligible rows never displace matches.

Visibility & Multi-Tenancy

Engram supports visibility scopes for both single-agent and multi-agent deployments.

Visibility Who can read
private Only the principal that wrote it
workspace Any principal in the same workspace
tenant Any principal in the organization
public Any authenticated caller, where enabled

Row Level Security is enforced at the Postgres level — one forgotten WHERE clause cannot cause a cross-tenant leak. The runtime service connects as a dedicated non-owner role (engram_app) with no BYPASSRLS, and every tenant-scoped table uses FORCE ROW LEVEL SECURITY, so isolation holds even if the connecting role is the table owner. App-layer visibility/workspace logic is still the primary semantic rule; RLS is defense-in-depth beneath it.

Caller-facing item mutations apply the same eligibility rules as reads. A caller can therefore mutate a private memory only when it owns that memory, and can mutate a workspace-visible memory only while eligible for that workspace. Missing and inaccessible item IDs both return 404 Not Found so item existence is not disclosed. This eligibility check does not itself grant privileged review transitions; transition authority and route scope enforcement are separate trust controls.

API-Key Scopes & Authorization

Every API key carries an explicit, validated list of scopes. Scopes answer "may this credential attempt this class of operation?" — a question orthogonal to tenant membership, item eligibility, and principal type (agent vs. user vs. admin), which answer "may this specific principal perform this specific action?". All four checks are independently enforced; none substitutes for another.

The canonical scope vocabulary is exactly:

Scope Grants
read recall, search, item list/detail, taxonomy, tunnels, KG query/timeline, diary reads, classification
write remember, feedback, item metadata/supersede/invalidate, KG writes, diary writes, tunnel creation, and collaborative review actions (dispute, self-withdrawal)
review review queues/hygiene, privileged review decisions (activate, reactivate, reject, non-author archival), human verification, conflict resolution, bulk archival
export GET /v1/export/cca
admin tenant/workspace/principal/API-key/memory-kind governance, and every operation above

admin is a super-scope — a key carrying admin satisfies every other scope requirement automatically; you never need to also list read, write, review, or export alongside it. Scopes otherwise do not imply one another: write does not imply read, review does not imply write or read, and export does not imply read. A caller missing a required scope gets 403 Forbidden with a body like {"detail": "Requires scope: write"} (or "Requires one of scopes: write, review") — scope denial always happens before any handler-level mutation or eligibility disclosure.

For a memory-profile-bound key, scopes authorize the operation but do not bypass the profile's active read/write revision. Even admin and export remain narrowed on MemoryItem-backed data. The server pins one ResolvedMemoryContext per request; there is no request field, header, SDK option, MCP argument, or Hermes setting that selects another profile. Unprofiled keys retain compatibility behavior. Fully omitted write scope uses the profile default; tenant/public writes require their explicit write flags, every workspace association requires a writable grant, and existing-item mutation requires both read and write eligibility. Private no-workspace creation remains governed by API write scope, not include_private.

POST /v1/items/{item_id}/review is a mixed-purpose endpoint: an agent with only write may dispute an item or withdraw its own still-proposed proposal, but activating, reactivating, rejecting, or archiving someone else's proposal is a privileged review decision that requires review — even for a human user principal the principal-type policy would otherwise allow. Scope and principal-type policy are both required; neither one alone is sufficient (an agent with review still cannot activate anything — agent principal restrictions still apply).

Every route's scope requirement is discoverable in the OpenAPI schema (GET /openapi.json) under the x-engram-scope-policy vendor extension, e.g.:

{ "all_of": ["write"], "admin_satisfies": true }
{ "any_of": ["write", "review"], "admin_satisfies": true,
  "conditional": { "privileged_review_transitions": "review" } }
{ "exempt": true, "reason": "liveness probe" }

Only GET /health and GET /ready are exempt from scope enforcement; every other application route declares an explicit all_of/any_of requirement, and a completeness test fails CI if a new route ships without one.

Issuing keys. New-key issuance validates the requested scope list against the vocabulary above, rejects unknown/misspelled scopes with 422, dedupes, and persists them in a canonical order (read, write, review, export, admin). An explicit empty scope list ("scopes": []) is allowed and authenticates, but such a key can only reach /health//ready. Omitting scopes entirely defaults to ["read", "write"].

# Full administrator (bootstraps the very first key):
engram bootstrap-key --scopes read,write,admin,export

# Or via the admin API, once you have an admin-scoped key:
curl -X POST http://localhost:8000/v1/admin/api-keys \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "tenant_id": "<tenant-uuid>",
        "principal_id": "<principal-uuid>",
        "scopes": ["review"],
        "label": "human-reviewer"
      }'

Common issuance patterns:

Key purpose scopes
Read-only recall/search agent ["read"]
Write-only capture agent ["write"]
Human reviewer ["review"]
Export-only integration ["export"]
Full administrator ["admin"]

Backward compatibility. Existing keys with read/write/export/admin continue working unchanged under the new matrix, and existing admin-only keys (including bootstrap keys) automatically satisfy review via the super-scope rule — no data migration is needed. Historical rows may contain scope strings that predate validation (e.g. from hand-inserted keys); those authenticate normally, and any unrecognized string simply confers no authority rather than crashing authentication.

Development mode. With ENGRAM_AUTH_ENABLED=false (the default for local dev), every request resolves the seeded default admin principal, which carries admin and therefore passes every scope gate — the same runtime guards run in both modes; auth-disabled mode is not special-cased per route.

Engram's scopes are a custom bearer-token vocabulary, not OAuth2 — there is no token endpoint, no refresh flow, and no third-party identity provider integration.

Classification

Two endpoints serve classification with distinct roles:

  • POST /v1/classify returns taxonomy suggestions plus independent retention evidence and stores a one-hour, server-attested receipt. confidence remains a deprecated alias of taxonomy_confidence.
  • POST /v1/remember may bind that receipt with classification_run_id; write scope is still required and the server validates principal, content, source, workspace, expiry, and kind.

Taxonomy and retention confidences are separately clamped to 0.0–0.95. Retention dispositions are: retain for atomic information useful beyond the current working moment; transient for session-local utility; noise for acknowledgements, status/process chatter, and repeated tool output; and uncertain when durable value cannot be judged safely. Retention confidence measures only the positive case for durable retention, so confidence that content is noise does not raise it. Retention evidence is the classifier recommending retention — it is not epistemic or factual confidence, does not establish that the proposition is correct, does not affect recall ranking, and does not authorize a write.

The classifier's suggested_visibility is applied downward only — it may narrow the requested scope (e.g. tenantprivate) but never widen it. Invalid or absent suggestions preserve the caller's requested visibility unchanged.

Every receipt-bound write records its receipt/version, both confidence layers, source prior, visibility decision, reason, and allowlisted, context-sanitized provider provenance. A receipt is permanently consumed even if its bound item is later deleted. Dedup binding requires matching source and governed kind and can only narrow the existing visibility. Content is never mutated. Promotion Path A v2 consumes this bound evidence through its independently gated retention-evidence lane.

Seed classification rules are intentionally conservative: "skip" rules are whole-message status-only matchers (bare ok, done, passed) that don't fire on status words inside meaningful sentences, and doctrine classification requires explicit policy/invariant phrasing rather than casual modal verbs.

Memory Kinds

kind is governed by a tenant-scoped memory_kinds registry, not a hard-coded enum. Every tenant is seeded with nine built-in kinds — fact, preference, doctrine, decision, invariant, observation, diary_entry, procedure, summary — each carrying behavior flags (singleton, requires_review, stays_in_recall_when_disputed, default_importance) that drive supersession, initial review status, and disputed-recall inclusion. Classification and /v1/remember validate against the tenant's enabled registry rows; an unknown or disabled kind is rejected with a 422, never silently coerced to fact.

Tenant admins can add governed custom kinds without a schema migration:

GET    /v1/admin/memory-kinds            # list all (enabled + disabled)
POST   /v1/admin/memory-kinds            # create a custom kind (admin scope)
PATCH  /v1/admin/memory-kinds/{name}     # edit flags, enable/disable

Custom kind names must match ^[a-z][a-z0-9_]{0,63}$ and may not shadow a built-in name. name is immutable after creation. Disabling a kind blocks new writes/classification into it but never touches existing memories of that kind — deletion is intentionally unsupported (disabling is sufficient). See docs/design.md § Memory kinds for the full built-in behavior table.

Architecture

  • Postgres 16 + pgvector — single storage backend
  • FastAPI — REST core
  • MCP adapter — exposes Engram to MCP-compatible clients
  • Python SDK — thin client wrapper over the REST API
  • Multi-tenant from day onetenant_id on every tenant-scoped table
  • Row Level Security — database-enforced isolation
  • Append-first content — memory content is never silently overwritten
  • Audited metadata events — review status, visibility, taxonomy, and lifecycle changes are logged
  • Separate embeddings table — model-keyed embeddings support re-embedding without schema churn
  • Full-text search — generated tsvector column and GIN index
  • Tenant-configurable policy — scoring weights, trust defaults, recall policy, and promotion thresholds
  • Postgres job queue/v1/remember enqueues embedding generation and LLM classification refinement; a worker (engram worker) drains the queue off the request path so memory stays cheap per write

REST is the core interface. MCP and SDK integrations are adapters.

Agent frameworks are clients.

MCP Adapter

Engram includes an MCP server adapter for MCP-compatible clients such as Hermes, Claude Desktop, and other agent runtimes.

The MCP adapter exposes tools for:

  • remembering memories
  • startup and semantic recall
  • keyword, semantic, and hybrid search
  • classification
  • knowledge-graph queries
  • knowledge-graph additions
  • private diary entries

See adapters/mcp-server/README.md for MCP setup and tool signatures.

Vocabulary

Engram uses evocative naming drawn from memory palace traditions.

Engram term Plain-language equivalent
Wing Domain / category
Room Subcategory
Memory item A stored memory
Tunnel Cross-category link
Diary Principal-private journal
Doctrine Standing instruction / operating rule
Invariant Must-remain-true constraint

Status

Engram's MVP is implemented and dogfood-deployed — not a skeleton or a plan. The canonical memory service and agent adapters are exercised by a live, network-verified deployment. The current trust workflow is extensively covered by PostgreSQL-backed CI; deployment verification predates the latest trust and concurrency remediation, so those changes are not yet claimed as live-proven.

What "verified" means here: implemented = code exists and is unit/integration tested (CI runs the full suite against Postgres 16 + pgvector 0.8); dogfood-verified = exercised against the running deployment recorded in docs/ops/dogfood-verification.md; deferred = explicitly post-MVP (see docs/plans/engram-mvp-backlog.md).

MVP capability matrix

Capability Implemented Dogfood-verified
Schema + migrations (19 tables), RLS, FTS, pgvector storage yes yes
POST /v1/remember (trust fields, dedup, supersession, secret guard) yes yes
Startup recall (scoring, pinned bypass, anti-feedback loop, reasons) yes yes
Semantic recall (mode=semantic, proposed items tagged unreviewed) yes over FTS*
Keyword / semantic / hybrid search yes over FTS*
Item CRUD, PATCH with audited item_events, review/verify/supersede yes yes
Write-time conflict detection + resolution yes yes
Feedback endpoint + recall explanations + warnings yes yes
Knowledge graph (visibility inheritance), taxonomy, tunnels, diary yes yes
Relationship-aware recall (bounded graph + tunnel expansion) yes over FTS*
Memory hygiene (stale detection, bulk-archive, stats) yes
LLM + rule-based classification yes
Auto-promotion — Path A (age + confidence + no conflict) yes
CCA export + importers (CCA, MemPalace — dry-run/apply) yes
API-key auth + admin endpoints (scopes, bootstrap flow) yes yes
Python SDK (async client over REST) yes yes
MCP adapter (stdio, all tools) yes yes
Profile-keyed embeddings + zero-downtime re-embedding yes mocked only
Postgres job queue + engram worker (async embeddings, classification refine, conflict check) yes
Deployment artifacts (Compose, .env.example, backup, init-db) yes yes
Dogfood deployment (auth-enabled, network-reachable, backed up) yes yes

* Semantic/vector recall and search are implemented and tested, but the dogfood deployment runs with ENGRAM_EMBEDDING_PROVIDER=none (intentionally disabled for initial dogfooding). Keyword/FTS recall and search are verified live over the network. The live OpenAI embedding path has not been recorded-verified yet — see docs/embeddings.md.

Explicitly deferred (post-MVP)

These are intentionally out of scope for the MVP and tracked in docs/plans/engram-mvp-backlog.md:

  • Production data migration runs (CCA + MemPalace --apply against the live instance) — BL-011. The importers are built; only the operational import is pending.
  • engram-hooks / Hermes automatic lifecycle capture — BL-012 / ENG-HERMES-001. The compatibility shim (native-hook detection, monkey-patch, guard, idempotent install(), structured status) is implemented and unit tested (pytest -q adapters/engram-hooks/tests), and a documented Hermes dogfood profile now loads it (profiles/hermes-engram-dogfood.yaml, docs/ops/hermes-dogfood-profile.md). What's not yet done: a recorded end-to-end run against a real Hermes checkout confirming an automatic write reaches the deployed Engram instance. Explicit MCP-driven dogfooding works today regardless.
  • Auto-promotion Path B (usage-validated quorum).
  • Hard delete + deletion_events tombstones + KG cascade.
  • PII-risk classification and sensitive-read audit logging.
  • Admin list/update/delete + console; multi-way conflict table.
  • Local (non-OpenAI) embedding provider; Helm/cloud artifacts.
  • Phase 3 open-source packaging hardening.

Dogfooding

Engram runs on a dedicated VM ("engram01"), reachable over a Tailscale mesh, with auth enabled, a bootstrap API key issued, nightly pg_dump backups, and a restore smoke test. The full sanitized verification record — deployment, network health, authenticated remember→recall round trips, MCP adapter smoke, backup/restore — lives in docs/ops/dogfood-verification.md.

The dogfood interface is the MCP adapter: agents call engram_remember, engram_recall, and engram_search over stdio against the deployed instance. See adapters/mcp-server/README.md for setup and the Hermes config example.

Hostnames, IPs, and API keys are deliberately omitted from the public repo. Operators hold the real values in a secret manager.

Roadmap

Engram is being built in layers. The first four are largely landed; layer five (open-source readiness) is the active frontier.

  1. Canonical memory MVPdone. Durable storage, core schema, REST foundation, recall primitives, import/export.
  2. Trustable memory workflowdone. Classification, review states, promotion (Path A), disputes, conflict detection, provenance.
  3. Rich memory topologydone. Knowledge graph, tunnels, taxonomy browser, and relationship-aware recall.
  4. Agent integrationSDK and MCP done and verified; Hermes automatic lifecycle capture (engram-hooks) written but unverified (post-MVP). MCP, SDK, startup + semantic recall, and pre-compression/sync-turn capture (the latter awaiting engram-hooks verification).
  5. Verification and open-source readinessin progress. Post-remediation trust closure, real Hermes lifecycle capture, live embeddings, quality evaluations, examples, deployment hardening, security review, and hosted-service preparation. See the active verification ledger in docs/plans/post-remediation-verification-2026-07.md.

See docs/design.md for the full design document and docs/plans/engram-mvp-backlog.md for the execution backlog and post-MVP work.

License

Apache License 2.0

About

Zutfen Engram is a memory service that gives teams of AI agents a shared, structured, durable brain. It is an organized knowledge system with taxonomy, relationships, temporal validity, and per-agent scoping.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages