Summary
On a self-hosted instance configured with EMBEDDING_PROVIDER=cloud-ensemble, every memory is accepted and stored but never embedded. embeddingStatus stays FAILED, the memories become permanently unsearchable, and nothing in /v1/health or /v1/stats indicates a problem. The only signal is a line in the container log.
We ran for roughly 24 hours believing memory was working before noticing.
Environment
- Engram self-hosted from this repo (
staging), cloned 2026-09-22
EDITION=local, Docker Compose as shipped
pgvector/pgvector:pg16, elasticsearch:8.14.0
EMBEDDING_PROVIDER=cloud-ensemble, OPENAI_API_KEY set
EMBEDDING_MODEL and VECTOR_SEARCH_MODEL not set (not mentioned in the quickstart)
What happens
resolveEmbeddingModelId() in src/vector/embedding-model.util.ts:
return (
process.env.EMBEDDING_MODEL ??
process.env.VECTOR_SEARCH_MODEL ??
'bge-base'
);
MODEL_DIMS['bge-base'] is 768. The cloud ensemble produces 1536-dimensional vectors (openai-small). The pre-insert guard in PgVectorProvider.upsert therefore rejects every write:
[PgVector] Dimension mismatch for model 'bge-base': expected 768 dims but got 1536.
Check EMBEDDING_PROVIDER / LOCAL_EMBED_MODEL / EMBEDDING_MODEL alignment.
Once the retry budget is exhausted the memories are dead-lettered permanently:
[EmbeddingRetry] DEAD LETTER: 2 memories will never be searchable
POST /v1/memories/embedding-retry no longer picks them up:
[Retry] Embedding retry complete: 0/0 succeeded, 0 discovered from DB, 2 exhausted
The only recovery is to delete the memories and write them again.
Why this is hard to notice
GET /v1/health reports "status": "healthy" and dependencies.engramEmbed.status: "up" the entire time.
- Startup logs report success:
Cloud ensemble initialized with 2 models: openai-small, openai-large
Embedding provider: cloud-ensemble (model: cloud-ensemble (openai-small), dims: 1536)
engram_recall still returns memories through a fallback path, so from an MCP client the system looks functional — results are just never actually ranked.
- The failure is visible only in
embeddingStatus on individual memories, or in the container log.
Steps to reproduce
- Clone the repo, set
EMBEDDING_PROVIDER: cloud-ensemble and OPENAI_API_KEY in the api service environment.
- Do not set
EMBEDDING_MODEL.
docker compose up -d --build
- Store any memory (MCP
engram_remember or POST /v1/memories).
GET /v1/memories → "embeddingStatus": "FAILED", "embeddingModel": null.
GET /v1/health → "status": "healthy".
Workaround
Setting EMBEDDING_MODEL: openai-small in the compose environment fixes it. Every subsequent write embeds correctly and ranking starts differentiating between records.
Suggested fixes
- Derive the default model id from the configured provider rather than hard-coding
'bge-base'. The provider already knows its own model name and dimension count — it logs both at startup.
- Fail fast at boot when the provider's dimension count does not match the resolved model's expected dimensions. All the information is available at that point.
- Expose embedding health in
/v1/health — for example a count of memories with embeddingStatus != COMPLETE. A dead-lettered memory is unrecoverable, so a silent failure here is expensive.
- Consider letting
POST /v1/memories/embedding-retry reset the retry counter, so operators can recover after fixing the configuration instead of losing the data.
Note
The comment above resolveEmbeddingModelId() documents this exact class of divergence ("Adversarial audit 2026-06-09 (Retrieval C1)"). The write/search divergence was fixed by centralising the helper, but the default value still points at the local model — so a cloud-provider install hits the same failure mode from the other side, and the docs do not mention that EMBEDDING_MODEL must be set.
Summary
On a self-hosted instance configured with
EMBEDDING_PROVIDER=cloud-ensemble, every memory is accepted and stored but never embedded.embeddingStatusstaysFAILED, the memories become permanently unsearchable, and nothing in/v1/healthor/v1/statsindicates a problem. The only signal is a line in the container log.We ran for roughly 24 hours believing memory was working before noticing.
Environment
staging), cloned 2026-09-22EDITION=local, Docker Compose as shippedpgvector/pgvector:pg16,elasticsearch:8.14.0EMBEDDING_PROVIDER=cloud-ensemble,OPENAI_API_KEYsetEMBEDDING_MODELandVECTOR_SEARCH_MODELnot set (not mentioned in the quickstart)What happens
resolveEmbeddingModelId()insrc/vector/embedding-model.util.ts:MODEL_DIMS['bge-base']is768. The cloud ensemble produces 1536-dimensional vectors (openai-small). The pre-insert guard inPgVectorProvider.upserttherefore rejects every write:Once the retry budget is exhausted the memories are dead-lettered permanently:
POST /v1/memories/embedding-retryno longer picks them up:The only recovery is to delete the memories and write them again.
Why this is hard to notice
GET /v1/healthreports"status": "healthy"anddependencies.engramEmbed.status: "up"the entire time.Cloud ensemble initialized with 2 models: openai-small, openai-largeEmbedding provider: cloud-ensemble (model: cloud-ensemble (openai-small), dims: 1536)engram_recallstill returns memories through a fallback path, so from an MCP client the system looks functional — results are just never actually ranked.embeddingStatuson individual memories, or in the container log.Steps to reproduce
EMBEDDING_PROVIDER: cloud-ensembleandOPENAI_API_KEYin theapiservice environment.EMBEDDING_MODEL.docker compose up -d --buildengram_rememberorPOST /v1/memories).GET /v1/memories→"embeddingStatus": "FAILED","embeddingModel": null.GET /v1/health→"status": "healthy".Workaround
Setting
EMBEDDING_MODEL: openai-smallin the compose environment fixes it. Every subsequent write embeds correctly and ranking starts differentiating between records.Suggested fixes
'bge-base'. The provider already knows its own model name and dimension count — it logs both at startup./v1/health— for example a count of memories withembeddingStatus != COMPLETE. A dead-lettered memory is unrecoverable, so a silent failure here is expensive.POST /v1/memories/embedding-retryreset the retry counter, so operators can recover after fixing the configuration instead of losing the data.Note
The comment above
resolveEmbeddingModelId()documents this exact class of divergence ("Adversarial audit 2026-06-09 (Retrieval C1)"). The write/search divergence was fixed by centralising the helper, but the default value still points at the local model — so a cloud-provider install hits the same failure mode from the other side, and the docs do not mention thatEMBEDDING_MODELmust be set.