Skip to content

EMBEDDING_MODEL defaults to 'bge-base' (768d), silently discarding every memory on a cloud-provider install #329

Description

@uknwmyname-sudo

Summary

On a self-hosted instance configured with EMBEDDING_PROVIDER=cloud-ensemble, every memory is accepted and stored but never embedded. embeddingStatus stays FAILED, the memories become permanently unsearchable, and nothing in /v1/health or /v1/stats indicates a problem. The only signal is a line in the container log.

We ran for roughly 24 hours believing memory was working before noticing.

Environment

  • Engram self-hosted from this repo (staging), cloned 2026-09-22
  • EDITION=local, Docker Compose as shipped
  • pgvector/pgvector:pg16, elasticsearch:8.14.0
  • EMBEDDING_PROVIDER=cloud-ensemble, OPENAI_API_KEY set
  • EMBEDDING_MODEL and VECTOR_SEARCH_MODEL not set (not mentioned in the quickstart)

What happens

resolveEmbeddingModelId() in src/vector/embedding-model.util.ts:

return (
  process.env.EMBEDDING_MODEL ??
  process.env.VECTOR_SEARCH_MODEL ??
  'bge-base'
);

MODEL_DIMS['bge-base'] is 768. The cloud ensemble produces 1536-dimensional vectors (openai-small). The pre-insert guard in PgVectorProvider.upsert therefore rejects every write:

[PgVector] Dimension mismatch for model 'bge-base': expected 768 dims but got 1536.
Check EMBEDDING_PROVIDER / LOCAL_EMBED_MODEL / EMBEDDING_MODEL alignment.

Once the retry budget is exhausted the memories are dead-lettered permanently:

[EmbeddingRetry] DEAD LETTER: 2 memories will never be searchable

POST /v1/memories/embedding-retry no longer picks them up:

[Retry] Embedding retry complete: 0/0 succeeded, 0 discovered from DB, 2 exhausted

The only recovery is to delete the memories and write them again.

Why this is hard to notice

  • GET /v1/health reports "status": "healthy" and dependencies.engramEmbed.status: "up" the entire time.
  • Startup logs report success:
    Cloud ensemble initialized with 2 models: openai-small, openai-large
    Embedding provider: cloud-ensemble (model: cloud-ensemble (openai-small), dims: 1536)
  • engram_recall still returns memories through a fallback path, so from an MCP client the system looks functional — results are just never actually ranked.
  • The failure is visible only in embeddingStatus on individual memories, or in the container log.

Steps to reproduce

  1. Clone the repo, set EMBEDDING_PROVIDER: cloud-ensemble and OPENAI_API_KEY in the api service environment.
  2. Do not set EMBEDDING_MODEL.
  3. docker compose up -d --build
  4. Store any memory (MCP engram_remember or POST /v1/memories).
  5. GET /v1/memories → "embeddingStatus": "FAILED", "embeddingModel": null.
  6. GET /v1/health → "status": "healthy".

Workaround

Setting EMBEDDING_MODEL: openai-small in the compose environment fixes it. Every subsequent write embeds correctly and ranking starts differentiating between records.

Suggested fixes

  1. Derive the default model id from the configured provider rather than hard-coding 'bge-base'. The provider already knows its own model name and dimension count — it logs both at startup.
  2. Fail fast at boot when the provider's dimension count does not match the resolved model's expected dimensions. All the information is available at that point.
  3. Expose embedding health in /v1/health — for example a count of memories with embeddingStatus != COMPLETE. A dead-lettered memory is unrecoverable, so a silent failure here is expensive.
  4. Consider letting POST /v1/memories/embedding-retry reset the retry counter, so operators can recover after fixing the configuration instead of losing the data.

Note

The comment above resolveEmbeddingModelId() documents this exact class of divergence ("Adversarial audit 2026-06-09 (Retrieval C1)"). The write/search divergence was fixed by centralising the helper, but the default value still points at the local model — so a cloud-provider install hits the same failure mode from the other side, and the docs do not mention that EMBEDDING_MODEL must be set.


Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions