Skip to content

fix(yeti): set the agents variables Yeti actually reads - #279

Merged
tomchop merged 2 commits into
google:mainfrom
tomchop:fix/yeti-agents-env-vars
Aug 31, 2026
Merged

tomchop merged 2 commits into
google:mainfrom
tomchop:fix/yeti-agents-env-vars

Conversation

@tomchop

@tomchop tomchop commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator

Follow-up to #277 and #278. The ChromaDB wiring in those is correct — I verified YETI_CHROMADB_HTTP_ROOT, /api/v2/heartbeat and mountPath: /data against chromadb/chroma:1.5.7 and yetiplatform/yeti:2.9.0. This fixes the equivalent problem on the agents side.

1. YETI_AGENTS_ENDPOINT is not a setting Yeti has

Yeti's config layer maps environment variables as YETI_<SECTION>_<KEY>, and the agents integration reads enabled, http_root and websocket_root (core/web/apiv2/agents.py:24-25). There is no endpoint key, so the variable was ignored.

It failed loudly rather than silently, because the image ships /app/yeti.conf. Running yetiplatform/yeti:2.9.0 with only the chart's variable set:

$ docker run -e YETI_AGENTS_ENDPOINT=http://rel-yeti-agents:8888 yetiplatform/yeti:2.9.0 ...
agents.enabled        = True
agents.http_root      = 'http://agents:8888'
agents.websocket_root = 'ws://agents:8888'

Those are the Docker Compose service names. In Kubernetes they do not resolve — and since enabled also falls back to True, the UI advertised the agents feature and it failed for every user.

-- name: YETI_AGENTS_ENDPOINT
-  value: "http://{{ .Release.Name }}-yeti-agents:8888"
+- name: YETI_AGENTS_ENABLED
+  value: "True"
+- name: YETI_AGENTS_HTTP_ROOT
+  value: "http://{{ .Release.Name }}-yeti-agents:8888"
+- name: YETI_AGENTS_WEBSOCKET_ROOT
+  value: "ws://{{ .Release.Name }}-yeti-agents:8888"

2. Disabling agents did not disable them either

Same root cause, opposite direction. With agents off the guard set no variable at all, so that same yeti.conf left enabled = True — the UI offered a feature with no backend deployed. The {{- else }} branch is as load-bearing as the if:

{{- else }}
- name: YETI_AGENTS_ENABLED
  value: "False"
{{- end }}

3. Agents required googleApiKeySecret but not global.yeti.apiKeySecret

apiKeySecret defaults to "", so with only a Google secret configured the workload deployed without YETI_API_KEY:

$ helm template ... --set agents.googleApiKeySecret=my-secret | grep -c "name: YETI_API_KEY"
0

The agents start, but semantic_search fails on every call — the tool being the entire reason they are wired to Yeti. Both credentials are now required before the workload renders.

The condition moves into a yeti.agentsEnabled helper rather than being repeated: the same answer decides the StatefulSet, its Service, and what the API is told about the feature, and those must not be able to disagree.

Verified

helm lint passes. All four configurations render as intended:

values agents workload YETI_AGENTS_ENABLED
defaults not rendered "False"
googleApiKeySecret only not rendered "False"
llmProvider=ollama + apiKeySecret rendered "True"
both secrets rendered "True"

With everything enabled, all five StatefulSets render (api … agents, arangodb, bloomcheck, chromadb, tasks) and the URLs are correct:

YETI_CHROMADB_HTTP_ROOT      http://t-yeti-chromadb:8000
YETI_AGENTS_HTTP_ROOT        http://t-yeti-agents:8888
YETI_AGENTS_WEBSOCKET_ROOT   ws://t-yeti-agents:8888

Deliberately not included

A readinessProbe on the agents workload. yeti-platform/yeti-agents#7 adds a /health endpoint reporting configuration state, which is the right probe target — but it is not merged, so yetiplatform/yeti-agents:dev does not serve it yet and a probe would fail every pod. Worth a follow-up once that image ships.

The yeti.envs scope. The agents pod receives YETI_ARANGODB_PASSWORD, YETI_AUTH_SECRET_KEY and YETI_USER_PASSWORD through the shared include and needs none of them. Out of scope here; better handled when that helper is split.

Suggestion

This is the second variable-name mismatch of this kind (YETI_CHROMA_HOST in #277 review, YETI_AGENTS_ENDPOINT here), and helm lint cannot catch either — they are semantically wrong, not malformed. A render-and-assert over the handful of variables Yeti actually reads would stop the next one:

helm template ... | grep -q 'name: YETI_AGENTS_HTTP_ROOT' || exit 1

YETI_AGENTS_ENDPOINT is not a setting Yeti has. Its config layer maps
YETI_<SECTION>_<KEY>, and the agents integration reads enabled, http_root and
websocket_root (core/web/apiv2/agents.py). The variable was therefore ignored,
and because the image ships a yeti.conf, the API fell back to its Compose
defaults: enabled=True with http_root=http://agents:8888, a name that does not
resolve in Kubernetes. The feature appeared in the UI and failed for every user.

The disabled path had the same cause and needed saying explicitly: with no
variable set at all, that same config file left the feature advertised with no
backend deployed.

Deploying agents also now requires global.yeti.apiKeySecret. The agents reach
Yeti over its API, so without those credentials the workload starts but every
tool call fails -- the tools being the reason it is connected to Yeti at all.
The condition moves into a yeti.agentsEnabled helper because the same answer
decides three things, the StatefulSet, its Service and what the API is told,
and those must not be able to disagree.
Enabling agents without credentials used to render nothing, so asking for the
service and not getting it looked the same as not asking. It now stops the
install and names the secret that is missing, with the command to create it:
the pod would otherwise start and be unable to answer, which is harder to work
out after the fact than a message at the point of the mistake.

agents.enabled therefore defaults to false. It cannot default to true while
this is fatal, because chart-testing installs with the chart's own defaults and
has no secrets to give it.

ChromaDB is enabled by default instead, which needs no credentials and which
Yeti otherwise falls back to an embedded index for -- one per process, so the
API and the task workers stop seeing the same data as soon as they are
scheduled apart.

Also adds a readiness probe on the agents /health endpoint, verified present in
yetiplatform/yeti-agents:dev. Readiness only: a configuration problem is not
fixed by restarting, and a liveness probe on the same endpoint would replace a
pod that can still be asked what is wrong with one that keeps dying.
@tomchop

tomchop commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator Author

Pushed two follow-ups after yeti-platform/yeti-agents#7 merged.

ChromaDB is now enabled by default

It needs no credentials, and without it Yeti falls back to an embedded index. That index is per-process, so the API and the task workers stop seeing the same data the moment they are scheduled onto different nodes — writes land in one and reads come back empty from the other, with no error on either side. Sharing a volume is not an alternative: the embedded backend is SQLite, whose locking is unreliable on network filesystems.

Asking for agents without credentials now fails the install

Previously agents.enabled: true with no secrets rendered nothing, so asking for the service and not getting it was indistinguishable from not asking. It now stops with the specific secret named and the command to create it:

$ helm template ... --set agents.enabled=true

agents.enabled is true but agents.googleApiKeySecret is not set.
The Yeti Agents service reads GOOGLE_API_KEY at startup and cannot run without it.

Create the secret and pass its name:
  kubectl create secret generic yeti-google-api-secret \
    --namespace default --from-literal=google-api-key=$GOOGLE_API_KEY
  helm upgrade ... --set agents.googleApiKeySecret=yeti-google-api-secret

Alternatively set agents.llmProvider=ollama to use a local model, or
agents.enabled=false to deploy Yeti without the agents service.

Missing global.yeti.apiKeySecret produces the equivalent message for that secret.

agents.enabled therefore defaults to false. It cannot default to true while this is fatal: ct install deploys with the chart's own defaults and has no secrets to supply, so every PR would fail CI. Defaulting to false keeps the default install working while making an explicit request for agents either succeed or explain itself.

Readiness probe on the agents service

yeti-agents#7 added /health, which reports whether the service found the configuration it needs. Verified against the published image rather than assumed:

$ docker run yetiplatform/yeti-agents:dev   # no configuration
$ curl localhost:8888/health
{"ready": false, "problems": [
  {"variable": "GOOGLE_API_KEY", "message": "...", "fatal": true},
  {"variable": "YETI_API_ROOT",  "message": "...", "fatal": false},
  {"variable": "YETI_API_KEY",   "message": "...", "fatal": false}
]}                                                            [HTTP 503]

Readiness only, no liveness probe: a configuration problem is not fixed by restarting, and a liveness probe on the same endpoint would replace a pod that can still be asked what is wrong with one that keeps dying.

Verified

helm lint passes. All five paths behave as intended:

values result
defaults (what ct install runs) renders; chromadb on, agents off
agents.enabled=true, no secrets fails, names googleApiKeySecret
agents.enabled=true + google secret only fails, names global.yeti.apiKeySecret
agents.enabled=true + llmProvider=ollama + api key renders
agents.enabled=true + both secrets renders, agents deployed

Chart versions bumped to yeti 2.6.0 / osdfir-infrastructure 2.12.0 — minor rather than patch, since enabling ChromaDB by default adds a StatefulSet and a 10Gi PVC to existing installs on upgrade.

@tomchop
tomchop requested a review from berggren August 31, 2026 14:37
@tomchop
tomchop merged commit 1e17df0 into google:main Aug 31, 2026
7 checks passed
@tomchop
tomchop deleted the fix/yeti-agents-env-vars branch August 31, 2026 15:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant