Repository navigation
fbuild-daemon grows to ~3.9GB while idle then fails health check; heap-profile entry points not wired #1360
Description
Activity
Folding in two points from #1361 (closed as a duplicate of this), both about the "make sure it actually works" half rather than the wiring half.
Symbolization is the thing most likely to be quietly broken, and Windows is where. A pprof heap dump that comes back as bare hex addresses is not a heap profile — it is a file that looks like one. Whether frames resolve is decided by
[profile.release]'s debug-info settings and PDB availability, not by the profiler, so this needs checking against a release daemon before the wiring is called done. Worth verifying on all three CI platforms, since the failure mode differs per platform (PDB on Windows, DWARF/split-debug on Linux, Mach-O + dSYM on macOS).Argument for wiring the allocator unconditionally rather than behind an fbuild-side feature. The plan in the description suggests gating so release builds are unaffected, which is the safe default — but the cost of that default is that the profiler is never compiled into the binary on the machine where the leak actually reproduces. This report is a good example: it took two observations in one session on someone's real Windows box, and it is explicitly not minimal-reproducible yet. A feature-gated profiler would mean asking that person to build a custom binary and then reproduce a slow leak again.
mimalloc_pprof::MiMallocis the same mimalloc underneath, and its sampling profiler is off untilprofis started, so the steady-state cost of shipping it unconditionally is close to the allocator swap alone. That is measurable: benchmark a build wall-clock with each allocator and let the number decide. If it is free, ship it always-on and the dump trigger becomes usable in the field; if it is not, gate it and accept the field cost. Either way the answer should come from a measurement, not from the default.Off-CPU: agreed the tokio-console path is separate, and worth noting it is the mode that matters most for this daemon — fbuild-daemon spends most of its wall clock waiting on subprocesses, sockets, and the filesystem, so an on-CPU-only profile is close to blind for its real latency. Also relevant to distinguishing this leak from a task that never completes and holds its allocations.
One more thing from the report worth not losing:
fbuild daemon stopreturned success but left the wedged process alive. That is its own bug independent of the leak — a stop that reports success without stopping anything will keep masking this class of failure even after the leak is fixed.More reproductions, and a second failure mode
This has now blocked four separate validation attempts in one session. Adding the data since the original report only had the health-check symptom.
Memory, sampled across the session — three independent occurrences, each after a series of platform builds:
fbuild-daemon.exe 3,881,972 K (idle, still climbing) fbuild-daemon.exe 3,850,964 K fbuild-daemon.exe 4,152,624 KEach time,
taskkill /F /IM fbuild-daemon.exerestored normal operation immediately.fbuild daemon stopreported success but did not clear the wedged process.Second failure mode — the daemon dies mid-build, not just at startup. The original report covered
daemon did not become healthy after 3 spawn attempts. This one is different and produces a much more confusing message, because it surfaces as a build error:ERROR fbuild::output: daemon error: lost connection to daemon mid-build (error decoding response body); the daemon process may have died — check ~/.fbuild/prod/daemon/daemon.log ERROR: Compilation failed for board samd21The
error decoding response bodywording also appears asbuild error: package error: failed to read response bodyduring toolchain downloads, so the same underlying symptom reads as a network problem in one place and a compile problem in another. Neither says "the daemon ran out of memory", which is what appears to have happened.Third observation — a 30-minute build timeout on esp32s3:
BUILD FAIL Process timed out after 1800 seconds: fbuild.exe .build/pio/esp32s3 build -e esp32s3 Compilation finished in 30m:14s with 1 error(s)Not proven to be the same root cause, but consistent with a daemon thrashing near its memory ceiling rather than compiling. Worth checking whether the timeout path and the mid-build-disconnect path share a cause.
Why this matters beyond the leak
All three surface as build failures on a specific board. Someone hitting this reasonably concludes their code broke that platform, and starts bisecting source that was never at fault. A message naming the daemon — and suggesting
fbuild daemon kill-all— would save that entirely.The heap-profiling wiring described in the original report would settle what is actually retaining memory; the entry points exist in the embedded zccache lib but are not reachable from the shipped binary.
Another occurrence, this time taking down a cold-provision build rather than a warm one:
ERROR fbuild::output: daemon error: lost connection to daemon mid-build (error decoding response body); the daemon process may have died Compilation finished in 24m:59s with 1 error(s) - Board samd21 failedContext worth noting: this run was also fighting a toolchain download that never completed (#1370 — restarts from zero at ~90MB), so the daemon was alive and idle for ~25 minutes while ~450MB was pulled and discarded. The daemon died before any compilation started.
That combination is the practical outcome: a cold
bash compile samd21on Windows cannot currently finish, and the two failures compound — the download keeps the daemon alive long enough for it to die. This blocked local verification of FastLED/FastLED#4015; I fell back to CI, where the toolchain is already provisioned and the daemon is short-lived.- added a commit that references this issue
on Aug 23, 2026 Status, so this stays open for the right reason.
Landed
- Heap profiling is wired and reachable — feat(daemon): adopt mimalloc-pprof and prove heap, on-CPU, and off-CPU profiling #1365 put
mimalloc_pprof::MiMallocbehind a feature infbuild-daemon, withFBUILD_HEAP_PROFILEread at startup and a loopback-gatedheap_dumphandler. That is the "suggested wiring" section of this issue. fbuild daemon stopno longer reports success without stopping anything — fix(daemon): makedaemon stopfollow the process, not the endpoint #1379. It knew two states (endpoint answers / no daemon) and this report was the third: alive, owning the port, not answering/health.stopcalled that "not running", deleted the pid/port/claim records and exited 0, which is whytaskkill /Fwas the only thing that worked and why the wedged process got harder to find afterwards. It now follows the process, terminates an unresponsive daemon, and confirms the process is gone before claiming it stopped.- The misleading surface from "Why this matters beyond the leak" — also fix(daemon): make
daemon stopfollow the process, not the endpoint #1379. The spawn failure now says up front that no compilation was attempted, and when a live daemon PID actually backs the claim it names the wedged case and points atfbuild daemon stop/kill-all. Without that evidence it says nothing rather than guessing over the version-mismatch explanation, which is the more common cause.
Still open — the headline
The ~3.9 GB idle growth itself is undiagnosed. The profiler exists now precisely so someone can answer it; nobody has run it against a daemon in this state and read the pprof output. Reproduction is still not minimal (interleaved
teensy40/wasmcompiles).Also unaddressed: off-CPU profiling.
tokio-consoleneeds--cfg tokio_unstableacross the build, which is a different change with a different blast radius than the heap-profile wiring was — worth splitting out rather than folding in.- Heap profiling is wired and reachable — feat(daemon): adopt mimalloc-pprof and prove heap, on-CPU, and off-CPU profiling #1365 put
Same failure on a Linux GitHub Actions runner, and the memory hypothesis does not apply there. This issue documents Windows 10 with a daemon at ~3.9 GB RSS climbing while idle. The identical error appeared today on
ubuntu-latestin FastLED CI:daemon error: daemon did not become healthy after 3 spawn attempts (10s each). This is a daemon problem, not a defect in the code being built -- no compilation was attempted. running-process broker: direct daemon fallback (failed to connect to broker: No such file or directory (os error 2))FastLED PR #4211, job
esp32s3_qemu_rmt / qemu_test, sketchBlinkParallel, run 34182903326.Why this is a different shape from the Windows case:
- Fresh runner, no accumulated state. A GitHub-hosted runner starts clean, so the daemon had no opportunity to leak to 3.9 GB while idle. The three spawn attempts span 30 seconds total on a machine that had just started.
- The daemon never became healthy at all, rather than degrading into unhealthiness after prolonged use. So "crosses a threshold after leaking" cannot be the mechanism here.
- The broker socket was absent:
failed to connect to broker: No such file or directory (os error 2). On Windows the daemon existed and was unhealthy; here it appears never to have come up.
So either this issue covers two distinct causes behind one message, or the health check fails for a reason unrelated to memory. Worth separating, because the Windows remedy (kill the bloated daemon) has no analogue on a runner where there is nothing to kill.
Two more daemon faults from the same day, on a NixOS bench, offered as related data rather than as this issue:
- The daemon died leaving a
<defunct>zombie process and a stale socket; clients gotConnection refusedon its WebSocket.fbuild daemon restartrecovered it. The user-visible symptom wasopen_port(/dev/ttyACM2) exceeded 3s; serial driver may be wedged-- i.e. it presented as dead hardware, not a dead daemon. - A second
RpcBenchclient on a port the daemon already held connected without error and was then completely inert, every call returningNone(FastLED#4207).
Three daemon faults, three different presentations. The CI message quoted above is the only one of the three that names the daemon as the culprit; the other two pointed at hardware and firmware respectively. Whatever the root causes turn out to be, that message wording is worth preserving -- it is what stopped me hunting for a regression in the PR.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTriage
Summary
fbuild-daemon.exegrows to ~3.9 GB RSS and then fails its own health check, wedging every build with:Killing the daemon restores normal operation immediately, so this is the daemon, not the build.
Evidence
Observed twice in one session on Windows 10 / x64. Sampling an idle daemon (no build in flight) three times in a row:
Memory is still climbing while idle — roughly 128 K per sample — so this reads as a slow leak that eventually crosses whatever threshold makes the health probe fail, rather than a legitimate working-set for a build.
Reproduction is not yet minimal. It appeared after a series of
bash compile teensy40 --examples Blinkruns interleaved withbash compile wasm --examples .... A second, healthy daemon process (~38 MB) coexisted with the bloated one at one point, which suggests a stale instance is not always reaped.Recovery that works today:
fbuild daemon stopreturned success but did not clear the wedged process.The profiling entry points exist but are not wired
The obvious next step is a heap profile, and zccache already ships the machinery — it just is not enabled in fbuild.
zccache/src/lib.rsgates it behind a feature and requires the embedding binary to choose the allocator:fbuild is half wired for this:
crates/fbuild-daemon/src/main.rs:1-2does install a global allocator:mimalloc(Cargo.toml:130—mimalloc = "0.1"), notmimalloc_pprof::MiMalloc, which is the oneheap_profileneeds.heap-profilefeature — every dep is a barezccache = { git = ..., rev = ... }with nofeatures.So the pprof snapshot path is present in the dependency tree and unreachable from the shipped binary.
Suggested wiring
features = ["heap-profile"]to the zccache dep on a crate the daemon links (e.g.crates/fbuild-build/Cargo.toml:57), ideally behind an fbuild-sideheap-profilefeature so release builds are unaffected.crates/fbuild-daemon/src/main.rs, swap the allocator tozccache::heap_profile::MiMallocunder that feature.fbuild daemon heap-dumpsubcommand, or an env var read at startup — that callsprofand writes a pprof snapshot.Off-CPU profiling is also not reachable
zccache's
ZCCACHE_DAEMON_PROFILE=tokio-console(zccache-daemon-core/src/daemon/entry.rs:594) initializes tokio-console tracing, which is the right tool for async stalls. But that lives inzccache-daemon-core, which drives the zccache daemon;fbuild-daemonis a separate binary with its own tracing setup and does not consult that variable. Wiring tokio-console intofbuild-daemonwould be a separate, similarly small change and would help distinguish "leaking memory" from "blocked on a task that never completes".Why this matters beyond the leak
The health-check failure surfaces as a build error (
Compilation failed for board teensy40), which sends people looking at their code. The message should distinguish "daemon unhealthy — tryfbuild daemon kill-all" from a genuine compile failure.