Skip to content

fbuild-daemon grows to ~3.9GB while idle then fails health check; heap-profile entry points not wired #1360

Description

@zackees

Summary

fbuild-daemon.exe grows to ~3.9 GB RSS and then fails its own health check, wedging every build with:

daemon error: daemon did not become healthy after 3 spawn attempts (10s each)
ERROR: Compilation failed for board teensy40

Killing the daemon restores normal operation immediately, so this is the daemon, not the build.

Evidence

Observed twice in one session on Windows 10 / x64. Sampling an idle daemon (no build in flight) three times in a row:

fbuild-daemon.exe 3,881,972 K
fbuild-daemon.exe 3,882,100 K
fbuild-daemon.exe 3,882,228 K

Memory is still climbing while idle — roughly 128 K per sample — so this reads as a slow leak that eventually crosses whatever threshold makes the health probe fail, rather than a legitimate working-set for a build.

Reproduction is not yet minimal. It appeared after a series of bash compile teensy40 --examples Blink runs interleaved with bash compile wasm --examples .... A second, healthy daemon process (~38 MB) coexisted with the bloated one at one point, which suggests a stale instance is not always reaped.

Recovery that works today:

taskkill /F /IM fbuild-daemon.exe

fbuild daemon stop returned success but did not clear the wedged process.

The profiling entry points exist but are not wired

The obvious next step is a heap profile, and zccache already ships the machinery — it just is not enabled in fbuild.

zccache/src/lib.rs gates it behind a feature and requires the embedding binary to choose the allocator:

/// Enable the `heap-profile` feature, install [`MiMalloc`] as the final
/// binary's `#[global_allocator]`, then use [`prof`] to start profiling and
/// write pprof-compatible snapshots. The allocator declaration belongs in
/// the embedding executable because a Rust library cannot select a global
/// allocator on its consumer's behalf.
#[cfg(feature = "heap-profile")]
pub mod heap_profile {
    pub use mimalloc_pprof::{
        enable_heap_profiling, enable_heap_profiling_with, prof, DumpFormat,
        MiMalloc, ProfConfig, ProfConfigMode,
    };
}

fbuild is half wired for this:

  • crates/fbuild-daemon/src/main.rs:1-2 does install a global allocator:
    #[global_allocator]
    static GLOBAL: mimalloc::MiMalloc = mimalloc::MiMalloc;
  • But it is plain mimalloc (Cargo.toml:130 — mimalloc = "0.1"), not mimalloc_pprof::MiMalloc, which is the one heap_profile needs.
  • And no crate enables the zccache heap-profile feature — every dep is a bare zccache = { git = ..., rev = ... } with no features.

So the pprof snapshot path is present in the dependency tree and unreachable from the shipped binary.

Suggested wiring

  1. Add features = ["heap-profile"] to the zccache dep on a crate the daemon links (e.g. crates/fbuild-build/Cargo.toml:57), ideally behind an fbuild-side heap-profile feature so release builds are unaffected.
  2. In crates/fbuild-daemon/src/main.rs, swap the allocator to zccache::heap_profile::MiMalloc under that feature.
  3. Expose a dump trigger — a fbuild daemon heap-dump subcommand, or an env var read at startup — that calls prof and writes a pprof snapshot.

Off-CPU profiling is also not reachable

zccache's ZCCACHE_DAEMON_PROFILE=tokio-console (zccache-daemon-core/src/daemon/entry.rs:594) initializes tokio-console tracing, which is the right tool for async stalls. But that lives in zccache-daemon-core, which drives the zccache daemon; fbuild-daemon is a separate binary with its own tracing setup and does not consult that variable. Wiring tokio-console into fbuild-daemon would be a separate, similarly small change and would help distinguish "leaking memory" from "blocked on a task that never completes".

Why this matters beyond the leak

The health-check failure surfaces as a build error (Compilation failed for board teensy40), which sends people looking at their code. The message should distinguish "daemon unhealthy — try fbuild daemon kill-all" from a genuine compile failure.

Activity

  1. zackees commented on Aug 23, 2026

    @zackees
    MemberAuthor

    Folding in two points from #1361 (closed as a duplicate of this), both about the "make sure it actually works" half rather than the wiring half.

    Symbolization is the thing most likely to be quietly broken, and Windows is where. A pprof heap dump that comes back as bare hex addresses is not a heap profile — it is a file that looks like one. Whether frames resolve is decided by [profile.release]'s debug-info settings and PDB availability, not by the profiler, so this needs checking against a release daemon before the wiring is called done. Worth verifying on all three CI platforms, since the failure mode differs per platform (PDB on Windows, DWARF/split-debug on Linux, Mach-O + dSYM on macOS).

    Argument for wiring the allocator unconditionally rather than behind an fbuild-side feature. The plan in the description suggests gating so release builds are unaffected, which is the safe default — but the cost of that default is that the profiler is never compiled into the binary on the machine where the leak actually reproduces. This report is a good example: it took two observations in one session on someone's real Windows box, and it is explicitly not minimal-reproducible yet. A feature-gated profiler would mean asking that person to build a custom binary and then reproduce a slow leak again.

    mimalloc_pprof::MiMalloc is the same mimalloc underneath, and its sampling profiler is off until prof is started, so the steady-state cost of shipping it unconditionally is close to the allocator swap alone. That is measurable: benchmark a build wall-clock with each allocator and let the number decide. If it is free, ship it always-on and the dump trigger becomes usable in the field; if it is not, gate it and accept the field cost. Either way the answer should come from a measurement, not from the default.

    Off-CPU: agreed the tokio-console path is separate, and worth noting it is the mode that matters most for this daemon — fbuild-daemon spends most of its wall clock waiting on subprocesses, sockets, and the filesystem, so an on-CPU-only profile is close to blind for its real latency. Also relevant to distinguishing this leak from a task that never completes and holds its allocations.

    One more thing from the report worth not losing: fbuild daemon stop returned success but left the wedged process alive. That is its own bug independent of the leak — a stop that reports success without stopping anything will keep masking this class of failure even after the leak is fixed.

  2. zackees commented on Aug 23, 2026

    @zackees
    MemberAuthor

    More reproductions, and a second failure mode

    This has now blocked four separate validation attempts in one session. Adding the data since the original report only had the health-check symptom.

    Memory, sampled across the session — three independent occurrences, each after a series of platform builds:

    fbuild-daemon.exe  3,881,972 K   (idle, still climbing)
    fbuild-daemon.exe  3,850,964 K
    fbuild-daemon.exe  4,152,624 K
    

    Each time, taskkill /F /IM fbuild-daemon.exe restored normal operation immediately. fbuild daemon stop reported success but did not clear the wedged process.

    Second failure mode — the daemon dies mid-build, not just at startup. The original report covered daemon did not become healthy after 3 spawn attempts. This one is different and produces a much more confusing message, because it surfaces as a build error:

    ERROR fbuild::output: daemon error: lost connection to daemon mid-build
      (error decoding response body); the daemon process may have died —
      check ~/.fbuild/prod/daemon/daemon.log
    ERROR: Compilation failed for board samd21
    

    The error decoding response body wording also appears as build error: package error: failed to read response body during toolchain downloads, so the same underlying symptom reads as a network problem in one place and a compile problem in another. Neither says "the daemon ran out of memory", which is what appears to have happened.

    Third observation — a 30-minute build timeout on esp32s3:

    BUILD FAIL Process timed out after 1800 seconds:
      fbuild.exe .build/pio/esp32s3 build -e esp32s3
    Compilation finished in 30m:14s with 1 error(s)
    

    Not proven to be the same root cause, but consistent with a daemon thrashing near its memory ceiling rather than compiling. Worth checking whether the timeout path and the mid-build-disconnect path share a cause.

    Why this matters beyond the leak

    All three surface as build failures on a specific board. Someone hitting this reasonably concludes their code broke that platform, and starts bisecting source that was never at fault. A message naming the daemon — and suggesting fbuild daemon kill-all — would save that entirely.

    The heap-profiling wiring described in the original report would settle what is actually retaining memory; the entry points exist in the embedded zccache lib but are not reachable from the shipped binary.

  3. zackees commented on Aug 23, 2026

    @zackees
    MemberAuthor

    Another occurrence, this time taking down a cold-provision build rather than a warm one:

    ERROR fbuild::output: daemon error: lost connection to daemon mid-build
      (error decoding response body); the daemon process may have died
    Compilation finished in 24m:59s with 1 error(s)
      - Board samd21 failed
    

    Context worth noting: this run was also fighting a toolchain download that never completed (#1370 — restarts from zero at ~90MB), so the daemon was alive and idle for ~25 minutes while ~450MB was pulled and discarded. The daemon died before any compilation started.

    That combination is the practical outcome: a cold bash compile samd21 on Windows cannot currently finish, and the two failures compound — the download keeps the daemon alive long enough for it to die. This blocked local verification of FastLED/FastLED#4015; I fell back to CI, where the toolchain is already provisioned and the daemon is short-lived.

  4. zackees commented on Aug 23, 2026

    @zackees
    MemberAuthor

    Status, so this stays open for the right reason.

    Landed

    • Heap profiling is wired and reachable — feat(daemon): adopt mimalloc-pprof and prove heap, on-CPU, and off-CPU profiling #1365 put mimalloc_pprof::MiMalloc behind a feature in fbuild-daemon, with FBUILD_HEAP_PROFILE read at startup and a loopback-gated heap_dump handler. That is the "suggested wiring" section of this issue.
    • fbuild daemon stop no longer reports success without stopping anything — fix(daemon): make daemon stop follow the process, not the endpoint #1379. It knew two states (endpoint answers / no daemon) and this report was the third: alive, owning the port, not answering /health. stop called that "not running", deleted the pid/port/claim records and exited 0, which is why taskkill /F was the only thing that worked and why the wedged process got harder to find afterwards. It now follows the process, terminates an unresponsive daemon, and confirms the process is gone before claiming it stopped.
    • The misleading surface from "Why this matters beyond the leak" — also fix(daemon): make daemon stop follow the process, not the endpoint #1379. The spawn failure now says up front that no compilation was attempted, and when a live daemon PID actually backs the claim it names the wedged case and points at fbuild daemon stop / kill-all. Without that evidence it says nothing rather than guessing over the version-mismatch explanation, which is the more common cause.

    Still open — the headline

    The ~3.9 GB idle growth itself is undiagnosed. The profiler exists now precisely so someone can answer it; nobody has run it against a daemon in this state and read the pprof output. Reproduction is still not minimal (interleaved teensy40 / wasm compiles).

    Also unaddressed: off-CPU profiling. tokio-console needs --cfg tokio_unstable across the build, which is a different change with a different blast radius than the heap-profile wiring was — worth splitting out rather than folding in.

  5. zackees commented on Sep 8, 2026

    @zackees
    MemberAuthor

    Same failure on a Linux GitHub Actions runner, and the memory hypothesis does not apply there. This issue documents Windows 10 with a daemon at ~3.9 GB RSS climbing while idle. The identical error appeared today on ubuntu-latest in FastLED CI:

    daemon error: daemon did not become healthy after 3 spawn attempts (10s each).
    This is a daemon problem, not a defect in the code being built --
    no compilation was attempted.
    running-process broker: direct daemon fallback
      (failed to connect to broker: No such file or directory (os error 2))
    

    FastLED PR #4211, job esp32s3_qemu_rmt / qemu_test, sketch BlinkParallel, run 34182903326.

    Why this is a different shape from the Windows case:

    • Fresh runner, no accumulated state. A GitHub-hosted runner starts clean, so the daemon had no opportunity to leak to 3.9 GB while idle. The three spawn attempts span 30 seconds total on a machine that had just started.
    • The daemon never became healthy at all, rather than degrading into unhealthiness after prolonged use. So "crosses a threshold after leaking" cannot be the mechanism here.
    • The broker socket was absent: failed to connect to broker: No such file or directory (os error 2). On Windows the daemon existed and was unhealthy; here it appears never to have come up.

    So either this issue covers two distinct causes behind one message, or the health check fails for a reason unrelated to memory. Worth separating, because the Windows remedy (kill the bloated daemon) has no analogue on a runner where there is nothing to kill.

    Two more daemon faults from the same day, on a NixOS bench, offered as related data rather than as this issue:

    1. The daemon died leaving a <defunct> zombie process and a stale socket; clients got Connection refused on its WebSocket. fbuild daemon restart recovered it. The user-visible symptom was open_port(/dev/ttyACM2) exceeded 3s; serial driver may be wedged -- i.e. it presented as dead hardware, not a dead daemon.
    2. A second RpcBench client on a port the daemon already held connected without error and was then completely inert, every call returning None (FastLED#4207).

    Three daemon faults, three different presentations. The CI message quoted above is the only one of the three that names the daemon as the culprit; the other two pointed at hardware and firmware respectively. Whatever the root causes turn out to be, that message wording is worth preserving -- it is what stopped me hunting for a regression in the PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions