Benchmark: https://fastled.github.io/fbuild/ (fbuild 10.46s vs PlatformIO 6.75s cold, ratio 1.551, stacks matched).
To locate the gap I instrumented both tools on the same box (16-core Linux, identical pins: espressif32@6.13.0, toolchain 8.4.0+2021r2-patch5, framework 3.20017.241212+sha.dcc1105b, fixture bench/blink). PlatformIO has no built-in per-phase timer, so stats came from -v command tracing plus timestamp wrappers installed over both toolchains' g++/gcc/ar/size/objcopy binaries (removed afterwards). fbuild stats from FBUILD_PERF_LOG_JSON. Raw logs: 3 cold trials each, median — pio 4.20s, fbuild 5.71s (1.36x), same shape as CI.
Timelines
PlatformIO (4.20s):
0.00-0.80 SCons eval + LDF scan + ino preprocess (serial)
0.80-2.45 all 47 compiles, 16-way parallel, cpu-sum 16.2s
2.57-3.84 link 1.27s
3.84-4.20 size + esptool image 0.36s
fbuild (5.71s):
0.00-1.67 resolve/boot/config, zero compilers running (serial)
1.67-3.51 46 core compiles (window 1.84s, cpu-sum 21.8s)
3.52-4.10 sketch compile, strictly SERIAAL 0.59s
4.17-5.31 link 1.14s (parity)
5.31-5.70 convert/size/image 0.39s (parity)
Cause 1 — bloated -I list: ~55% slower per compiler invocation (est. 1.5-2.0s on CI)
fbuild passes 385 -I dirs; PlatformIO passes 199 (220 fbuild-only, all existing dirs — SDK component fan-out from get_sdk_include_dirs in crates/fbuild-library/src/library/esp32_framework/sdk_paths.rs:75 plus variant/src/toolchain extras). Toolchain and flags are otherwise identical (fbuild even omits pio's -ggdb, which should make it faster).
2x2 A/B on one TU (FunctionalInterrupt.cpp, 3 reps, min):
| command |
-I count |
time |
| fbuild flags + fbuild includes |
385 |
0.566s |
| fbuild flags + pio includes |
199 |
0.357s |
| pio flags + pio includes |
199 |
0.386s |
| pio flags + fbuild includes |
385 |
0.596s |
Swapping only the include list moves compile time ~+55% in both directions. Whole-build: identical 47-TU cpu-sum goes 16.2s (pio) -> 21.8s (fbuild, --jobs 16). PlatformIO's 199-dir set compiles the entire Arduino core, so it is a proven-sufficient baseline.
Fix: trim the SDK include fan-out to the parsed flags/includes set (what PlatformIO uses), dedupe against toolchain/variant dirs.
Cause 2 — sketch compile is serialized (~0.6s; matches CI compile-sketch: 589ms)
The ESP orchestrator runs strictly sequential phases: compile-core-variant -> core-cache-store -> compile-sketch -> link (crates/fbuild-build-esp/src/esp32/orchestrator/build.rs:742-790). PlatformIO compiles the sketch concurrently with the core files — zero wall cost. fbuild pays a full extra TU slot.
Fix: schedule sketch compile concurrently with the core fan-out (same jobs pool).
Cause 3 — pre-compile serial overhead: 1.67s vs pio's 0.80s (~+0.9s)
pioarduino-resolve ~150ms and boot-artifacts ~110ms every cold build
- ~0.7s in no perf phase at all: a constant gap between the
framework-libs cache and framework core cache log lines (0.21 -> 0.89; also present on warm builds)
- CI-side, Σ phases = 9.73s vs 10.46s total => ~0.73s unphased
Fix: attribute the unphased gap (likely package/cache metadata IO before the PerfTimer starts) and fold it into phases.
1.75 + 0.6 + 0.9 ≈ CI's 3.7s gap.
Ruled out by measurement
- Toolchain/flags: identical versions; flag-only diff favors fbuild (no
-ggdb).
- Link: fbuild 1.14s vs pio 1.27s — parity.
- Job oversubscription: default
jobs = ncpu*2 (32 on 16 cores) looked suspicious, but a controlled --jobs 16 cold run changed nothing (5.60s vs 5.57s wall).
Side finding (separate issue-worthy)
Local warm rebuilds take 1.8s while CI publishes 142ms. fast-path-check returns in 0.006ms but the full pipeline still runs (resolve 151ms + 46 up-to-date checks ~724ms + esptool ~0.15s). Stable across runs with a hot daemon and unaffected by toolchain mtimes — the fast-path early-return does not engage locally.
Benchmark: https://fastled.github.io/fbuild/ (fbuild 10.46s vs PlatformIO 6.75s cold, ratio 1.551, stacks matched).
To locate the gap I instrumented both tools on the same box (16-core Linux, identical pins:
espressif32@6.13.0, toolchain8.4.0+2021r2-patch5, framework3.20017.241212+sha.dcc1105b, fixturebench/blink). PlatformIO has no built-in per-phase timer, so stats came from-vcommand tracing plus timestamp wrappers installed over both toolchains'g++/gcc/ar/size/objcopybinaries (removed afterwards). fbuild stats fromFBUILD_PERF_LOG_JSON. Raw logs: 3 cold trials each, median — pio 4.20s, fbuild 5.71s (1.36x), same shape as CI.Timelines
PlatformIO (4.20s):
fbuild (5.71s):
Cause 1 — bloated
-Ilist: ~55% slower per compiler invocation (est. 1.5-2.0s on CI)fbuild passes 385
-Idirs; PlatformIO passes 199 (220 fbuild-only, all existing dirs — SDK component fan-out fromget_sdk_include_dirsincrates/fbuild-library/src/library/esp32_framework/sdk_paths.rs:75plus variant/src/toolchain extras). Toolchain and flags are otherwise identical (fbuild even omits pio's-ggdb, which should make it faster).2x2 A/B on one TU (
FunctionalInterrupt.cpp, 3 reps, min):-IcountSwapping only the include list moves compile time ~+55% in both directions. Whole-build: identical 47-TU cpu-sum goes 16.2s (pio) -> 21.8s (fbuild,
--jobs 16). PlatformIO's 199-dir set compiles the entire Arduino core, so it is a proven-sufficient baseline.Fix: trim the SDK include fan-out to the parsed
flags/includesset (what PlatformIO uses), dedupe against toolchain/variant dirs.Cause 2 — sketch compile is serialized (~0.6s; matches CI
compile-sketch: 589ms)The ESP orchestrator runs strictly sequential phases:
compile-core-variant->core-cache-store->compile-sketch-> link (crates/fbuild-build-esp/src/esp32/orchestrator/build.rs:742-790). PlatformIO compiles the sketch concurrently with the core files — zero wall cost. fbuild pays a full extra TU slot.Fix: schedule sketch compile concurrently with the core fan-out (same jobs pool).
Cause 3 — pre-compile serial overhead: 1.67s vs pio's 0.80s (~+0.9s)
pioarduino-resolve~150ms andboot-artifacts~110ms every cold buildframework-libs cacheandframework core cachelog lines (0.21 -> 0.89; also present on warm builds)Fix: attribute the unphased gap (likely package/cache metadata IO before the
PerfTimerstarts) and fold it into phases.1.75 + 0.6 + 0.9 ≈ CI's 3.7s gap.
Ruled out by measurement
-ggdb).jobs = ncpu*2(32 on 16 cores) looked suspicious, but a controlled--jobs 16cold run changed nothing (5.60s vs 5.57s wall).Side finding (separate issue-worthy)
Local warm rebuilds take 1.8s while CI publishes 142ms.
fast-path-checkreturns in 0.006ms but the full pipeline still runs (resolve 151ms + 46 up-to-date checks ~724ms + esptool ~0.15s). Stable across runs with a hot daemon and unaffected by toolchain mtimes — the fast-path early-return does not engage locally.