Skip to content

Add justfile + GPU benchmark, validate CUDA/OpenCL, fix CI TBB linking - #4

Merged
leonardoaraujosantos merged 1 commit into
mainfrom
feat/justfile-bench-ci-fix
Jul 5, 2026
Merged

leonardoaraujosantos merged 1 commit into
mainfrom
feat/justfile-bench-ci-fix

Conversation

@leonardoaraujosantos

Copy link
Copy Markdown
Contributor

Summary

Adds a just-based developer workflow (modeled on the SciPP project), a GPU throughput benchmark, and fixes the Linux CI build that has been failing on every run since the GPU backends landed.

Along the way, the CUDA and OpenCL D3Q19 wind-tunnel solvers — previously authored without an NVIDIA machine and never compiled — were built and validated on real hardware (RTX 5060, Blackwell).

CI fix (the failing build)

The Linux job configured with -DCMAKE_EXE_LINKER_FLAGS="-ltbb". CMake places linker flags before the object files on the link line, so under the default --as-needed linker no symbols are yet pending when libtbb is seen and it gets dropped — every build failed with undefined reference to tbb::detail::r1::*. Adding -Wl,--no-as-needed pins libtbb regardless of order. This mirrors the justfile's configure recipe so local and CI builds link identically.

justfile

Recipe Purpose
just bootstrap Build + install NumPP into .deps/
just build / just test Configure + compile + full CTest (CPU)
just gpu-detect Probe host for CUDA / OpenCL / Metal, recommend a recipe
just gpu Auto-enable every backend present, build + test
just cuda / just opencl / just metal Build + run one backend's GPU test
just bench [nx ny nz steps] CPU vs enabled GPUs, in GLUPS
just spec / just ci / just clean / just debug OpenSpec validate / CI / cleanup / debug

Platform quirks are encoded once: Linux TBB linking (-Wl,--no-as-needed -ltbb), and CUDA built for sm_90 (SASS + PTX) so the driver JITs onto newer GPUs like Blackwell sm_120 with nvcc < 13.

Benchmark

bench/wind_tunnel_bench.cpp (behind -DCYBERFLUIDS_BUILD_BENCH=ON) times an identical wind-tunnel run per backend and reports GLUPS plus speed-up. Backend-guarded, so CPU-only hosts still build and get a CPU number.

Measured on RTX 5060 + i9-12900K (24 threads):

Backend Precision GLUPS vs CPU
CPU (par_unseq) fp64 ~0.038 1.0x
CUDA fp32 ~1.15 ~30x
OpenCL fp32 ~1.15 ~30x

CUDA ≈ OpenCL because the D3Q19 step is memory-bandwidth-bound (~78% of the card's DRAM peak) — both hit the same roofline.

Docs

  • docs/backends.md: CUDA marked ✅ validated with RTX 5060 numbers + a benchmark-reproduction section.
  • README.md: new just quick-start and GPU sections; corrected the stale "CUDA/OpenCL are still stubs" claim.

Testing

  • just cuda, just opencl, just gpu, just bench, just gpu-detect, just spec all pass locally on the RTX 5060 box.
  • GPU solvers agree with the fp64 CPU oracle to ~1e-5 (Linf/Uin).
  • The CI linker fix is verified by this PR's own CI run going green (the failure it fixes is a linker-flag issue, not covered by a unit test).

Developer workflow and GPU throughput tooling, plus a fix for the Linux CI
build that has been failing since the GPU backends landed.

justfile (mirrors the SciPP project):
- bootstrap, configure, build, test, spec, ci, clean
- gpu-detect: probe the host for CUDA / OpenCL / Metal and recommend a recipe
- cuda / opencl / metal / gpu: build and test individual or all GPU backends
- bench: build + run the throughput benchmark
- Encodes the platform quirks: Linux links the par_unseq backend with
  `-Wl,--no-as-needed -ltbb`; CUDA builds for sm_90 (SASS + PTX) so the driver
  JITs onto newer GPUs (e.g. Blackwell sm_120) with nvcc < 13.

bench/wind_tunnel_bench.cpp (+ CYBERFLUIDS_BUILD_BENCH option):
- Times an identical D3Q19 wind-tunnel run per backend, reports GLUPS and the
  speed-up over the CPU baseline. Backend-guarded, so CPU-only hosts still build
  and get a number. On RTX 5060: CUDA/OpenCL ~1.15 GLUPS, ~30x the 24-thread CPU.

CI fix (.github/workflows/ci.yml):
- The Linux job passed `-DCMAKE_EXE_LINKER_FLAGS="-ltbb"`, but linker flags land
  before the object files, so the default --as-needed dropped libtbb and every
  build failed with undefined `tbb::detail::r1::*` references. Add
  `-Wl,--no-as-needed` so libtbb is pinned regardless of link order, matching the
  justfile so local and CI builds link identically.

docs: mark CUDA validated (RTX 5060, Blackwell via PTX JIT, ~1e-5 vs the fp64
oracle) in docs/backends.md with a benchmark-reproduction section; refresh the
README GPU/quick-start sections and correct the stale "CUDA/OpenCL are stubs" note.
@leonardoaraujosantos
leonardoaraujosantos merged commit f3630d3 into main Jul 5, 2026
3 checks passed
@leonardoaraujosantos
leonardoaraujosantos deleted the feat/justfile-bench-ci-fix branch July 5, 2026 11:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant