Repository navigation
Add justfile + GPU benchmark, validate CUDA/OpenCL, fix CI TBB linking - #4
Merged
Merged
Conversation
Developer workflow and GPU throughput tooling, plus a fix for the Linux CI build that has been failing since the GPU backends landed. justfile (mirrors the SciPP project): - bootstrap, configure, build, test, spec, ci, clean - gpu-detect: probe the host for CUDA / OpenCL / Metal and recommend a recipe - cuda / opencl / metal / gpu: build and test individual or all GPU backends - bench: build + run the throughput benchmark - Encodes the platform quirks: Linux links the par_unseq backend with `-Wl,--no-as-needed -ltbb`; CUDA builds for sm_90 (SASS + PTX) so the driver JITs onto newer GPUs (e.g. Blackwell sm_120) with nvcc < 13. bench/wind_tunnel_bench.cpp (+ CYBERFLUIDS_BUILD_BENCH option): - Times an identical D3Q19 wind-tunnel run per backend, reports GLUPS and the speed-up over the CPU baseline. Backend-guarded, so CPU-only hosts still build and get a number. On RTX 5060: CUDA/OpenCL ~1.15 GLUPS, ~30x the 24-thread CPU. CI fix (.github/workflows/ci.yml): - The Linux job passed `-DCMAKE_EXE_LINKER_FLAGS="-ltbb"`, but linker flags land before the object files, so the default --as-needed dropped libtbb and every build failed with undefined `tbb::detail::r1::*` references. Add `-Wl,--no-as-needed` so libtbb is pinned regardless of link order, matching the justfile so local and CI builds link identically. docs: mark CUDA validated (RTX 5060, Blackwell via PTX JIT, ~1e-5 vs the fp64 oracle) in docs/backends.md with a benchmark-reproduction section; refresh the README GPU/quick-start sections and correct the stale "CUDA/OpenCL are stubs" note.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a
just-based developer workflow (modeled on the SciPP project), a GPU throughput benchmark, and fixes the Linux CI build that has been failing on every run since the GPU backends landed.Along the way, the CUDA and OpenCL D3Q19 wind-tunnel solvers — previously authored without an NVIDIA machine and never compiled — were built and validated on real hardware (RTX 5060, Blackwell).
CI fix (the failing build)
The Linux job configured with
-DCMAKE_EXE_LINKER_FLAGS="-ltbb". CMake places linker flags before the object files on the link line, so under the default--as-neededlinker no symbols are yet pending whenlibtbbis seen and it gets dropped — every build failed withundefined reference to tbb::detail::r1::*. Adding-Wl,--no-as-neededpins libtbb regardless of order. This mirrors the justfile'sconfigurerecipe so local and CI builds link identically.justfile
just bootstrap.deps/just build/just testjust gpu-detectjust gpujust cuda/just opencl/just metaljust bench [nx ny nz steps]just spec/just ci/just clean/just debugPlatform quirks are encoded once: Linux TBB linking (
-Wl,--no-as-needed -ltbb), and CUDA built forsm_90(SASS + PTX) so the driver JITs onto newer GPUs like Blackwellsm_120withnvcc < 13.Benchmark
bench/wind_tunnel_bench.cpp(behind-DCYBERFLUIDS_BUILD_BENCH=ON) times an identical wind-tunnel run per backend and reports GLUPS plus speed-up. Backend-guarded, so CPU-only hosts still build and get a CPU number.Measured on RTX 5060 + i9-12900K (24 threads):
par_unseq)CUDA ≈ OpenCL because the D3Q19 step is memory-bandwidth-bound (~78% of the card's DRAM peak) — both hit the same roofline.
Docs
docs/backends.md: CUDA marked ✅ validated with RTX 5060 numbers + a benchmark-reproduction section.README.md: newjustquick-start and GPU sections; corrected the stale "CUDA/OpenCL are still stubs" claim.Testing
just cuda,just opencl,just gpu,just bench,just gpu-detect,just specall pass locally on the RTX 5060 box.Linf/Uin).