Fix GPU review findings (OpenCL RAII, CUDA 64-bit index, doc honesty) - #3
Merged
Merged
Conversation
Adversarial review of the GPU wind-tunnel solvers surfaced four issues: - OpenCL: constructor was not exception-safe — a clCheck throw mid-construction leaked the already-acquired cl objects. Moved all release into ~Impl (RAII), so a partially-constructed instance still cleans up. - CUDA: kernels used `long` for linear index arithmetic, which is 32-bit on LLP64 (Windows) and overflows for large grids. Switched to `long long`. - CUDA: added cudaGetLastError() after the kernel-launch loop so a launch fault is caught with step attribution (parity with the OpenCL clCheck path). - docs/backends.md: corrected a false "CUDA validated on NVIDIA hardware" claim (the CUDA path is authored without nvcc and validated on an NVIDIA machine). OpenCL wind tunnel still matches the CPU oracle (Linf/Uin ~ 9e-6). Archives the add-gpu-wind-tunnel OpenSpec change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adversarial-review fixes for the GPU wind-tunnel solvers: OpenCL constructor RAII (no leak on throw), CUDA
long->long long(Windows LLP64 index overflow), CUDA post-launchcudaGetLastError, and corrected a false 'CUDA validated' claim in docs/backends.md. OpenCL still matches the CPU oracle (Linf/Uin ~ 9e-6). Archives the add-gpu-wind-tunnel OpenSpec change.