Repository navigation
feat(bench): validate CPU/GPU 2-opt selection - #12
Merged
Merged
Conversation
Signed-off-by: Guilherme Pedroza <guilhermebarb0sa@proton.me>
Signed-off-by: Guilherme Pedroza <guilhermebarb0sa@proton.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a correctness-first, reproducible comparison of the existing CPU and GPU 2-opt move-selection APIs. The runner validates all candidate deltas, the selected move, and the cost after reversal before recording full public-API latency. It changes no public library, kernel, or PTX code.
Key Changes
Validation runner
vrp-gpu-benchpackage.+inf, and compares selected indices exactly.f64after applying a selected reversal. Rejects invalid or divergent results before writing a report.Reproducible timing and results
Observed Result
The recorded RTX 5060 Ti run passed all delta, selection, and reversal-cost checks on nine synthetic cases and ten C101 routes. For the 2048-customer synthetic route, median complete API latency was 25.52 ms on CPU and 156.17 ms on GPU. The current GPU API was slower for every measured route with at least two customers. These measurements include GPU context/module setup, validation, transfers, kernels, synchronization, and download; they do not isolate kernel time or compare converged solvers.
Commits Included
feat(bench): validate CPU GPU 2-opt selectiondocs(bench): record CPU GPU validation resultsVerification
just ci: format, all-feature check, strict Clippy, 43 library tests, five runner tests, one doctest, rustdoc, package verification, and cargo-deny.Signed-off-bytrailers.Scope
This PR validates one-move selection and its full public-API latency. It does not implement GPU route convergence, compare equal-budget complete solvers, measure kernel-only throughput, or claim support beyond the tested hardware.