Summary
While bringing up an mjlab / mujoco_warp RL training stack on MI300X (gfx942) with this ROCm Warp build (built from amd-integration, reports Warp 1.13.0+rocm.0), I hit two ROCm capability gaps that block otherwise CUDA-only downstream code. Both are understood and worked around on my side; filing here so they are tracked and, ideally, surfaced more clearly to downstream users.
The training itself works end to end on MI300X once these are handled (physics steps and the PPO update run on the GPU, sustained 100% utilization), so this is about closing capability gaps, not a fundamental blocker.
1. wp.Texture2D (hipArray textures) unsupported on ROCm
wp.Texture2D(...) fails on ROCm with:
RuntimeError: Failed to create CUDA texture: Warp CUDA error 801: operation not supported (in wp_texture_create_device, texture.cpp:171)
This surfaces in mujoco_warp's create_render_context, which materializes model textures even for raycast-only contexts. I have opened a mujoco_warp PR to guard that path with its existing use_textures flag (google-deepmind/mujoco_warp#1619), which avoids the call when no rendering is requested. But any downstream code that legitimately needs textures on ROCm (camera sensors) will still hit the underlying limitation.
Questions:
- Is hipArray texture-object support (
wp_texture_create_device) on the roadmap for this fork, or is there a documented reason it cannot be supported on gfx942?
- If it is a hard limitation, would it be possible to raise a clearer, HIP-specific error message (as is already done for conditional graphs, see below) so downstream users immediately understand it is a backend capability gap rather than a usage error?
2. Conditional graph nodes: is_cuda is True on ROCm, so downstream gating does not trip
The fork correctly and clearly refuses conditional graph nodes:
RuntimeError: Conditional graph nodes are not supported on HIP/ROCm
(assert_conditional_graph_support in warp/_src/context.py). The clear message is appreciated.
The interop trap is that a ROCm device reports as CUDA. mujoco_warp's solver gates conditional-graph use on wp.get_device().is_cuda (solver.py in main), assuming is_cuda == True implies conditional graphs are available. On ROCm, is_cuda is True (the device is exposed as cuda:0), so the gate stays open and the solver calls wp.capture_while, which then raises the error above. Downstream I work around it by forcing opt.graph_conditional = False.
Question:
- Would it make sense to expose a capability query (for example
wp.is_conditional_graph_supported(device) returning False on ROCm) so downstream libraries can gate on the actual capability instead of inferring it from is_cuda? That would let mujoco_warp/mjlab select the non-conditional solver path automatically on ROCm.
Environment
- GPU: AMD Instinct MI300X (gfx942), ROCm 10.0
- Warp: built from
amd-integration, banner Warp 1.13.0+rocm.0
- Downstream: mujoco_warp 3.8.1, mjlab 1.3.0, PyTorch 2.10.0+rocm7.0
Happy to provide full repros or test any changes on MI300X hardware.
Summary
While bringing up an mjlab / mujoco_warp RL training stack on MI300X (gfx942) with this ROCm Warp build (built from
amd-integration, reportsWarp 1.13.0+rocm.0), I hit two ROCm capability gaps that block otherwise CUDA-only downstream code. Both are understood and worked around on my side; filing here so they are tracked and, ideally, surfaced more clearly to downstream users.The training itself works end to end on MI300X once these are handled (physics steps and the PPO update run on the GPU, sustained 100% utilization), so this is about closing capability gaps, not a fundamental blocker.
1.
wp.Texture2D(hipArray textures) unsupported on ROCmwp.Texture2D(...)fails on ROCm with:This surfaces in mujoco_warp's
create_render_context, which materializes model textures even for raycast-only contexts. I have opened a mujoco_warp PR to guard that path with its existinguse_texturesflag (google-deepmind/mujoco_warp#1619), which avoids the call when no rendering is requested. But any downstream code that legitimately needs textures on ROCm (camera sensors) will still hit the underlying limitation.Questions:
wp_texture_create_device) on the roadmap for this fork, or is there a documented reason it cannot be supported on gfx942?2. Conditional graph nodes:
is_cudais True on ROCm, so downstream gating does not tripThe fork correctly and clearly refuses conditional graph nodes:
(
assert_conditional_graph_supportinwarp/_src/context.py). The clear message is appreciated.The interop trap is that a ROCm device reports as CUDA. mujoco_warp's solver gates conditional-graph use on
wp.get_device().is_cuda(solver.py inmain), assumingis_cuda == Trueimplies conditional graphs are available. On ROCm,is_cudaisTrue(the device is exposed ascuda:0), so the gate stays open and the solver callswp.capture_while, which then raises the error above. Downstream I work around it by forcingopt.graph_conditional = False.Question:
wp.is_conditional_graph_supported(device)returningFalseon ROCm) so downstream libraries can gate on the actual capability instead of inferring it fromis_cuda? That would let mujoco_warp/mjlab select the non-conditional solver path automatically on ROCm.Environment
amd-integration, bannerWarp 1.13.0+rocm.0Happy to provide full repros or test any changes on MI300X hardware.