Autonomous GPU Migration Platform — Starting with CUDA → ROCm. Designed for a future of portable GPU computing.
Built during the 🏆 AMD Developer Hackathon 2026 — An AI-powered platform that helps developers move CUDA applications to AMD ROCm using multi-agent reasoning.
The biggest challenge preventing organizations from adopting AMD GPUs isn't hardware. It isn't performance. It isn't software quality.
It's migration.
Thousands of CUDA applications remain locked to the NVIDIA ecosystem because migrating production GPU software is expensive, risky, time-consuming, and requires highly specialized expertise.
Kernel Olympics changes that.
Instead of acting as another AI coding assistant, Kernel Olympics functions as an Autonomous GPU Migration Platform that understands an entire CUDA project, plans the migration, performs intelligent code transformations, verifies correctness, benchmarks performance, explains every change, and produces a production-ready migration report.
Reduce GPU migration from weeks of engineering work to an AI-assisted workflow that developers can trust.
AMD GPUs (MI300X) outperform NVIDIA on price/performance. Yet enterprises stay on NVIDIA because 20% of CUDA code won't port to ROCm — custom kernels, warp-sensitive logic, library-specific calls. hipify handles the easy 80%. The remaining 20% is a manual, weeks-long slog per project.
AMD's #1 adoption blocker isn't hardware — it's software migration friction.
The broader market is bigger: GPU architectures multiply (NVIDIA CUDA, AMD ROCm, Intel oneAPI, Apple Metal, custom NPUs) while the talent pool doesn't. Every hardware generation creates a $2B+ migration tax across the industry — teams rewriting kernels by hand instead of building new products.
Kernel Olympics is not:
- ❌ another AI chatbot
- ❌ another GitHub Copilot
- ❌ another wrapper around hipify
- ❌ simple prompt engineering
Instead, Kernel Olympics behaves like an experienced GPU engineering team. Multiple specialized AI agents collaborate to understand an entire repository before making any modifications. Rather than translating files one by one, the platform reasons about architecture, dependencies, compatibility, performance implications, unsupported APIs, testing strategy, documentation, and migration risks.
Every decision is transparent. Every modification is explainable. Every migration produces evidence.
Kernel Olympics doesn't process kernels independently. Before modifying a single line, the platform analyzes the entire project context to identify:
- CUDA kernels and their call sites
- Runtime API usage and device memory operations
- Build configuration and dependencies
- Shared utilities and include hierarchy
- Unsupported CUDA features and migration complexity
Instead of blindly translating files, the platform understands the codebase as a whole — which dramatically improves migration quality.
Today, migrating GPU software is difficult because developers must manually:
- Understand large CUDA codebases
- Identify unsupported APIs
- Rewrite kernels
- Replace memory management
- Verify correctness
- Debug compilation failures
- Benchmark performance
- Write migration documentation
Even experienced GPU developers spend days or weeks doing this. The process is repetitive, error-prone, expensive, and hard to scale.
Kernel Olympics introduces an AI-native migration workflow. Instead of treating migration as file conversion, the platform treats it as an engineering reasoning problem. The system first understands the repository, identifies architectural patterns, analyzes dependencies, detects unsupported CUDA features, proposes migration strategies, executes intelligent transformations, validates results, benchmarks performance, and generates a comprehensive migration report explaining every decision.
This creates a migration pipeline that is explainable, repeatable, and significantly easier for developers to trust.
🐙 Live: kernel-olympics-production.up.railway.app — Upload a CUDA kernel and watch the pipeline port it live.
📡 Real AMD MI300X via notebooks.amd.com/hackathon — CUDA source → 4-LLM loop ports → hipcc compile → GPU run → PASSED ✓
| Metric | Value |
|---|---|
| Pipeline budget | 1,800s (30 min) |
| Max iterations | 10 (compile-fix loop) |
| LLM cost per run | ~$0.09 |
| Cache hit speed | ~0.2ms |
| Tests | 665 passing |
| CI/CD | ✅ Automated (GitHub Actions) |
| Hardware target | AMD MI300X via notebooks.amd.com/hackathon (192GB HBM3, CDNA3) |
┌─────────────────────────────────┐
│ CUDA Kernel (.cu) │
└──────────────┬──────────────────┘
│
▼
┌─────────────────────────────┐
│ 1. Risk Classifier │
│ (RED/YELLOW/GREEN) │
└──────────────┬──────────────┘
│
▼
┌─────────────────────────────┐
│ 2. Pattern Memory Cache │
│ (trigram ~0.2ms, 60,000×) │
└──────────────┬──────────────┘
│
╔══════════════╧══════════════╗
║ LLM Agent Loop ║
║ (full auto, no human) ║
║ ║
║ 3. DeepSeek-v4-Pro ║
║ → architecture plan ║
║ ║
║ 4. GLM-5.2 ║
║ → HIP kernel code gen ║
║ ║
║ 5. Kimi K2.7 ║
║ → 3-gate eval + refine ║
║ (compile-fix loop) ║
║ ║
║ 6. Gemma 4 / DeepSeek-v4 ║
║ → final verification ║
╚══════════════╧══════════════╝
│
▼
┌─────────────────────────────┐
│ 7. REAL AMD GPU │
│ hipcc + run + numerical │
│ diff verification │
└──────────────┬──────────────┘
│
▼
┌─────────────────────────────┐
│ ✅ HIP Kernel + Proof │
└─────────────────────────────┘
Every modification includes reasoning. Nothing is hidden. Every decision is transparent.
Original CUDA API
↓
Reason for replacement
↓
New ROCm implementation
↓
Performance implications
↓
Documentation
↓
Confidence Score
Instead of asking "What changed?", developers see why every change was made — making migration reviewable, auditable, and trustworthy.
Instead of calling expensive LLMs for every kernel, we cache porting patterns as trigram vectors:
| Metric | Without Cache | With Cache | Speedup |
|---|---|---|---|
| Pattern lookup | N/A | 0.2ms | — |
| LLM call (simulated) | ~12s | 0.2ms | 60,000× |
| Verified with live API | — | ✓ measured | ✓ |
The long-term vision extends far beyond CUDA → ROCm. Future versions can support:
- CUDA → ROCm (✅ currently supported)
- CUDA → SYCL
- CUDA → Vulkan Compute
- CUDA → OpenCL
- CUDA → Metal
- CUDA → DirectML
The platform is designed as a universal GPU migration engine rather than a single-purpose converter. Adding new target paths is a configuration change, not a rewrite.
→ kernel-olympics-production.up.railway.app — Upload a .cu file, see the autonomous pipeline port it to HIP in real time.
git clone https://github.com/indrad3v4/Kernel-Olympics.git
cd Kernel-Olympics
pip install -r requirements.txt
# Run the full pipeline on a sample kernel
make port CU_FILE=sample_kernels/cuda/warp_reduce.cu
# Or try a more complex warp-level scan kernel
make port CU_FILE=sample_kernels/cuda/nvidia_shfl_scan.cuFull guide: AMD_STACK_USAGE.md — but here's the 30-second version:
- AMD GPU (MI300X, MI250, RX 7900 XTX) with ROCm 6+ installed
- No AMD hardware? Use the AMD AI Notebooks portal — free MI300X instances for hackathon participants
hipccin PATH (hipcc --version)- Python 3.11+, git
# 1. Clone
git clone https://github.com/indrad3v4/Kernel-Olympics.git
cd Kernel-Olympics
# 2. Install deps
pip install -r requirements.txt
# 3. Set your API key (any OpenAI-compatible LLM)
export FIREWORKS_API_KEY="your-key"
# 4. Run the full pipeline on a real kernel (port → compile → run → verify)
make pipeline CU_FILE=sample_kernels/cuda/warp_reduce.cu
# Or with extended budget (30 min, 10 iterations)
make pipeline-heavy CU_FILE=sample_kernels/cuda/nvidia_shfl_scan.cu| Stage | Time | Output |
|---|---|---|
| Port (4-LLM loop) | 1–5 min | HIP code in ported_kernels/ |
| Compile (hipcc) | 2–5 sec | Binary in build/ |
| Run on AMD GPU | 1–3 sec | Numerical output |
| Verify | 0.5 sec | PASS ✓ or FAIL |
A successful run ends with:
✅ Pipeline complete — RESULT: PASS
The pipeline still ports + verifies. Just skip compile/run:
make port CU_FILE=sample_kernels/cuda/warp_reduce.cu
make inspect CU_FILE=sample_kernels/cuda/warp_reduce.cu # view ported codemake doctor # check ROCm, hipcc, API key
make test # 665 tests| Target | Description |
|---|---|
help |
Show all available targets |
install |
Create venv + install deps |
port |
Pipeline on one kernel: make port CU_FILE=path.cu |
port-all |
Pipeline on ALL sample kernels |
compile |
hipcc proof harness + compile |
run |
Run compiled binary on AMD GPU |
pipeline |
Full cycle: port → compile → run |
pipeline-heavy |
Extended budget (1,800s) |
test |
Run 665 pytest tests |
demo |
Live demo with recording |
inspect |
Inspect spec/ported kernel/proof |
debug-kernel |
Interactive kernel explorer |
retry |
Re-run a single pipeline stage |
Three levels of debugging for when things go wrong:
# Inspect specs and artifacts
make inspect CU_FILE=sample_kernels/cuda/warp_reduce.cu
make inspect PORTED=ported_kernels/warp_reduce.hip.cpp
# Interactive kernel exploration
make debug-kernel CU_FILE=sample_kernels/cuda/warp_reduce.cu
# Retry a single stage
make retry CU_FILE=sample_kernels/cuda/warp_reduce.cu STAGE=portmake test # full suite
make test-verbose # with progress| CUDA Intrinsic | HIP Equivalent | Action |
|---|---|---|
__shfl_up_sync(mask, val, d, w) |
__shfl_up(val, d, width) |
Mask dropped |
cudaMalloc() |
hipMalloc() |
1:1 rename |
cudaMemcpy() |
hipMemcpy() |
1:1 rename |
findCudaDevice() |
hipGetDevice() |
SDK strip |
sdkCreateTimer() |
— | Removed (NOP) |
threadIdx.x |
hipThreadIdx.x |
Namespace add |
- NVIDIA SDK symbols — auto-detected and stripped by
verifier.py - Wave64 divergence —
warpSizeconstant used instead of hardcoded 64 - SIGSEGV from host-code symbols — sanitizer in verifier catches at compile time
| Role | Member | Focus |
|---|---|---|
| 🚀 Lead | indradev_ | Architecture, pipeline orchestrator |
| 🔬 Kernel & Docs | Bromine185 | CUDA kernel analysis, warp primitives, documentation, demo recording |
| 🔧 Infra & Verification | icodemun44 | CI/CD, AMD cloud, Jupyter integration, verifier core, semantic repair engine, static analysis, tooling |
| 🧪 CI | meteorite67 | GitHub Actions, test suite |
| ✅ Verification | paparehan | Structural/lexical validation gates, safe writer, typed diagnostics |
| 🏭 AMD | Aahil-Riyaz (Satoru) | AMD MI300X testing, ROCm debugging, CI/CD |
MIT
Built for the AMD Developer Hackathon ACT II · Track 3 (Open Innovation)