Skip to content

Latest commit

 

History

335 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🚀 Kernel Olympics

Autonomous GPU Migration Platform — Starting with CUDA → ROCm. Designed for a future of portable GPU computing.

Built during the 🏆 AMD Developer Hackathon 2026 — An AI-powered platform that helps developers move CUDA applications to AMD ROCm using multi-agent reasoning.


🌍 Why Kernel Olympics Exists

The biggest challenge preventing organizations from adopting AMD GPUs isn't hardware. It isn't performance. It isn't software quality.

It's migration.

Thousands of CUDA applications remain locked to the NVIDIA ecosystem because migrating production GPU software is expensive, risky, time-consuming, and requires highly specialized expertise.

Kernel Olympics changes that.

Instead of acting as another AI coding assistant, Kernel Olympics functions as an Autonomous GPU Migration Platform that understands an entire CUDA project, plans the migration, performs intelligent code transformations, verifies correctness, benchmarks performance, explains every change, and produces a production-ready migration report.

Reduce GPU migration from weeks of engineering work to an AI-assisted workflow that developers can trust.


The $10B Problem

AMD GPUs (MI300X) outperform NVIDIA on price/performance. Yet enterprises stay on NVIDIA because 20% of CUDA code won't port to ROCm — custom kernels, warp-sensitive logic, library-specific calls. hipify handles the easy 80%. The remaining 20% is a manual, weeks-long slog per project.

AMD's #1 adoption blocker isn't hardware — it's software migration friction.

The broader market is bigger: GPU architectures multiply (NVIDIA CUDA, AMD ROCm, Intel oneAPI, Apple Metal, custom NPUs) while the talent pool doesn't. Every hardware generation creates a $2B+ migration tax across the industry — teams rewriting kernels by hand instead of building new products.


✨ What Makes Kernel Olympics Different?

Kernel Olympics is not:

  • ❌ another AI chatbot
  • ❌ another GitHub Copilot
  • ❌ another wrapper around hipify
  • ❌ simple prompt engineering

Instead, Kernel Olympics behaves like an experienced GPU engineering team. Multiple specialized AI agents collaborate to understand an entire repository before making any modifications. Rather than translating files one by one, the platform reasons about architecture, dependencies, compatibility, performance implications, unsupported APIs, testing strategy, documentation, and migration risks.

Every decision is transparent. Every modification is explainable. Every migration produces evidence.


🔍 Repository Intelligence

Kernel Olympics doesn't process kernels independently. Before modifying a single line, the platform analyzes the entire project context to identify:

  • CUDA kernels and their call sites
  • Runtime API usage and device memory operations
  • Build configuration and dependencies
  • Shared utilities and include hierarchy
  • Unsupported CUDA features and migration complexity

Instead of blindly translating files, the platform understands the codebase as a whole — which dramatically improves migration quality.


🧠 The Problem

Today, migrating GPU software is difficult because developers must manually:

  • Understand large CUDA codebases
  • Identify unsupported APIs
  • Rewrite kernels
  • Replace memory management
  • Verify correctness
  • Debug compilation failures
  • Benchmark performance
  • Write migration documentation

Even experienced GPU developers spend days or weeks doing this. The process is repetitive, error-prone, expensive, and hard to scale.


💡 Our Solution

Kernel Olympics introduces an AI-native migration workflow. Instead of treating migration as file conversion, the platform treats it as an engineering reasoning problem. The system first understands the repository, identifies architectural patterns, analyzes dependencies, detects unsupported CUDA features, proposes migration strategies, executes intelligent transformations, validates results, benchmarks performance, and generates a comprehensive migration report explaining every decision.

This creates a migration pipeline that is explainable, repeatable, and significantly easier for developers to trust.


🎬 Demo

🐙 Live: kernel-olympics-production.up.railway.app — Upload a CUDA kernel and watch the pipeline port it live.

Kernel Olympics Pipeline Demo
📡 Real AMD MI300X via notebooks.amd.com/hackathon — CUDA source → 4-LLM loop ports → hipcc compile → GPU run → PASSED ✓


📈 Quick Stats

Metric Value
Pipeline budget 1,800s (30 min)
Max iterations 10 (compile-fix loop)
LLM cost per run ~$0.09
Cache hit speed ~0.2ms
Tests 665 passing
CI/CD ✅ Automated (GitHub Actions)
Hardware target AMD MI300X via notebooks.amd.com/hackathon (192GB HBM3, CDNA3)

🧠 Multi-Agent Architecture

                    ┌─────────────────────────────────┐
                    │      CUDA Kernel (.cu)           │
                    └──────────────┬──────────────────┘
                                   │
                                   ▼
                    ┌─────────────────────────────┐
                    │  1. Risk Classifier          │
                    │  (RED/YELLOW/GREEN)          │
                    └──────────────┬──────────────┘
                                   │
                                   ▼
                    ┌─────────────────────────────┐
                    │  2. Pattern Memory Cache     │
                    │  (trigram ~0.2ms, 60,000×)   │
                    └──────────────┬──────────────┘
                                   │
                    ╔══════════════╧══════════════╗
                    ║      LLM Agent Loop         ║
                    ║  (full auto, no human)      ║
                    ║                              ║
                    ║  3. DeepSeek-v4-Pro          ║
                    ║     → architecture plan      ║
                    ║                              ║
                    ║  4. GLM-5.2                  ║
                    ║     → HIP kernel code gen    ║
                    ║                              ║
                    ║  5. Kimi K2.7                ║
                    ║     → 3-gate eval + refine   ║
                    ║       (compile-fix loop)     ║
                    ║                              ║
                    ║  6. Gemma 4 / DeepSeek-v4    ║
                    ║     → final verification     ║
                    ╚══════════════╧══════════════╝
                                   │
                                   ▼
                    ┌─────────────────────────────┐
                    │  7. REAL AMD GPU            │
                    │  hipcc + run + numerical    │
                    │  diff verification           │
                    └──────────────┬──────────────┘
                                   │
                                   ▼
                    ┌─────────────────────────────┐
                    │   ✅ HIP Kernel + Proof     │
                    └─────────────────────────────┘

🔬 Explainable AI

Every modification includes reasoning. Nothing is hidden. Every decision is transparent.

Original CUDA API
    ↓
Reason for replacement
    ↓
New ROCm implementation
    ↓
Performance implications
    ↓
Documentation
    ↓
Confidence Score

Instead of asking "What changed?", developers see why every change was made — making migration reviewable, auditable, and trustworthy.


🔬 Key Innovation: Pattern Memory Cache

Instead of calling expensive LLMs for every kernel, we cache porting patterns as trigram vectors:

Metric Without Cache With Cache Speedup
Pattern lookup N/A 0.2ms —
LLM call (simulated) ~12s 0.2ms 60,000×
Verified with live API — ✓ measured ✓

🗺️ Future Roadmap

The long-term vision extends far beyond CUDA → ROCm. Future versions can support:

  • CUDA → ROCm (✅ currently supported)
  • CUDA → SYCL
  • CUDA → Vulkan Compute
  • CUDA → OpenCL
  • CUDA → Metal
  • CUDA → DirectML

The platform is designed as a universal GPU migration engine rather than a single-purpose converter. Adding new target paths is a configuration change, not a rewrite.


🚀 Quick Start

Live Demo (no install required)

→ kernel-olympics-production.up.railway.app — Upload a .cu file, see the autonomous pipeline port it to HIP in real time.

Local Setup

git clone https://github.com/indrad3v4/Kernel-Olympics.git
cd Kernel-Olympics
pip install -r requirements.txt

# Run the full pipeline on a sample kernel
make port CU_FILE=sample_kernels/cuda/warp_reduce.cu

# Or try a more complex warp-level scan kernel
make port CU_FILE=sample_kernels/cuda/nvidia_shfl_scan.cu

🏁 For Judges: Running on AMD GPU

Full guide: AMD_STACK_USAGE.md — but here's the 30-second version:

Prerequisites

  • AMD GPU (MI300X, MI250, RX 7900 XTX) with ROCm 6+ installed
  • No AMD hardware? Use the AMD AI Notebooks portal — free MI300X instances for hackathon participants
  • hipcc in PATH (hipcc --version)
  • Python 3.11+, git

Quick Run

# 1. Clone
git clone https://github.com/indrad3v4/Kernel-Olympics.git
cd Kernel-Olympics

# 2. Install deps
pip install -r requirements.txt

# 3. Set your API key (any OpenAI-compatible LLM)
export FIREWORKS_API_KEY="your-key"

# 4. Run the full pipeline on a real kernel (port → compile → run → verify)
make pipeline CU_FILE=sample_kernels/cuda/warp_reduce.cu

# Or with extended budget (30 min, 10 iterations)
make pipeline-heavy CU_FILE=sample_kernels/cuda/nvidia_shfl_scan.cu

What to Expect

Stage Time Output
Port (4-LLM loop) 1–5 min HIP code in ported_kernels/
Compile (hipcc) 2–5 sec Binary in build/
Run on AMD GPU 1–3 sec Numerical output
Verify 0.5 sec PASS ✓ or FAIL

A successful run ends with:

✅ Pipeline complete — RESULT: PASS

Without an AMD GPU

The pipeline still ports + verifies. Just skip compile/run:

make port CU_FILE=sample_kernels/cuda/warp_reduce.cu
make inspect CU_FILE=sample_kernels/cuda/warp_reduce.cu  # view ported code

Verify the Setup

make doctor           # check ROCm, hipcc, API key
make test             # 665 tests

📋 Makefile Targets

Target Description
help Show all available targets
install Create venv + install deps
port Pipeline on one kernel: make port CU_FILE=path.cu
port-all Pipeline on ALL sample kernels
compile hipcc proof harness + compile
run Run compiled binary on AMD GPU
pipeline Full cycle: port → compile → run
pipeline-heavy Extended budget (1,800s)
test Run 665 pytest tests
demo Live demo with recording
inspect Inspect spec/ported kernel/proof
debug-kernel Interactive kernel explorer
retry Re-run a single pipeline stage

🐞 Debug Mode

Three levels of debugging for when things go wrong:

# Inspect specs and artifacts
make inspect CU_FILE=sample_kernels/cuda/warp_reduce.cu
make inspect PORTED=ported_kernels/warp_reduce.hip.cpp

# Interactive kernel exploration
make debug-kernel CU_FILE=sample_kernels/cuda/warp_reduce.cu

# Retry a single stage
make retry CU_FILE=sample_kernels/cuda/warp_reduce.cu STAGE=port

🧪 Running Tests

make test              # full suite
make test-verbose      # with progress

🔧 Pipeline Architecture Details

CUDA → HIP Transformations

CUDA Intrinsic HIP Equivalent Action
__shfl_up_sync(mask, val, d, w) __shfl_up(val, d, width) Mask dropped
cudaMalloc() hipMalloc() 1:1 rename
cudaMemcpy() hipMemcpy() 1:1 rename
findCudaDevice() hipGetDevice() SDK strip
sdkCreateTimer() — Removed (NOP)
threadIdx.x hipThreadIdx.x Namespace add

Known Issues Handled

  • NVIDIA SDK symbols — auto-detected and stripped by verifier.py
  • Wave64 divergence — warpSize constant used instead of hardcoded 64
  • SIGSEGV from host-code symbols — sanitizer in verifier catches at compile time

👥 Team

Role Member Focus
🚀 Lead indradev_ Architecture, pipeline orchestrator
🔬 Kernel & Docs Bromine185 CUDA kernel analysis, warp primitives, documentation, demo recording
🔧 Infra & Verification icodemun44 CI/CD, AMD cloud, Jupyter integration, verifier core, semantic repair engine, static analysis, tooling
🧪 CI meteorite67 GitHub Actions, test suite
✅ Verification paparehan Structural/lexical validation gates, safe writer, typed diagnostics
🏭 AMD Aahil-Riyaz (Satoru) AMD MI300X testing, ROCm debugging, CI/CD

📄 License

MIT


Built for the AMD Developer Hackathon ACT II · Track 3 (Open Innovation)
lablab.ai


built by indradev_ · ☕ support

About

CUDA→ROCm Migration Copilot

Resources

Contributing

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages