Senior Staff Engineer | AI Infrastructure · GPU Platforms · Kubernetes · Cloud · Linux Systems
Building AI infrastructure from the kernel to the cluster — and from the cluster to the model.
I build GPU platforms and production ML infrastructure across Linux, Kubernetes, networking, and storage.
- GPU Platforms — NVIDIA GPUs, CUDA, DeepStream, DCGM, training and inference
- Cluster Networking — InfiniBand, RDMA/RoCE, Clos fabrics, BGP/ECMP, NetBox
- Kubernetes — GPU workloads, CNI, Helm, GitOps
- Linux Reliability — eBPF, cgroups, OOM analysis, kernel diagnostics
- MLOps — model lifecycle, evaluation, lineage, serving, observability
- Automation — AWS, Terraform, Ansible, GitHub Actions, CI/CD
Currently exploring distributed GPU communication, AI storage paths, and GPU/Linux performance engineering.
| Area | Project | Description |
|---|---|---|
| AI Factory | ai-data-center-systems | GPU, networking, storage, training, and inference reference |
| AI Factory | ai-factory-network-twin | NetBox-driven network digital twin for validating BGP, ECMP, isolation, and failure recovery with Containerlab |
| AI Industry | capex-lens | Dashboard tracking the AI infrastructure investment cycle through market indicators |
| GPU | gds-lab | Storage-to-GPU data paths and GPUDirect Storage experiments |
| GPU | ghostmem | eBPF diagnostics for shared-memory and CUDA-related OOM pressure |
| GPU | ansible-role-nvidia-container-toolkit | Automated NVIDIA Container Toolkit provisioning |
| DevOps | crashshoot | Containerized Linux kernel debugging and vmcore postmortem analysis |
| DevOps | tail-lifter | eBPF connectivity from Tailscale clients to Kubernetes services |
| DevOps | kernel-lens | Evidence-linked AI summaries of Linux kernel development |
| DevOps | linux-boot-optimization-lab | Boot and latency analysis for embedded Linux and physical AI |
| Data Pipelines | drizzler | Adaptive data collection and LLM summarization with API and Helm deployment |
| MLOps | model-port-gateway | Model intake, fine-tuning, validation, and rollout reference |
| MLOps | vision-mlops | Vision lifecycle reference from annotation to ONNX deployment |
| MLOps | nimbus | GPU-aware orchestration and rollout for edge AI workloads |
| MLOps | vrs | Video reasoning with perception, VLM verification, and Kubernetes |
| MLOps | cp-fp-mining-lab | Reviewed false-positive mining and detector retraining |
| MLOps | ml-helmfiles | Kubernetes training and Triton inference with Helmfile |
| LLMOps | rag-adapt-lab | Reproducible RAG, SFT, and RAFT evaluation |
| LLMOps | llm-serving-lab | GPU benchmarks and observability for LLM serving engines |
| LLMOps | FlareGraph | Cloudflare-native knowledge graphs and MCP retrieval |
| LLMOps | skills | Reusable agent skills for engineering workflows |
| LLM Applications | jev-iab-explorer | Video classification with resumable processing and provider safeguards |
| Personal | akbo | Offline-ready piano sheet music viewer for iPad and desktop |
| Personal | speak-loop | English speaking coach for engineers with real-time voice practice, feedback, and spaced review |





