I build low-latency AI inference systems for realtime speech and voice products.
Current: Member of Technical Staff at Stellon Labs, focused on Text to Speech and Speech to text model serving, GPU benchmarking, native inference packaging, and production deployment.
- Realtime ASR/STT/TTS serving
- GPU inference on various data center GPUs
- vLLM, Modal, RunPod, Vast, Replicate
- C++/CUDA/Metal/ONNX Runtime inference
- Android/Termux, macOS Metal, Windows ARM64/x86, Linux ARM64/manylinux
- Benchmarking: TTFB, RTF, route time, lock wait, RSS/VRAM, CPU perf, energy/power
Professional account. Personal projects: @Haaziq386
