AI Inference Engineer at Perplexity in London, working on serving large language models fast and efficiently. MEng Information & Computer Engineering, University of Cambridge.
I work across ML systems and research: LLM inference, model merging and parameter-efficient fine-tuning, evaluating LLM research agents, and scientific machine learning. Currently building a from-scratch LLM inference engine.
- Model merging in Speech LLMs: MEng thesis. LoRA adapters in a 3B Speech LLM quietly break each other's tasks; calibrated layer-wise merging recovers single-task performance on average across 7 speech tasks without joint retraining, using 16× less training data than multi-task training. Thesis PDF
- sciagent: a benchmark for LLM research agents in which ground truth is a structural edit to an executable programme rather than a label. It is bit-exact deterministic and content-addressed, with over 1,000 tests and
mypy --strictthroughout. - Recurrent neural operator for visco-plasticity: learns a history-dependent constitutive law from unit-cell simulations, reaching test R² 0.996 with 16.6k parameters.
- Physics-informed networks and neural operators: PINNs for 2D elasticity and a Fourier Neural Operator for Darcy flow, with ~30× lower test error than a CNN baseline.
- Interior-point solver: Newton, KKT, log-barrier and Phase I written from first principles in NumPy, then used to train hard- and soft-margin SVMs.
- Cart-pole control from learned dynamics: kernel-regression dynamics and gradient-based policy search through a differentiable JAX simulator.