"Bridging the gap between bits and atoms."
Undergraduate in Computer Science at The University of Hong Kong and Vision Group member of the HKU RoboMaster Team. I work on real-world perception systems, vision-language reasoning, autonomous agents, and embodied AI.
My current interests sit at the intersection of:
- Vision-language models: spatial reasoning, visual grounding, and geometry-aware multimodal inference.
- Computer vision: detection, segmentation, depth estimation, calibration, and robust state estimation.
- Autonomous systems: perception-to-action pipelines, closed-loop control, and real-time deployment.
|
I am interested in how LLMs and VLMs reason, plan, use tools, and maintain context over long-horizon tasks.
|
I study whether explicit visual geometry can make VLM reasoning more reliable on real images.
|
I build and optimize perception pipelines for dynamic physical environments.
|
|
Geometry-aware visual reasoning pipeline for evaluating whether object detection, SAM2 segmentation, Depth Anything V2 relative depth, and object-level geometry can improve VLM spatial reasoning. The project builds a staged benchmark pipeline: Current features include YOLO detection, label normalization, SAM2 mask generation, depth estimation, object-level geometry extraction, reasoning prompt generation, a geometry-only rule baseline, and failure analysis that separates upstream perception errors from reasoning errors. |
Real-time deep learning aim bot for Aimlab Sixshot. The system uses screen capture, target detection, and hardware mouse control at 100+ FPS. Built a custom MiniUNet with approximately 467K parameters for Gaussian heatmap regression, plus an end-to-end workflow from interactive labeling to real-time inference and adaptive flick control. |
|
PyTorch · OpenCV · CUDA |
Python · C++ · C# · Java · JavaScript |
|
Linux · ROS · Docker · Git · MySQL · LaTeX |
React · Vue · TypeScript · Unity |


