Computer vision and multimodal AI, with an end-to-end engineering range.
I work across multimodal generative models, physical AI, and
geometric computer vision. My role typically spans research design,
data and training systems, evaluation, optimization, and product
integration. I am most interested in spatial systems that must
remain reliable under real data, hardware, and infrastructure
constraints.
Current focus Multimodal post-training, real-to-sim, digital twins, and robot calibration.
Research foundation Robust and efficient 3D reconstruction, neural rendering, and multiview geometry.
Selected work
Representative contributions
Luma AI2025—Present
Multimodal models and physical AI
Core contributor to Uni-1; drove post-training data and recipes for multiview rendering, spatial control, and image editing.
Built a Ray, Qwen-VL, and vLLM curation system that filtered 97 million candidates into 16 million high-quality SFT rows.
Independently owned embedding-based corpus normalization and representation-similarity tooling.
Research high-fidelity real-to-sim and calibration across hand-eye, robot-world, and joint-encoder bias.