Profile

Computer vision and multimodal AI, with an end-to-end engineering range.

I work across multimodal generative models, physical AI, and geometric computer vision. My role typically spans research design, data and training systems, evaluation, optimization, and product integration. I am most interested in spatial systems that must remain reliable under real data, hardware, and infrastructure constraints.

Current focus Multimodal post-training, real-to-sim, digital twins, and robot calibration.

Research foundation Robust and efficient 3D reconstruction, neural rendering, and multiview geometry.

Selected work

Representative contributions

Luma AI 2025—Present

Multimodal models and physical AI

  • Core contributor to Uni-1; drove post-training data and recipes for multiview rendering, spatial control, and image editing.
  • Built a Ray, Qwen-VL, and vLLM curation system that filtered 97 million candidates into 16 million high-quality SFT rows.
  • Independently owned embedding-based corpus normalization and representation-similarity tooling.
  • Research high-fidelity real-to-sim and calibration across hand-eye, robot-world, and joint-encoder bias.
Uni-1 technical specifications ↗

Meta Reality Labs 2024—2025

Egocentric 4D vision and reconstruction at scale

  • First author of Ego-1K, a CVPR 2026 benchmark with 956 videos and 514,000 frames from a synchronized 12+4-camera rig.
  • Initiated Gaussian-splatting reconstruction across thousands of indoor LiDAR and RGB captures for sensor simulation.
  • Improved dynamic multiview depth in throughput, online calibration, fusion accuracy, and completeness.
  • Led creation of approximately 300,000 photorealistic 3D-LLM post-training examples and designed automated geometric quality metrics.

Apple 2024

On-device vision ML for Persona

  • Directly responsible for Persona enrollment relighting and light normalization on Apple Vision Pro.
  • Owned data, preprocessing, architecture, training, evaluation, quantization analysis, runtime optimization, and integration.
  • Shipped the component for visionOS under on-device memory and inference-latency constraints.

Reconstruct Inc. 2015—2023 · Intermittent

3D computer vision systems and product engineering

  • Built production foundations across AWS, authentication, large-file workflows, load-balanced services, databases, and Docker.
  • Developed interactive WebGL tools for point clouds, BIM, meshes, 360 imagery, measurement, annotation, and capture navigation.
  • Built and integrated SfM, bundle adjustment, multiview stereo, semantic mapping, and neural-rendering pipelines.
  • Co-invented a productized multiview measurement method granted as U.S. Patent 12,597,201.

Technical scope

Methods and systems

Multimodal ML
Post-training data, VLM curation, diffusion and flow matching, image and video generation, editing, spatial control, evaluation, and ablation design.
Geometry
Robot and camera calibration, real-to-sim, SfM, bundle adjustment, MVS, stereo, RGB-D, SLAM, NeRF, Gaussian Splatting, and 3D/4D reconstruction.
ML systems
PyTorch, Ray, vLLM, DDP, FSDP, tensor and context parallelism, quantization, runtime profiling, and distributed model serving.
Engineering
Python, C/C++, CUDA, OpenGL/WebGL, AWS, GCP, Docker, backend systems, databases, and interactive 3D applications.

Research

Selected publications

Representative first-authored and collaborative work. Complete list on Google Scholar ↗

Loading publications…

Background

Experience and education

Dec. 2025—Present

Luma AI

Research Scientist / Engineer

Nov. 2024—Dec. 2025

Meta Reality Labs

Computer Vision Engineer

Feb.—Nov. 2024

Apple

Machine Learning Engineer, Video Computer Vision

2018—2024

University of Illinois Urbana-Champaign

Ph.D. in Computer Science · Advisor: Professor Derek Hoiem Thesis: Toward Robust and Efficient 3D Reconstruction

2019—2022 · summers

Research internships

Meta Reality Labs Research Amazon Go Microsoft HoloLens & Research

Additional

Patents and recognition

  • 2026

    U.S. Patent 12,597,201 — Interactive measurements using co-registered images and 3D points.

  • 2024

    U.S. Patent 12,154,312 — Pixel correspondence via patch-based neighborhood consensus.

  • 2022

    CVPR Best Paper Finalist and Oral presentation for DIVeR.