PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
A minimalist pixel-space diffusion transformer that predicts dense 3D point maps from a single image, with no VAE and no hybrid architecture.
Researcher at KE:SAI Open Science Lab
Postdoc at ETH Zurich
I am a researcher at KE:SAI Open Science Lab and a postdoc at ETH Zurich, advised by Andreas Geiger and Marc Pollefeys.
I received my PhD from ETH Zurich under the supervision of Marc Pollefeys and Andreas Geiger. During my PhD, I interned at Google, where I worked with Michael Niemeyer and Federico Tombari. Previously, I worked with Jianfei Cai and Hamid Rezatofighi at Monash University. I earned my master's degree from University of Science and Technology of China, where I worked with Juyong Zhang. During my master's, I also studied as an exchange student at NTU and completed a research internship at Microsoft Research Asia.
I am honored to have been named a 2025 Apple Scholar in AI/ML and to have received the Gold Reviewer Award at ICML 2026, the Top Reviewer Award at NeurIPS 2024, and the Outstanding Reviewer Award at CVPR 2022.
I have broad interests in computer vision and deep learning. I have worked on depth, stereo, optical flow, tracking, feed-forward NeRF/3DGS, and generative models. I am actively exploring new ideas, and I am currently interested in representation learning and generative models for efficient and scalable world models and physical AI. Please see my full publication list on Google Scholar.
A minimalist pixel-space diffusion transformer that predicts dense 3D point maps from a single image, with no VAE and no hybrid architecture.
Test-time scaling for feed-forward Gaussian splatting.
A simple yet powerful framework for large-displacement optical flow and point tracking.
The first data-driven multi-view 3D point tracker for tracking arbitrary 3D points across multiple cameras.
Cross-task interactions between feed-forward Gaussian splatting and depth.
4K panorama synthesis with a single feed-forward inference.
Unposed 3DGS reconstruction made easy.
A cost volume representation for efficiently predicting 3D Gaussians from sparse multi-view images in a single feed-forward inference.
A unified dense correspondence matching formulation enables three motion and 3D perception tasks to be solved with a unified model.
Factorizing 2D optical flow with 1D attention and 1D correlation enables 4K resolution optical flow estimation on standard GPUs.
A sparse points-based cost aggregation method leads to an efficient and accurate stereo matching architecture without any 3D convolutions.
Key open-source projects developed through my research and collaborations.