
VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
Hyojun Go, Dominik Narnhofer, Goutam Bhat, Prune Truong, Federico Tombari, Konrad Schindler
International Conference on Learning Representations (ICLR) 2026 Oral
Unified a pretrained video diffusion model and a feed-forward 3D reconstruction model into a single end-to-end latent diffusion model that directly generates 3D worlds from text.





















