Ze Wang
Research Scientist at Luma AI
I work on frontier multimodal generative foundation models — joint video, audio, and synchronized audiovisual generation at scale. At Luma, I am a primary contributor to Ray 3.5, Luma’s first joint audio‑video generation model, and to the unified Omni foundation model family.
Previously, I was a Research Scientist at AMD GenAI, where I led the development of vision generation and understanding models in the fully open-source Instella model family, including Instella‑T2I. Before that, I was a Postdoctoral Research Associate at Purdue University. I received my Ph.D. from Purdue in 2023, advised by Prof. Qiang Qiu and Prof. Guillermo Sapiro (starting at Duke University), and my B.E. from Beihang University in 2017.
My research interests include diffusion, autoregressive, and hybrid models for efficient image, video, and audio generation; multimodal large language models for unified understanding and generation; and post-training of foundation models.
Experience
-
2026 – now
Luma AI · Research Scientist
Joint audio‑video generation (Ray 3.5), the Omni foundation model family, and Luma’s audio and audiovisual modeling stack.
-
2024 – 2026
AMD · Research Scientist, GenAI
Led Instella‑T2I (text‑to‑image), Instella‑Nexus (unified multimodal), and Instella‑T2V (text‑to‑video) in the open-source Instella model family.
-
2023 – 2024
Purdue University · Postdoctoral Research Associate
Efficient diffusion models, generative model evaluation, and adaptation of vision and language foundation models.
-
earlier
Research internships
Microsoft Azure AI (2022–2024), Meta AI (2021), Volvo AI Research (2020), Horizon Robotics (2017).
Selected Publications
-
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Representation and Generation
preprint 2025 code
-
CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models
preprint 2026
-
Instella: Fully Open Language Models with Stellar Performance
preprint 2025 code
-
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
CVPR 2025
-
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
ICML 2025
-
Binary Latent Diffusion
CVPR 2023 code
-
Energy-Inspired Self-Supervised Pretraining for Vision Models
ICLR 2023 Spotlight
-
Few-Shot Fast-Adaptive Anomaly Detection
NeurIPS 2022
-
Continual Learning with Filter Atom Swapping
ICLR 2022 Spotlight
-
Image Generation Using Continuous Filter Atoms
NeurIPS 2021 Spotlight
-
Adaptive Convolutions with Per-Pixel Dynamic Filter Atom
ICCV 2021 code
-
Range Adaptation for 3D Object Detection in LiDAR
ICCV-W 2019 Best Paper Award
Full list on Google Scholar.