Siddharth Gururani

Siddharth Gururani

Senior Research Scientist, NVIDIA

I work on world foundation models and generative AI at NVIDIA — building models that understand and simulate the physical world, alongside generative systems for audio, speech, and music. Most recently I led joint audio-visual generation for Cosmos 3, NVIDIA's open omnimodal world foundation model for physical AI.

About

I am a Senior Research Scientist in the Cosmos Lab at NVIDIA. My current work spans world foundation models for physical AI, generative audio, speech and singing synthesis, video-to-audio generation, and multimodal understanding. Most recently, I was the tech lead for joint audio-visual generation in Cosmos 3, NVIDIA's open omnimodal world foundation model that jointly reasons over and generates language, image, video, audio, and action.

Before NVIDIA, I was an AI Scientist at Electronic Arts working on expressive speech synthesis. I received my Ph.D. in Music Technology from Georgia Tech, where I worked on weakly supervised methods for identifying musical instruments in audio. I also hold B.Tech. and M.Tech. degrees in Computer Science and Engineering from IIT Kharagpur.

My path into machine learning began with music — I still play guitar, and that interest in understanding and creating sound runs through much of my research.

Publications

Technical Report 2026

Cosmos 3: An Open Omnimodal World Foundation Model for Physical AI

NVIDIA — Siddharth Gururani and the Cosmos team · Tech Lead (Audio)

An open frontier world foundation model that unifies physical reasoning, world generation, and action generation across language, image, video, audio, and action, via a two-tower mixture-of-transformers architecture.

CVPR 2026 2026

Benchmarking Single-Factor Physical Video-to-Audio Generation

Siddharth Gururani et al. — Tingle Li, Kevin J. Shih, Gantavya Bhatt, Sang-gil Lee, Zhifeng Kong, Arushi Goel, Gopala Anumanchipalli, Ming-Yu Liu

FlatSounds tests whether video-to-audio systems capture physical factors and timing rather than relying only on captions or surface realism.

arXiv 2026

Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Zhifeng Kong, Siddharth Gururani, Sang-gil Lee, Jaehyeon Kim, Aya Aljafari, Chao-Han Huck Yang, Sungwon Kim, Ramani Duraiswami, Dinesh Manocha, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping

An open audio-language model family for long-form understanding and reasoning over speech, sound, and music.

arXiv 2026

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

Arushi Goel, Sreyan Ghosh, Vatsal Agarwal, Nishit Anand, Kaousheik Jayakumar, Lasha Koroshinadze, Yao Xu, Katie Lyons, James Case, Karan Sapra, Kevin J. Shih, Siddharth Gururani, Abhinav Shrivastava, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Bryan Catanzaro, Mohammad Shoeybi, Wei Ping

A benchmark for evaluating models that reason jointly over visual, audio, and textual signals in long, real-world videos.

Technical Report 2025

Cosmos-Predict2.5: A Flow-Based World Foundation Model for Physical AI

NVIDIA — Siddharth Gururani and the Cosmos team · Core Contributor

A flow-based world foundation model for physical AI, enabling controllable world simulation and video generation for downstream training.

ICLR 2025 2025

Fugatto 1: Foundational Generative Audio Transformer Opus 1

Rafael Valle, Rohan Badlani, Zhifeng Kong, Sang-gil Lee, Arushi Goel, Sungwon Kim, Joao Santos, Shuqi Dai, Siddharth Gururani, Aya Aljafari, Alexander Liu, Kevin Shih, Ryan Prenger, Wei Ping, Chao-Han Huck Yang, Bryan Catanzaro

A foundational generative audio model for flexible text, audio, and music generation tasks.

Technical Report 2025

Cosmos-Reason1: From Physical Common Sense to Embodied Reasoning

NVIDIA — Siddharth Gururani and the Cosmos team · Core Contributor

Multimodal reasoning models and benchmarks for physical common sense and embodied decision-making.

Technical Report 2025

Cosmos-Predict1: Cosmos World Foundation Model Platform for Physical AI

NVIDIA — Siddharth Gururani and the Cosmos team · Core Contributor

A platform of world foundation models, tokenizers, data pipelines, and post-training tools for physical AI.

Technical Report 2024

Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

NVIDIA — Yuval Atzmon, Maciej Bala, Yogesh Balaji, Tiffany Cai, Yin Cui, Jiaojiao Fan, Yunhao Ge, Siddharth Gururani, Jacob Huffman, Ronald Isaac, Pooya Jannaty, Tero Karras, Grace Lam, and others

A pixel-space diffusion family for high-resolution image generation and controllable visual content creation.

ACM Multimedia 2024 2024

ExpressiveSinger: Multilingual and Multi-Style Score-Based Singing Voice Synthesis with Expressive Performance Control

Shuqi Dai, Ming-Yu Liu, Rafael Valle, Siddharth Gururani

A singing voice synthesis system for multilingual and multi-style score-based generation with explicit performance controls.

ICML 2024 2024

Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion

Yujia Huang, Adishree Ghatare, Yuanzhe Liu, Ziniu Hu, Qinsheng Zhang, Chandramouli S. Sastry, Siddharth Gururani, Sageev Oore, Yisong Yue

Diffusion-based symbolic music generation guided by musical rules that are useful but not differentiable.

ICCV 2023 2023

SPACE: Speech-Driven Portrait Animation with Controllable Expression

Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, Ming-Yu Liu

A speech-driven portrait animation system with controllable pose, emotion, and expression intensity.

INTERSPEECH 2023 2023

RAD-MMM / Multilingual Multiaccented Multispeaker TTS with RADTTS

Rohan Badlani, Rafael Valle, Kevin J. Shih, Joao Felipe Santos, Siddharth Gururani, Bryan Catanzaro

A multilingual and multiaccented text-to-speech system with explicit controls over speaker, accent, language, pitch, and energy.

arXiv 2021

An Interdisciplinary Review of Music Performance Analysis

Alexander Lerch, Claire Arthur, Ashis Pati, Siddharth Gururani

A survey of music performance analysis across measurement, performer intent, listener perception, and MIR.

arXiv 2019

Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features

Siddharth Gururani, Kilol Gupta, Dhaval Shah, Zahra Shakeri, Jervis Pinto

An approach to text-to-speech prosody transfer using interpretable pitch and loudness features.

ISMIR 2019 2019

An Attention Mechanism for Musical Instrument Recognition

Siddharth Gururani, Mohit Sharma, Alexander Lerch

Weakly supervised multi-label instrument recognition using attention to identify relevant regions in audio.

ISMIR 2018 2018

Instrument Activity Detection in Polyphonic Music Using Deep Neural Networks

Siddharth Gururani, Cameron Summers, Alexander Lerch

A neural approach to detecting regions of activity for multiple instruments in polyphonic music.

arXiv 2022

Anomalous Behaviour in Loss-Gradient Based Interpretability Methods

Vinod Subramanian, Siddharth Gururani, Emmanouil Benetos, Mark B. Sandler

Analyzes failure modes of loss-gradient saliency methods used to interpret audio classification models.

IEEE ISM 2021 2021

Semi-Supervised Audio Classification with Partially Labeled Data

Siddharth Gururani, Alexander Lerch

Semi-supervised methods for multi-label audio classification when only part of the labels are available.

ISMIR 2020 2020

dMelodies: A Music Dataset for Disentanglement Learning

Ashis Pati, Siddharth Gururani, Alexander Lerch

A dataset of simple melodies with known factors of variation for benchmarking disentanglement learning.

ISMIR 2020 2020

Score-Informed Networks for Music Performance Assessment

Jiawen Huang, Yun-Ning Hung, Ashis Pati, Siddharth Gururani, Alexander Lerch

Neural models that use score information to assess the quality of student music performances.

arXiv 2020

Visual Attention for Musical Instrument Recognition

Karn N. Watcharasupat, Siddharth Gururani, Alexander Lerch

Studies attention over time-frequency representations for multi-label instrument recognition.

ISMIR 2017 2017

Automatic Sample Detection in Polyphonic Music

Siddharth Gururani, Alexander Lerch

Detects sampled audio segments that are reused across polyphonic music tracks.

AES Semantic Audio 2017 2017

Objective Descriptors for the Assessment of Student Music Performances

Amruta Vidwans, Siddharth Gururani, Chih-Wei Wu, Vinod Subramanian, Rupak Swaminathan, Alexander Lerch

Audio features for automatically assessing the quality of student instrumental performances.

ISMIR 2016 2016

Automatic Practice Logging: Introduction, Dataset & Preliminary Study

R. Michael Winters, Siddharth Gururani, Alexander Lerch

A dataset and preliminary methods for automatically logging instrumental practice sessions.