At SIGGRAPH 2025, NVIDIA unveiled a range of software and research innovations supporting the development of physical AI, a field that combines neural graphics, synthetic data generation, physics-based simulation, reinforcement learning and reasoning capabilities. The company says these technologies can support robotics, autonomous vehicles, smart spaces and content creation applications.
SIGGRAPH is being held in Vancouver through August 14, where NVIDIA Research leaders are delivering a special presentation on graphics and simulation innovations that enable physical and spatial AI. Sanja Fidler, vice president of AI research at NVIDIA, said that AI is advancing simulation capabilities, while simulation is advancing AI systems, describing the relationship between the two fields as a powerful and distinctive coupling.
Libraries and Models for Developing Physical AI
At the conference, NVIDIA announced new software libraries for physical AI, including the NVIDIA Omniverse NuRec libraries for large-scale world reconstruction using 3D Gaussian splatting technology. The company also showcased updates to the NVIDIA Metropolis platform for vision intelligence, along with NVIDIA Cosmos and NVIDIA Nemotron, two reasoning models.
Among the announcements was Cosmos Reason, a new vision-language reasoning model for physical AI. According to NVIDIA, the model enables robots and vision intelligence agents to reason in a human-like manner by relying on prior knowledge and an understanding of physics and common sense.
Building Virtual Worlds for Training
NVIDIA believes that developing physical AI begins with creating high-fidelity 3D environments that adhere to the laws of physics. These environments enable advanced systems, such as humanoid robots, to be trained safely in simulation while improving the chances that acquired skills will transfer to the real world.
Ming-Yu Liu, vice president of research at NVIDIA, said that these virtual worlds require real-time rendering, computer vision, physical motion simulation, 2D and 3D generative AI, as well as reasoning capabilities. She added that NVIDIA Research has been working in these areas for nearly two decades.
The team’s work includes neural reconstruction and rendering, which use AI to transform data captured from cameras or sensors into realistic 3D representations. Fidler’s Spatial Intelligence Lab also presented ViPE, or Video Pose Engine, a pipeline for geometric 3D annotation of videos developed in collaboration with the Dynamic Vision Lab and the NVIDIA Isaac team. The tool estimates camera motion and creates detailed depth maps from amateur recordings, car cameras and cinematic footage.
More Than 12 Research Papers at SIGGRAPH
NVIDIA Research is presenting more than 12 research papers at the conference on neural rendering, real-time ray tracing, synthetic data generation and reinforcement learning. One paper addresses physics-aware 3D geometry reconstruction from images or video, aiming to avoid producing shapes that appear visually correct but lack structural stability within simulation.
Another paper presents a method for imparting physically accurate motion to simulated characters by combining a motion generator with a physics-based tracking controller. The resulting data can be used to develop virtual characters or train humanoid robots in complex movements, such as parkour and traversing difficult terrain.
Other research addresses light and material simulation. The NVIDIA team presented an AI assistant that uses diffusion models and a physics-based differentiable renderer to modify material texture maps through text commands, helping add details such as weathering or aging effects to 3D models. Another paper explores a differentiable vision query aimed at accelerating and improving the accuracy of 3D geometry reconstruction from images and video clips.
These works connect forward rendering, which converts 3D models into 2D images, with inverse rendering, which extracts 3D models from images. NVIDIA sees this integration as providing essential components for building virtual worlds and synthetic data that can be used to train physical AI systems.