On May 31, 2026, NVIDIA announced Cosmos 3, which it describes as an open world foundation model for Physical AI and its first fully open omnimodel, able to understand and generate several data types at once. The model uses a mixture-of-transformers architecture that pairs a reasoning transformer with an expert generation transformer, so it can reason about physical interactions before predicting what happens next. According to NVIDIA it was trained on one of the largest multimodal physical-AI datasets, spanning billions of samples across text, image, video, sound and action trajectories, and it can act as a vision-language model, a world model, a video foundation model and a backbone for teaching robots specific tasks. NVIDIA also introduced Super, Nano and (coming soon) Edge variants, alongside a Cosmos Coalition of partners including Agile Robots, Black Forest Labs, Generalist, LTX, Runway and Skild AI.
For Ergonet, this reflects the same direction we are building toward: machine vision and world models that let robots and automation systems understand physical environments before acting in them.
Source: NVIDIA