NVIDIA explains omni-models: unified text, image, audio, video

NVIDIA Cosmos Lab VP Ming-Yu Liu explains omni-models, a unified architecture that processes text, images, audio, video, and actions together. The video walks through how a single model handles inputs across modalities instead of separate specialist models.
Featured · Ming-Yu Liu
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- NVIDIA demos local Hermes agent debugging with RTX Spark
- Anthropic Plans to Change Data Retention Policy for Advanced AI
- ATDev updates progress on autonomous wheelchair with robotic arm
- Claude Code demo generates full MP4 videos with Qwen3-TTS and FLUX.2
- Grok chatbot spews gibberish to users in rare glitch