Frontier Foundation Models for Video Understanding
Media understanding
Referring Any Pixel from Image and Video
Bringing MLLMs into Embodied World
VideoRefer x VideoLLaMA3
Create short AI videos from text or first and last frames, with synchronized dialogue, effects, and ambience.