A collection of Foundation Vision Models that combine multiple models (CLIP, DINOv2, SAM, etc.).
Create short AI videos from text or first and last frames, with synchronized dialogue, effects, and ambience.