Overview of Wan2.1 Video Generative Models
mainWan2.1 is an open suite of large-scale video foundation models designed for high-performance video generation. Key capabilities include:
- Multiple Tasks: Supports Text-to-Video (T2V), Image-to-Video (I2V), Video Editing, Text-to-Image, and Video-to-Audio.
- Consumer-grade GPU Support: The T2V-1.3B model is optimized for consumer hardware, requiring approximately 8.19 GB of VRAM. On an RTX 4090, it can generate a 5-second 480P video in about 4 minutes (without quantization).
- Visual Text Generation: Capable of generating robust Chinese and English text within videos.
- Powerful Video VAE: The
Wan-VAEefficiently encodes and decodes 1080P videos of any length while preserving temporal information.