Overview of Wan2.2 Video Generative Models
mainWan2.2 is a suite of advanced large-scale video generative models featuring a Mixture-of-Experts (MoE) architecture. Key capabilities include:
- MoE Architecture: Uses specialized expert models for different denoising stages to increase capacity without increasing computational cost.
- Cinematic Aesthetics: Trained on curated data with detailed labels for lighting, composition, and color for precise style control.
- High-Definition Hybrid TI2V: A 5B model using a 16×16×4 compression ratio VAE, supporting Text-to-Video (T2V) and Image-to-Video (I2V) at 720P resolution (24fps). It is optimized to run on consumer-grade hardware like the NVIDIA RTX 4090.
- Specialized Models: Includes Wan2.2-Animate-14B for character animation/replacement and Wan2.2-S2V-14B for audio-driven (Speech-to-Video) generation.