Improved Video VAE for Latent Video Diffusion Model
ByteDance/Tongyi Lab's IV-VAE introduces Keyframe-based Temporal Compression and Group Causal Convolution to achieve SOTA video reconstruction with balanced inter-frame performance
2 posts tagged with "latent-compression"
ByteDance/Tongyi Lab's IV-VAE introduces Keyframe-based Temporal Compression and Group Causal Convolution to achieve SOTA video reconstruction with balanced inter-frame performance
ByteDance's FSVideo achieves 42.3x speedup over Wan2.1 via 64×64×4 compression autoencoder, layer memory self-attention DIT, and few-step latent upsampler