The code and weight for LoVA. LoVA is a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models.
Project page for AV-Phys Bench: Do Joint Audio-Video Generation Models Understand Physics?
共 22 条 · 第 2 / 2 页