video
177 个项目 · ⭐ 1086.8kMulti-Modal Transformer for Video Retrieval
General video interaction platform based on LLMs, including Video ChatGPT
Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
Recrafting Video Ads with Generative AI
Python API & command-line tool to easily transcribe speech-based video files into clean text
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video (ICML 2023)
Obsidian plugin to create high-quality transcriptions from markdown linked audio files
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
Instantly create video clips from LLM prompts
Easiest way of fine-tuning HuggingFace video classification models
[ICLR 2026] DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
Implementation of a multimodal diffusion transformer in Pytorch
AI-native studio for multi-device shows that happen inside phones.
共 177 条 · 第 7 / 9 页