Grounding Language Models for Compositional and Spatial Reasoning
[WACV 2026] Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
Identifying reasons for human actions in lifestyle vlogs.
[ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs
[ICLR 2026🔥] SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
Multi-Modal Hate Speech Detection using Deep Learning.
Open Translator: Speech To Speech and Speech to text Translator with voice cloning and other cool features
This repo contains the original implementation of VAuLT, the Vision-and-Augmented-Language Transformer. We provide instructions to download some multimodal social-media datasets, and scripts to experiment with. VAuLT is a stack of Transformers, a LM like BERT that preprocesses the text input of ViLT
共 40707 条 · 第 1978 / 2036 页