vision-and-language
40 个项目 · ⭐ 59.4kResearch code for EMNLP 2020 paper "HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training"
PyTorch code for "VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks" (CVPR2022)
OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
PyTorch code for “TVLT: Textless Vision-Language Transformer” (NeurIPS 2022 Oral)
[EMNLP'21] Visual News: Benchmark and Challenges in News Image Captioning
🥶Vilio: State-of-the-art VL models in PyTorch & PaddlePaddle
GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection (AAAI 2024)
source code and pre-trained/fine-tuned checkpoint for NAACL 2021 paper LightningDOT
[ACL 2021] Learning Relation Alignment for Calibrated Cross-modal Retrieval
[ICLR 2025] Official code repository for "TULIP: Token-length Upgraded CLIP"
Code for Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? [COLM 2024]
Grounding Language Models for Compositional and Spatial Reasoning
共 40 条 · 第 2 / 2 页