vision-language
29 个项目 · ⭐ 31.2k[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
A Framework of Small-scale Large Multimodal Models
Official implementation of SEED-LLaMA (ICLR 2024).
Official Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
💐Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
MixGen: A New Multi-Modal Data Augmentation
共 29 条 · 第 1 / 2 页