vision-and-language
40 个项目 · ⭐ 59.4kLAVIS - A One-stop Library for Language-Vision Intelligence
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
[EMNLP-2024] Build multimodal language agents for fast prototype and production
Codebase for Aria - an Open Multimodal Native MoE
Research code for ECCV 2020 paper "UNITER: UNiversal Image-TExt Representation Learning"
This repository is a curated collection of the most exciting and influential CVPR 2023 papers. 🔥 [Paper + Code]
Creating a software for automatic monitoring in online proctoring
PyTorch code for "Unifying Vision-and-Language Tasks via Text Generation" (ICML 2021)
HPT - Open Multimodal LLMs from HyperGAI
[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Code/Data for the paper: "LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding"
共 40 条 · 第 1 / 2 页