multimodal
972 个项目 · ⭐ 708.1k[CVPR2020] Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation
The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]
[ACL 2025] Code and data for OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
Embed arbitrary modalities (images, audio, documents, etc) into large language models.
KDD Cup 2020 Challenges for Modern E-Commerce Platform: Multimodalities Recall first place
The AdEMAMix Optimizer: Better, Faster, Older.
CVPR 2021: "Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE"
An official implementation for "X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval"
Winner system (DAMO-NLP) of SemEval 2022 MultiCoNER shared task over 10 out of 13 tracks.
FairyTailor: Multimodal Generative Framework for Storytelling
[NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis
[ICCV 2025] Perspective-Invariant 3D Object Detection
[ACL 2026 Oral] UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
[ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
共 972 条 · 第 15 / 49 页