multimodal
972 个项目 · ⭐ 708.1kLanguage Models Can See: Plugging Visual Controls in Text Generation
Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images - ICCV 2021
Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
EPNet: Enhancing Point Features with Image Semantics for 3D Object Detection(ECCV 2020)
MedEvalKit: A Unified Medical Evaluation Framework
[CVPR 2022 Oral & TPAMI 2023] Learning Optical Flow and Scene Flow with Bidirectional Camera-LiDAR Fusion
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
A Continuously Updated Library for Advanced Models for Multimodal Recommendation
A comprehensive code domain benchmark review of LLM researches.
Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
[ICLR 2026] Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
OmniFusion — a multimodal model to communicate using text and images
💐Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
DeepVerse: 4D Autoregressive Video Generation as a World Model
共 972 条 · 第 13 / 49 页