multimodal
972 个项目 · ⭐ 708.1kImproving Chest X-Ray Report Generation by Leveraging Warm-Starting
C++ framework to develop multimodal path planning requests
The largest multilingual image-text classification dataset. It contains fashion products.
An benchmark for evaluating the capabilities of large vision-language models (LVLMs)
[NeurIPS'25 Spotlight] Official implementation of "JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation"
My implementation of Kosmos2.5 from the paper: "KOSMOS-2.5: A Multimodal Literate Model"
一条命令,把微信变成任何 AI Agent 的入口 | Connect WeChat to any AI Agent with one command
A high-performance, universal serving framework for any-to-any models.
Image Classification Testing with LLMs
GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection (AAAI 2024)
ARES - Automatic Robot Evaluation System. A simple, scalable solution for robotics research
Baselines for CCKS 2022 Task "Link Prediction for Multimodal Product Knowledge Graph"
共 972 条 · 第 23 / 49 页