multimodal
972 个项目 · ⭐ 708.0k[ICLR 2026] Official code repository for "⚡️VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration"
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
[ICLR 2025] MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
Democratizing Augmented Reality research and development.
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
GAKG is a multimodal Geoscience Academic Knowledge Graph (GAKG) framework by fusing papers' illustrations, text, and bibliometric data.
Efficient Multimodal Foundation Model Adaptation for Recommendation
Unified interface for interacting with various LLMs hundreds of models, caching, fallback mechanisms, and enhanced reliability.
Official Implementation of VisualClaw: A Real-Time, Personalized Agent for the Physical World
OpenCV Spatial AI Competition 3rd prize winner project "Automatic Lawn Mower Navigation"
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
专利侵权分析系统 —— 输入专利公开号,产出竞品侵权分析报告;同时打包成 skill,可被任意 agent(codex,claude code 等) 调用。
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
共 972 条 · 第 26 / 49 页