vlm
117 个项目 · ⭐ 386.2k[ICLR 2026🔥] SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
Convert pdf and image files into markdown
EgoNormia | Benchmarking Physical Social Norm Understanding in VLMs
A full-modal personal knowledge base built on the Karpathy LLM Wiki concept.
The official implement of Unified Reasoning Emotion Generalist OneEmo
[WACV 2026 🔥] GAEA is a multimodal model with a new dataset and benchmark for context-aware image geolocation and QA.
GDB: GraphicDesignBench - A real-world benchmark for evaluating AI on graphic design tasks
Automated comic cataloging tool that identifies issues directly from cover images using a vision-language model, then cross-references results with the Grand Comics Database and the ComicVine API to generate structured, high-confidence collection data with minimal manual entry.
共 117 条 · 第 6 / 6 页