[ACMMM'24] MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm Detection
[WACV 2026 🔥] GAEA is a multimodal model with a new dataset and benchmark for context-aware image geolocation and QA.
A multimodal live AI assistant designed to enhance the browsing experience using Gemini.
Claude Code hook toolkit that gives vision-blind models a text description of pasted/tool-produced images.
Vision-Language Models on AMD GPUs — LLaVA, MiniGPT-4, Idefics on ROCm 🚀
[ICCV 2025] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
nvtop for vLLM — an interactive terminal dashboard for vLLM serving performance (concurrency, throughput, cache & KV memory, latency, spec-decode, GPU)
共 40673 条 · 第 2022 / 2034 页