[ACMMM'24] MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm Detection
[WACV 2026 🔥] GAEA is a multimodal model with a new dataset and benchmark for context-aware image geolocation and QA.
A multimodal live AI assistant designed to enhance the browsing experience using Gemini.
Vision-Language Models on AMD GPUs — LLaVA, MiniGPT-4, Idefics on ROCm 🚀
[ICCV 2025] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
nvtop for vLLM — an interactive terminal dashboard for vLLM serving performance (concurrency, throughput, cache & KV memory, latency, spec-decode, GPU)
在nano-vllm基础上支持moe模型以及Speculative Decoding技术
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
共 7255 条 · 第 359 / 363 页