[ACMMM'24] MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm Detection
[WACV 2026 🔥] GAEA is a multimodal model with a new dataset and benchmark for context-aware image geolocation and QA.
A multimodal live AI assistant designed to enhance the browsing experience using Gemini.
Vision-Language Models on AMD GPUs — LLaVA, MiniGPT-4, Idefics on ROCm 🚀
[ICCV 2025] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
nvtop for vLLM — an interactive terminal dashboard for vLLM serving performance (concurrency, throughput, cache & KV memory, latency, spec-decode, GPU)
在nano-vllm基础上支持moe模型以及Speculative Decoding技术
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
Spark Pulse is a web control plane for spark-vllm-docker
共 9100 条 · 第 451 / 455 页