A multimodal live AI assistant designed to enhance the browsing experience using Gemini.
Claude Code hook toolkit that gives vision-blind models a text description of pasted/tool-produced images.
Vision-Language Models on AMD GPUs — LLaVA, MiniGPT-4, Idefics on ROCm 🚀
[ICCV 2025] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
nvtop for vLLM — an interactive terminal dashboard for vLLM serving performance (concurrency, throughput, cache & KV memory, latency, spec-decode, GPU)
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs
Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
共 26058 条 · 第 1296 / 1303 页