vllm
351 个项目 · ⭐ 216.0kvLLM-5090: Docker Container for RTX 5090 + OpenCode
Your AI intranet: network the computers you already own for inference and training.
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
Official implementation of "DoRA: Weight-Decomposed Low-Rank Adaptation"
A High-Performance LLM Inference Engine with vLLM-Style Continuous Batching
Convert PowerPoint files into semantically rich text using vision language models
Data Center and Client workload and software optimizations for Intel hardware.
The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O on the request path, cache-affinity routing, RBAC, budgets and a 13-screen UI in the binary.
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
共 351 条 · 第 6 / 18 页