vllm
353 个项目 · ⭐ 216.4kModern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.
An endpoint server for efficiently serving quantized open-source LLMs for code.
High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
The open-source AI platform for enterprises that can't send data to the cloud. OpenAI-compatible API, full management dashboard, zero data egress.
Real-time streaming TTS server for XTTS-v2 on vLLM — OpenAI-compatible API, ~0.5s TTFB, Docker
DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding
[ICPP'25] TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Arks is a cloud-native inference framework running on Kubernetes
TurboQuant KV cache compression plugin for vLLM — asymmetric K/V, 8 models validated, consumer GPUs
Batch Deployment for Document Parsing with AWS Batch & Qwen-2.5-VL
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
A high-performance RDMA distributed file system for fast LLM Inference and GPU Training.
Android keyboard with local AI (Ollama, Whisper, MCP) or cloud (Gemini, Groq, OpenAI)
共 353 条 · 第 9 / 18 页