sglang
39 个项目 · ⭐ 44.0kA Datacenter Scale Distributed Inference Serving Framework
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Control panel for VLLM, Sglang, llama.cpp, exllamav3
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
MOVA: Towards Scalable and Synchronized Video–Audio Generation
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
共 39 条 · 第 1 / 2 页