llm-serving
49 个项目 · ⭐ 237.9k🪶 Lightweight OpenAI drop-in replacement for Kubernetes
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: 🚀 AI Gateway (LiteLLM) 🤖 LLM Serving (vLLM, SGLang, Ollama) 📊 Vector Databases, 🔍 Embedding Models (TEI) 📈 Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes
Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play".
A High-Efficiency System of Large Language Model Based Search Agents
A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
npm like package ecosystem for Prompts 🤖
[⛔️ DEPRECATED] Friendli: the fastest serving engine for generative AI
A high-performance RDMA distributed file system for fast LLM Inference and GPU Training.
A hybrid router that uses Spot GPU instances to reduce costs and Serverless GPUs for making Cold Starts faster.
面向开发者与初学者的 nano-vLLM 交互式源码教程 - 通过 13 个 HTML 互动实验 + 13 章中文教程理解大模型推理引擎
共 49 条 · 第 2 / 3 页