vllm
351 个项目 · ⭐ 216.0kWelcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
A Datacenter Scale Distributed Inference Serving Framework
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
A programmable Mixture-of-Models router for heterogeneous LLM inference
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
共 351 条 · 第 1 / 18 页