llm-serving
49 个项目 · ⭐ 237.8kA high-throughput and memory-efficient inference and serving engine for LLMs
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Superduper: End-to-end framework for building custom AI applications and agents.
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
Community maintained hardware plugin for vLLM on Ascend
MoBA: Mixture of Block Attention for Long-Context LLMs
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
This is suite of the hands-on training materials that shows how to scale CV, NLP, time-series forecasting workloads with Ray.
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
共 49 条 · 第 1 / 3 页