model-serving
38 个项目 · ⭐ 149.2kA high-throughput and memory-efficient inference and serving engine for LLMs
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
A framework for efficient model inference with omni-modality models
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
Community maintained hardware plugin for vLLM on Ascend
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
The missing bridge between your ML models and your AI agents.
共 38 条 · 第 1 / 2 页