inference-engine
27 个项目 · ⭐ 37.1k✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
OneDiff: An out-of-the-box acceleration library for diffusion models.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
https://wavespeed.ai/ Context parallel attention that accelerates DiT model inference with dynamic caching
PyTorch library for cost-effective, fast and easy serving of MoE models.
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
The inference engine the open-source world built for itself.
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
A High-Performance LLM Inference Engine with vLLM-Style Continuous Batching
共 27 条 · 第 1 / 2 页