model-serving
38 个项目 · ⭐ 149.2kBentoDiffusion: A collection of diffusion models served with BentoML
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.
Deploy, serve, and run a Hugging Face model on a Raspberry Pi with just a few lines of code
Find the optimal model serving solution for 🤗 Hugging Face models 🚀
Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.
An optimized FastAPI server for OpenAI's Whisper whisper-large-v3-turbo model using MLX optimization
Honeycomb Lab — hex map + OpenAI gateway control plane for a home AI fleet
Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs
共 38 条 · 第 2 / 2 页