← 返回专题广场
autoscaling
6 个项目 · ⭐ 10.9k1
1 天前
最近推送
2
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go
⭐ 196
⑂ 30
Apache-2.0
· 2 小时前推送
2 小时前
最近推送
3
3 天前
最近推送
4
Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.
Go
⭐ 94
⑂ 10
Apache-2.0
· 2026-05-14推送
2026-05-14
最近推送
5
This repository shows various ways of deploying a vision model (TensorFlow) from 🤗 Transformers.
Jupyter Notebook
⭐ 30
⑂ 3
Apache-2.0
· 2022-08-22推送
2022-08-22
最近推送
6
Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs
Go
⭐ 11
⑂ 5
Apache-2.0
· 3 天前推送
3 天前
最近推送