🚢 Yet another operator for running large language models on Kubernetes with ease. Powered by Ollama! 🐫
Kubernetes Copilot powered by AI (OpenAI/Claude/Gemini/etc)
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
🪶 Lightweight OpenAI drop-in replacement for Kubernetes
A diverse, simple, and secure all-in-one LLMOps platform
Hands-on MLOps projects to explore and learn the practical aspects of machine learning engineering for production.
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: 🚀 AI Gateway (LiteLLM) 🤖 LLM Serving (vLLM, SGLang, Ollama) 📊 Vector Databases, 🔍 Embedding Models (TEI) 📈 Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes
⚡Instant Stable Diffusion on k8s(Kubernetes) with Helm
Batteries Included is a Kubernetes based software platform for database, ai, web, monitoring, and more.
Template for AI chatbots & document management using Retrieval-Augmented Generation with vector search and FastAPI.
Arks is a cloud-native inference framework running on Kubernetes
共 324 条 · 第 15 / 17 页