nvidia
93 个项目 · ⭐ 164.5kKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
Stable Diffusion UI: Diffusers (CUDA/ONNX)
Multi-language agent runtime and library for execution scope management, lifecycle events, and middleware on tool and LLM calls.
Deploy stable diffusion model with onnx/tenorrt + tritonserver
One-click install for StabilityAI's Stable-Diffusion with AUTOMATIC1111's webui
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
共 93 条 · 第 3 / 5 页