cuda
132 个项目 · ⭐ 526.1kStable and Efficient Reinforcement Learning for Trillion-Parameter LLMs
🪶 Lightweight OpenAI drop-in replacement for Kubernetes
vLLM-5090: Docker Container for RTX 5090 + OpenCode
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
Gradio based tool to run opensource LLM models directly from Huggingface
Docker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by faster-whisper. Supports all Whisper models, NVIDIA GPU (CUDA) acceleration, JSON/SRT/VTT output, SSE streaming, offline mode, and multi-arch (amd64, arm64).
⚡Instant Stable Diffusion on k8s(Kubernetes) with Helm
An efficient implementation of RNN-T Prefix Beam Search in C++/CUDA.
共 132 条 · 第 5 / 7 页