cuda
132 个项目 · ⭐ 526.1kDGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
Generate textured, segmented and rigged GLB assets entirely on your own GPU.
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
A high-performance RDMA distributed file system for fast LLM Inference and GPU Training.
useful Gentoo overlay Curated ebuilds, AI, tools & science
GPU 性能与 AI Infra 学习项目:CUDA/Triton 算子、NCU/NSYS、vLLM/SGLang/TRT-LLM/ms-swift、PyTorch/DeepSpeed/ms-swift 训练、并行架构
Docker image for a self-hosted WhisperLive real-time speech-to-text server, powered by faster-whisper. Provides WebSocket streaming for live audio transcription and an OpenAI-compatible REST API. Supports all Whisper models, VAD, NVIDIA GPU (CUDA) acceleration, offline mode, and multi-arch (amd64, arm64).
A "standard library" of Triton kernels.
🎹 Instruct.KR 2025 Summer Meetup: 오픈소스 LLM, vLLM으로 Production까지 🎹
共 132 条 · 第 6 / 7 页