cuda
132 个项目 · ⭐ 526.1kSGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.
Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. Join our Discord communty: https://discord.com/invite/TgHXuSJEk6
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
🚀 200倍速!AI时代的下载神器 | Docker/PyPI/HuggingFace/CRAN 全加速 | 并行分片+智能缓存,让下载飞起来
AUTOMATIC1111/stable-diffusion-webui for CUDA and ROCm on NixOS
GGNN: State of the Art Graph-based GPU Nearest Neighbor Search
共 132 条 · 第 4 / 7 页