cuda
132 个项目 · ⭐ 526.1kA flexible framework of neural networks for deep learning
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.
NVIDIA cuML: GPU-Accelerated Machine Learning
Optimized primitives for collective multi-GPU communication
Tengine is a lite, high performance, modular inference engine for embedded device
Lightning fast C++/CUDA neural network framework
A retargetable MLIR-based machine learning compiler and runtime toolkit.
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Single C file, Realtime CPU/GPU Profiler with Remote Web Viewer
Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.
Open source neural network chess engine with GPU acceleration and broad hardware support.
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
共 132 条 · 第 2 / 7 页