inference
187 个项目 · ⭐ 698.5kA high-throughput and memory-efficient inference and serving engine for LLMs
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Making large AI models cheaper, faster and more accessible
Cross-platform, customizable ML solutions for live and streaming media.
SGLang is a high-performance serving framework for large language models and multimodal models.
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Large Language Model Text Generation Inference
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
共 187 条 · 第 1 / 10 页