inference
187 个项目 · ⭐ 698.5kCommunity maintained hardware plugin for vLLM on Ascend
Fast ML inference & training for ONNX models in Rust
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
High-efficiency floating-point neural network inference operators for mobile, server, and Web
Turn any computer or edge device into a command center for your computer vision projects.
Mastering Applied AI, One Concept at a Time
Communicate with an LLM provider using a single interface
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
Manages Unified Access to Generative AI Services built on Envoy Gateway
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
TensorFlow template application for deep learning
Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models 🚀
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
共 187 条 · 第 3 / 10 页