serving
21 个项目 · ⭐ 90.5kRay is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
A flexible, high-performance serving system for machine learning models
An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models
learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen
Serve, optimize and scale PyTorch models in production
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
Database system for AI-powered apps
TensorFlow template application for deep learning
A comprehensive guide to building RAG-based LLM applications for production.
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.
[⛔️ DEPRECATED] Friendli: the fastest serving engine for generative AI
ElasticMM: Elastic and Efficient MLLM Serving System
This hands-on lab walks you through a step-by-step approach to efficiently serving and fine-tuning large-scale Korean models on AWS infrastructure.
🎹 Instruct.KR 2025 Summer Meetup: 오픈소스 LLM, vLLM으로 Production까지 🎹
共 21 条 · 第 1 / 2 页