inference
187 个项目 · ⭐ 698.6kComprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP
Intelligent load balancer for distributed vLLM server clusters 分布式 vLLM 服务器集群的智能负载均衡器
Network-faithful simulation of LLM serving and training deployments
Deferred Continuous Batching in Resource-Efficient Large Language Model Serving (EuroMLSys 2024)
H.E.I.M.D.A.L.L looks at fleet telemetry and gives you natural-language insights. GPU data loading (cuDF), local LLM inference (Gemma 2), and production NIM on GKE. Open the notebooks, run cells, get answers! Quick start should not take longer than 10 minutes and the T4 path is completely free!
Know your VRAM before you run. Instant GPU memory estimates for any HuggingFace model.
LLM inference engine built from scratch in C++. No PyTorch, no frameworks.
Benchmark OpenAI-compatible AI endpoints and AI Accelerators in a reproducible structured way
Collection of ChatGPT alternatives & LLM tuning methods
共 187 条 · 第 9 / 10 页