inference-engine
27 个项目 · ⭐ 37.1khummingbird is a lightweight, zero dependency runtime for massive open source Mixture of Experts (MoE) language models. It unifies SSD, RAM, and VRAM into a single intelligent memory hierarchy, enabling inference of models like GPT-OSS 120B, GLM, DeepSeek, Qwen, and more on consumer hardware
[⛔️ DEPRECATED] Friendli: the fastest serving engine for generative AI
Comprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP
AI-Inference-Managed-by-AI: Go binary for managing AI inference on edge devices
共 27 条 · 第 2 / 2 页