inference
187 个项目 · ⭐ 698.5kA tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
A collection of Kotlin-based examples featuring AI frameworks such as Spring AI, LangChain4j, and more — complete with Kotlin notebooks for hands-on learning.
Summarization, translation, sentiment-analysis, text-generation and more at blazing speed using a T5 version implemented in ONNX.
Provider-pluggable orchestration runtime for multi-model AI inference. ( Sakana Fugu style )
Local LLM Testing & Benchmarking for Apple Silicon
Deep Learning based Automatic Speech Recognition with attention for the Nvidia Jetson.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
TypeScript SDK for different in-browser AI model providers, built to make client-side AI integration simpler and more consistent across vendors.
Inference and fine-tuning examples for vision models from 🤗 Transformers
[deprecated] AI Gateway - core infrastructure stack for building production-ready AI Applications
⚡️ A fast and flexible PyTorch inference server that runs locally, on any cloud or AI HW.
共 187 条 · 第 5 / 10 页