inference
187 个项目 · ⭐ 698.5kVarious installation guides for Large Language Models
A high-performance, universal serving framework for any-to-any models.
whisper-cpp-serve Real-time speech recognition and c+ of OpenAI's Whisper model in C/C++
Build an Autonomous Web3 AI Trading Agent (BASE + Uniswap V4 example)
Android native AI inference library, bringing text, image, video, STT, TTS inference
DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding
run ollama & gguf easily with a single command
[⛔️ DEPRECATED] Friendli: the fastest serving engine for generative AI
Arks is a cloud-native inference framework running on Kubernetes
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
Find the optimal model serving solution for 🤗 Hugging Face models 🚀
First open-source TurboQuant KV cache compression for LLM inference. Drop-in for HuggingFace. pip install turboquant.
ElasticMM: Elastic and Efficient MLLM Serving System
[MobiCom 2022] InFi is a library for building input filters for resource-efficient inference.
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
共 187 条 · 第 7 / 10 页