llama-cpp
104 个项目 · ⭐ 108.6kRun AI ✨ assistant locally! with simple API for Node.js 🚀
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.
This repo is to showcase how you can run a model locally and offline, free of OpenAI dependencies.
Boot your PC straight into an LLM. Rust, UEFI-resident, no operating system underneath.
On-device memory layer for AI agents. Claude Code, Hermes and OpenClaw. Hooks + MCP server + hybrid RAG search.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Input text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.
Booster - open accelerator for LLM models. Better inference and debugging for AI hackers
共 104 条 · 第 2 / 6 页