gguf
78 个项目 · ⭐ 101.8kRun AI ✨ assistant locally! with simple API for Node.js 🚀
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
Joy Caption is a ComfyUI node using the LLaVA model to generate stylized image captions, supporting batch processing and GGUF models.
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Download models from the Ollama library, without Ollama
InferrLM - On-device AI for iOS & Android
Local Qwen 3.8 27B uncensored Q4_K_M + harvested SYSTEM pack. Official 3.8 weights, not a 3.6 retitle.
Gradio based tool to run opensource LLM models directly from Huggingface
共 78 条 · 第 2 / 4 页