quantization
114 个项目 · ⭐ 232.3kPMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
[CVPR 2024 Highlight & TPAMI 2025] This is the official PyTorch implementation of "TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models".
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
152 open-source tools to run LLMs 100% locally – no cloud, no API keys, no censorship
Docker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by faster-whisper. Supports all Whisper models, NVIDIA GPU (CUDA) acceleration, JSON/SRT/VTT output, SSE streaming, offline mode, and multi-arch (amd64, arm64).
Open source subtitling platform 💻 for transcribing and translating videos/audios in Indic languages.
🛠️ Tools for Transformers compression using PyTorch Lightning ⚡
Repository for the companion Colab notebook of the Domain-Specific Small Language Models book.
Roy: A lightweight, model-agnostic framework for crafting advanced multi-agent systems using large language models.
共 114 条 · 第 3 / 6 页