inference
187 个项目 · ⭐ 698.6kConvert Hugging Face tokenizers to ONNX models for cross-language compatibility (.NET, Java, Python) with embedding models
Docker image for a self-hosted WhisperLive real-time speech-to-text server, powered by faster-whisper. Provides WebSocket streaming for live audio transcription and an OpenAI-compatible REST API. Supports all Whisper models, VAD, NVIDIA GPU (CUDA) acceleration, offline mode, and multi-arch (amd64, arm64).
Instruction Fine-Tuning of Meta Llama 3.2-3B Instruct on Kannada Conversations. Tailoring the model to follow specific instructions in Kannada, enhancing its ability to generate relevant, context-aware responses based on conversational inputs. Using the Kannada Instruct dataset for fine-tuning! Happy Finetuning 🎋
A python package to run inference with HuggingFace language and vision-language checkpoints wrapping many convenient features.
A web-based memory usage and performance calculator for Huggingface GGUF models
Typescript wrapper for the Hugging Face Inference API.
A hybrid router that uses Spot GPU instances to reduce costs and Serverless GPUs for making Cold Starts faster.
A "standard library" of Triton kernels.
Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text
[ICML 2025] Efficiently Serving Large Multimodal Models Using EPD Disaggregation
共 187 条 · 第 8 / 10 页