inference
187 个项目 · ⭐ 698.5kYour AI intranet: network the computers you already own for inference and training.
CLI & Python API to easily summarize text-based files with transformers
Accelerated NLP pipelines for fast inference on CPU. Built with Transformers and ONNX runtime.
Deploy stable diffusion model with onnx/tenorrt + tritonserver
Full in-browser Semantic Search with Huggingface Transformers.js and ElectricSQL's PGlite!
The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O on the request path, cache-affinity routing, RBAC, budgets and a 13-screen UI in the binary.
152 open-source tools to run LLMs 100% locally – no cloud, no API keys, no censorship
Serve scikit-learn, XGBoost, TensorFlow, and PyTorch models with AWS Lambda container images support.
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
Nano vLLM with vLLM v1's request scheduling strategy and chunked prefill
Docker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by faster-whisper. Supports all Whisper models, NVIDIA GPU (CUDA) acceleration, JSON/SRT/VTT output, SSE streaming, offline mode, and multi-arch (amd64, arm64).
Open source subtitling platform 💻 for transcribing and translating videos/audios in Indic languages.
Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.
Qualcomm Cloud AI SDK (Platform and Apps) enable high performance deep learning inference on Qualcomm Cloud AI platforms delivering high throughput and low latency across Computer Vision, Object Detection, Natural Language Processing and Generative AI models.
A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.
共 187 条 · 第 6 / 10 页