inference
187 个项目 · ⭐ 698.5kBuild your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Example notebooks for working with SageMaker Studio Lab. Sign up for an account at the link below!
AI Productivity Tool - Free and open source, improve user productivity, and protect privacy and data security. Including but not limited to: built-in local exclusive ChatGPT, DeepSeek, Phi, Qwen and other models, one-click batch intelligent processing of pictures, videos, audio, etc.
Accurate, large-scale, and extensible simulator for LLM inference Systems
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. Join our Discord communty: https://discord.com/invite/TgHXuSJEk6
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
⚠️ Legacy repository for Geti v2.x. For Geti v3.0+, visit https://github.com/open-edge-platform/geti
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.
Guideline following Large Language Model for Information Extraction
https://wavespeed.ai/ Context parallel attention that accelerates DiT model inference with dynamic caching
Super performant RAG pipelines for AI apps. Summarization, Retrieve/Rerank and Code Interpreters in one simple API.
共 187 条 · 第 4 / 10 页