llm-inference
197 个项目 · ⭐ 317.4kData Center and Client workload and software optimizations for Intel hardware.
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
Gradio based tool to run opensource LLM models directly from Huggingface
Nano vLLM with vLLM v1's request scheduling strategy and chunked prefill
Granite Switch — Build AI models like you build software
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: 🚀 AI Gateway (LiteLLM) 🤖 LLM Serving (vLLM, SGLang, Ollama) 📊 Vector Databases, 🔍 Embedding Models (TEI) 📈 Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes
Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play".
Completion After Prompt Probability. Make your LLM make a choice
Flowchart-like UI to interconnect LLM's and Huggingface models, and deploy them as a REST API with little to no code.
VindexLLM is a pure Delphi, GPU-powered LLM inference engine that uses Vulkan compute shaders to run GGUF models entirely on the GPU. It performs full transformer inference without relying on Python, CUDA, or other external runtimes, requiring only vulkan-1.dll, which is typically included with modern GPU drivers.
[SOICT 2024] LLM-Powered Video Search: A Comprehensive Multimedia Retrieval System
Android native AI inference library, bringing text, image, video, STT, TTS inference
共 197 条 · 第 6 / 10 页