llm-inference
197 个项目 · ⭐ 317.4kStreamlines and simplifies prompt design for both developers and non-technical users with a low code approach.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for custom training and inferencing.
LLMFlows - Simple, Explicit and Transparent LLM Apps
GPU environment and cluster management with LLM support
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space
A command-line interface tool for serving LLM using vLLM.
Self-hosted personalized AI in a mirror.
共 197 条 · 第 3 / 10 页