vllm
351 个项目 · ⭐ 216.0kgpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。
A PyTorch native library for training speculative decoding models
A minimal interface for AI Companion that runs entirely in your browser.
A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
vLLM Documentation in Chinese Simplified / vLLM 中文文档
Open-source tools for training and evaluating Vision Language Models for OCR
Fully-featured, beautiful web interface for vLLM - built with NextJS.
[ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark
Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Booster - open accelerator for LLM models. Better inference and debugging for AI hackers
Automated system for LLM evaluation via agents. Doc as below:
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
共 351 条 · 第 5 / 18 页