← 返回专题广场
vllm
351 个项目 · ⭐ 216.0k201
2023-11-04
最近推送
202
2026-05-11
最近推送
203
Serve Llama 3.3 70B (with AWQ quantization) using vLLM and deploy it on BentoCloud.
Python
⭐ 30
⑂ 9
Apache-2.0
· 2025-01-28推送
2025-01-28
最近推送
205
2 天前
最近推送
206
6 天前
最近推送
207
2025-06-11
最近推送
208
3 天前
最近推送
209
16 天前
最近推送
210
A web-based memory usage and performance calculator for Huggingface GGUF models
JavaScript
⭐ 28
⑂ 1
· 2026-07-20推送
2026-07-20
最近推送
211
A hybrid router that uses Spot GPU instances to reduce costs and Serverless GPUs for making Cold Starts faster.
Python
⭐ 27
⑂ 7
MIT
· 2026-03-14推送
2026-03-14
最近推送
212
面向开发者与初学者的 nano-vLLM 交互式源码教程 - 通过 13 个 HTML 互动实验 + 13 章中文教程理解大模型推理引擎
JavaScript
⭐ 27
⑂ 0
MIT
· 2 天前推送
2 天前
最近推送
213
a simple lightweight large language model pipeline framework.
Python
⭐ 27
⑂ 2
Apache-2.0
· 2025-04-25推送
2025-04-25
最近推送
214
2024-02-17
最近推送
215
Orpheus TTS Server with streaming support (TTFB ~160ms)
Python
⭐ 26
⑂ 5
Apache-2.0
· 2025-09-21推送
2025-09-21
最近推送
216
2025-04-19
最近推送
217
A "standard library" of Triton kernels.
Python
⭐ 26
⑂ 3
Apache-2.0
· 2025-10-03推送
2025-10-03
最近推送
218
Official implementation for Text Generation Beyond Discrete Token Sampling
Python
⭐ 26
⑂ 4
Apache-2.0
· 2025-08-12推送
2025-08-12
最近推送
219
2026-05-09
最近推送
共 351 条 · 第 11 / 18 页