← 返回专题广场
llm-inference
197 个项目 · ⭐ 317.4k181
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
Python
⭐ 13
⑂ 1
Apache-2.0
· 10 天前推送
10 天前
最近推送
182
13 天前
最近推送
183
The Continuous Verification and Optimization layer for self-hosted LLMs
Python
⭐ 13
⑂ 0
Apache-2.0
· 2026-07-18推送
2026-07-18
最近推送
184
3 天前
最近推送
185
6 天前
最近推送
186
19 小时前
最近推送
187
11 天前
最近推送
188
2025-03-06
最近推送
189
2026-07-19
最近推送
190
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
Python
⭐ 11
⑂ 0
Apache-2.0
· 12 天前推送
12 天前
最近推送
191
Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs
Go
⭐ 11
⑂ 5
Apache-2.0
· 3 天前推送
3 天前
最近推送
192
2026-07-07
最近推送
193
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
Python
⭐ 11
⑂ 2
Apache-2.0
· 2025-04-10推送
2025-04-10
最近推送
194
1 天前
最近推送
195
2025-10-29
最近推送
196
3 天前
最近推送
197
9 天前
最近推送
共 197 条 · 第 10 / 10 页