← 返回专题广场
inference-optimization
7 个项目 · ⭐ 2.8k1
High-efficiency floating-point neural network inference operators for mobile, server, and Web
C
⭐ 2.4k
⑂ 545
NOASSERTION
· 20 小时前推送
20 小时前
最近推送
2
4 天前
最近推送
3
A physics-grounded, cost-aware optimization loop for vLLM
Rust
⭐ 64
⑂ 8
NOASSERTION
· 4 天前推送
4 天前
最近推送
4
6 天前
最近推送
5
TurboQuant KV cache compression plugin for vLLM — asymmetric K/V, 8 models validated, consumer GPUs
Python
⭐ 49
⑂ 7
Apache-2.0
· 2026-04-10推送
2026-04-10
最近推送
6
11 天前
最近推送
7
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
Python
⭐ 13
⑂ 1
Apache-2.0
· 10 天前推送
10 天前
最近推送