← 返回专题广场
paged-attention
4 个项目 · ⭐ 6.3k1
8 天前
最近推送
2
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
⭐ 655
⑂ 102
Apache-2.0
· 1 天前推送
1 天前
最近推送
3
A High-Performance LLM Inference Engine with vLLM-Style Continuous Batching
C++
⭐ 119
⑂ 7
MIT
· 2026-01-03推送
2026-01-03
最近推送
4
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
Python
⭐ 13
⑂ 1
Apache-2.0
· 10 天前推送
10 天前
最近推送