← 返回专题广场
efficient-inference
9 个项目 · ⭐ 5.0k1
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
Python
⭐ 1.9k
⑂ 136
MIT
· 2024-07-10推送
2024-07-10
最近推送
2
EfficientFormerV2 [ICCV 2023] & EfficientFormer [NeurIPs 2022]
Python
⭐ 1.1k
⑂ 95
NOASSERTION
· 2023-08-13推送
2023-08-13
最近推送
3
[CVPR 2024] DeepCache: Accelerating Diffusion Models for Free
Python
⭐ 970
⑂ 52
Apache-2.0
· 2024-06-27推送
2024-06-27
最近推送
4
On-device LLM Inference Powered by X-Bit Quantization
Python
⭐ 317
⑂ 26
Apache-2.0
· 11 天前推送
11 天前
最近推送
5
Explorations into some recent techniques surrounding speculative decoding
Python
⭐ 307
⑂ 24
MIT
· 2024-12-23推送
2024-12-23
最近推送
6
[NeurIPS 2024] AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
Python
⭐ 214
⑂ 13
Apache-2.0
· 2025-09-27推送
2025-09-27
最近推送
7
[NeurIPS'24] Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
Python
⭐ 73
⑂ 6
Apache-2.0
· 2025-01-23推送
2025-01-23
最近推送
8
[ICML 2025] RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
Python
⭐ 51
⑂ 10
NOASSERTION
· 2025-08-07推送
2025-08-07
最近推送
9
11 天前
最近推送