← 返回专题广场
awq
6 个项目 · ⭐ 1.2k1
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
Python
⭐ 742
⑂ 77
Apache-2.0
· 2026-05-14推送
2026-05-14
最近推送
2
[ICML 2026] GRACE-VLM: deployable INT4 Qwen3-VL via quantization-aware distillation.
Jupyter Notebook
⭐ 216
⑂ 6
Apache-2.0
· 7 天前推送
7 天前
最近推送
3
4 天前
最近推送
4
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Python
⭐ 51
⑂ 4
Unlicense
· 2026-05-10推送
2026-05-10
最近推送
5
1 天前
最近推送
6
An OpenAI Compatible API which integrates LLM, Embedding and Reranker. 一个集成 LLM、Embedding 和 Reranker 的 OpenAI 兼容 API
Python
⭐ 18
⑂ 1
Apache-2.0
· 2025-08-21推送
2025-08-21
最近推送