← 返回专题广场
dflash
5 个项目 · ⭐ 2.1k1
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
Python
⭐ 1.5k
⑂ 172
NOASSERTION
· 15 天前推送
15 天前
最近推送
2
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.
Python
⭐ 449
⑂ 47
Apache-2.0
· 2026-07-03推送
2026-07-03
最近推送
3
DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding
Python
⭐ 54
⑂ 9
Apache-2.0
· 2026-06-28推送
2026-06-28
最近推送
4
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Python
⭐ 51
⑂ 4
Unlicense
· 2026-05-10推送
2026-05-10
最近推送