hec-ovi

hec-ovi/vllm-awq4-qwen

vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.

⭐ 51 ⑂ 4 Python Unlicense · 2026-05-10推送
51
Watchers
0
贡献者
0
Commits
0
Releases
3
Open Issues
2026-05-10
最近推送
原文 中文
暂无 README