guqiong96

guqiong96/Lvllm

LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.

⭐ 446 ⑂ 39 Python Apache-2.0 · 3 天前推送
446
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
3 天前
最近推送
原文 中文
暂无 README