kekzl

kekzl/imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

⭐ 36 ⑂ 2 Cuda MIT · 2 小时前推送
36
Watchers
0
贡献者
0
Commits
0
Releases
91
Open Issues
2 小时前
最近推送
原文 中文
暂无 README