kekzl/imp
From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.
⭐ 36
⑂ 2
Cuda
MIT
· 2 小时前推送
36
Watchers
0
贡献者
0
Commits
0
Releases
91
Open Issues
2 小时前
最近推送
原文
中文
暂无 README