ferrumox

ferrumox/fox

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.

⭐ 189 ⑂ 28 Rust NOASSERTION · 10 小时前推送
189
Watchers
0
贡献者
0
Commits
0
Releases
4
Open Issues
10 小时前
最近推送
原文 中文
暂无 README