← 返回专题广场
model-evaluation
8 个项目 · ⭐ 2.3k1
6 小时前
最近推送
3
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
Python
⭐ 102
⑂ 11
NOASSERTION
· 16 小时前推送
16 小时前
最近推送
4
Ollama Model Test - Figure out the best model for the task
Python
⭐ 51
⑂ 1
MIT
· 2026-06-04推送
2026-06-04
最近推送
5
OpenLLM Monitor is a plug-and-play, real-time observability dashboard for monitoring and debugging LLM API calls across OpenAI, Ollama, OpenRouter, and more. Tracks tokens, latency, cost, retries, and lets you replay prompts — fully open-source and self-hostable.
JavaScript
⭐ 43
⑂ 8
MIT
· 2026-07-07推送
2026-07-07
最近推送
6
2026-06-15
最近推送
7
2026-02-09
最近推送
8
2024-01-31
最近推送