40442
--
质量分
40443
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
Python
⭐ 11
⑂ 0
Apache-2.0
· 13 天前推送
--
质量分
40444
Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs
Go
⭐ 11
⑂ 5
Apache-2.0
· 5 天前推送
--
质量分
40445
--
质量分
40446
Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling
Python
⭐ 11
⑂ 1
Apache-2.0
· 2026-01-25推送
--
质量分
40447
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
Python
⭐ 11
⑂ 2
Apache-2.0
· 2025-04-10推送
--
质量分
40448
--
质量分
40449
📑 npm pacakge to Craft files into Markdown with ease
JavaScript
⭐ 11
⑂ 1
Apache-2.0
· 2025-01-03推送
--
质量分
40450
--
质量分
40451
Spark Pulse is a web control plane for spark-vllm-docker
Python
⭐ 11
⑂ 4
NOASSERTION
· 2026-06-22推送
--
质量分
40452
--
质量分
40453
--
质量分
40454
--
质量分
40455
--
质量分
40456
--
质量分
40458
--
质量分
40460
--
质量分
共 40673 条 · 第 2023 / 2034 页