llm-inference
197 个项目 · ⭐ 317.4kAn endpoint server for efficiently serving quantized open-source LLMs for code.
run ollama & gguf easily with a single command
CompanionLLM - A framework to finetune LLMs to be your own sentient conversational companion
[ICPP'25] TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Testing speed and accuracy of RAG with, and without Cross Encoder Reranker.
[⛔️ DEPRECATED] Friendli: the fastest serving engine for generative AI
Collection of Basic Prompt Templates for Various Chat LLMs (Chat LLM 的基础提示模板集合)
A high-performance Llama3 implementation using Micronaut and GraalVM Native Image
共 197 条 · 第 7 / 10 页