llm-inference
197 个项目 · ⭐ 317.4k✨ AI interface for tinkerers (Ollama, Haystack RAG, Python)
Explore the unknown, build the future, own your data.
This is suite of the hands-on training materials that shows how to scale CV, NLP, time-series forecasting workloads with Ray.
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
Context-Engine MCP - Agentic Context Compression Suite
Embedding Studio is a framework which allows you transform your Vector Database into a feature-rich Search Engine.
PyTorch library for cost-effective, fast and easy serving of MoE models.
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
A library to communicate with ChatGPT, Claude, Copilot, Gemini, HuggingChat, and Pi
CPU inference for the DeepSeek family of large language models in C++
On-device LLM Inference Powered by X-Bit Quantization
🌱 EcoLogits tracks the energy consumption and environmental footprint of using generative AI models through APIs.
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
Run generative AI models in sophgo BM1684X/BM1688
共 197 条 · 第 4 / 10 页