gpu
164 个项目 · ⭐ 715.5kGPU environment and cluster management with LLM support
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon
RAG (Retrieval-augmented generation) ChatBot that provides answers based on contextual information extracted from a collection of Markdown files.
A work in progress to build out solutions in Rust for MLOPs
Vector search engine inside Milvus, integrating FAISS, HNSW, DiskANN.
🚀 200倍速!AI时代的下载神器 | Docker/PyPI/HuggingFace/CRAN 全加速 | 并行分片+智能缓存,让下载飞起来
Local LLM Testing & Benchmarking for Apple Silicon
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Ollama Cloud is a Highly Scalable Cloud-native Stack for Ollama
GGNN: State of the Art Graph-based GPU Nearest Neighbor Search
RapidFire AI: Rapid AI Customization from RAG to Fine-Tuning
共 164 条 · 第 5 / 9 页