llm-inference
197 个项目 · ⭐ 317.4kSuperduper: End-to-end framework for building custom AI applications and agents.
Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Optimizing inference proxy for LLMs
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.
Ultrafast serverless GPU inference, sandboxes, and background jobs
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser
共 197 条 · 第 2 / 10 页