qwen3
41 个项目 · ⭐ 178.4kA high-throughput and memory-efficient inference and serving engine for LLMs
🔥 MaxKB is an open-source platform for building enterprise-grade agents. 强大易用的开源企业级智能体平台。
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
MimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Support
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
a magical LLM desktop client that makes it easy for *anyone* to use LLMs and MCP
Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
共 41 条 · 第 1 / 3 页