vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Ollama with intel (i)GPU acceleration in docker and benchmark
Ollama Model Test - Figure out the best model for the task
The new DARNA. HI Version2.5. Aspires to be your self hosted Personal Health concierge on Linux. pwd 'health'
86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.
CRISP — The missing layer between vibe coding and building something people actually want.
Train Llama 3 models from scratch. Any scale, any personality. By Arianna Method.
Physics-inspired transformer modules based on mean-field dynamics of vector-spin models in JAX
[TITS] reliable dense matching based point cloud registration for autonomous driving
Train GEMMA on TPU/GPU! (Codebase for training Gemma-Ko Series)
Interpretable Pre-Trained Transformers for Heart Time-Series Data
Human parsing model for fashion and virtual try-on applications
[⛔️ DEPRECATED] Friendli: the fastest serving engine for generative AI
共 9100 条 · 第 381 / 455 页