dgx-spark
30 个项目 · ⭐ 2.7ksparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.
Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)
共 30 条 · 第 1 / 2 页