FlashAttention-style custom attention backend for vLLM on AMD MI50/MI60/Radeon VII (gfx906). Downstream fork of mixa3607/ML-gfx906 with replacement HIP kernels and a vllm.general_plugins entry point.
Open-source local AI server configs, GFX906 runtime maintenance, reproducible benchmarks, and QC methods for affordable AI research infrastructure.