Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
-
Updated
Sep 19, 2026 - Python
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
使用稠密算力卡对稀疏大显存主机进行加速,主机包括 AI Max+ 395、M3 512G、DGX Spark 等。
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
llama.cpp with adaptive speculative decoding (--spec-draft-adaptive) and a Vulkan backend tuned for AMD Strix Halo. 4.7x on structured output, 1.9x mainline prefill on MoE.
The fastest way to run Qwen3.8 27B on Strix Halo (gfx1151)
Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI + systemd.
This is a mirror of the Strix Halo HomeLab wiki, to browse the wiki click on the link below
High-performance Lemonade alternative for AMD Strix Halo & Radeon — latest MoE models (Ling-3.0-Flash, Qwen 3.8 Flash Next, Ornith 1.5), bleeding-edge RDNA 3.5 Wave64/ROCmFP4 kernels, and silicon-tuned profiles.
Experimental support for many TTS/STT LLMs wrapped in a Wyoming API for consumption via Homeassistant
Local text to textured GLB on an AMD Strix Halo iGPU (gfx1151): FLUX.2 klein, a Vulkan-only TRELLIS.2 engine, and humanoid auto-rigging with SkinTokens on ROCm. No Blender, no CUDA.
A two-day, non-expert guide to running local LLMs on an AMD Strix Halo (Ryzen AI Max+ 395) box — full ~120GB memory pool, ROCm backend, NPU in parallel via FastFlowLM, one OpenAI-compatible endpoint.
A comprehensive guide to running Linux (Omarchy/Arch) on the 2025 ASUS ROG Flow Z13 (AMD Strix Halo). Includes CachyOS Kernel setup, Tablet Mode fixes, and Power Management for the Ryzen AI Max
Docker Compose for llama.cpp GGUF servers on AMD Strix Halo: Qwen, Gemma, and Laguna packages (abliterated and quantized), stock Vulkan plus ROCmFP4/MTP and ROCmFPX, parallel slots, with prefill/decode and quality metrics measured on this rig.
Self-hosted LLM/RAG stack in one command — AMD Strix Halo / x86_64 (ROCm/Vulkan, Docker Compose)
DeepSeek-V4-Flash inference for AMD Strix Halo
To associate your repository with the strix-halo topic, visit your repo's landing page and select "manage topics."