Qwen3.8-Flash-Next on the EVO-X2: 125B Parameters in 96GB of Unified Memory
What it took to run a 125B MoE model at 17-25 tok/s on a $3K mini PC — the BIOS carve-out, ROCm 7.2.4 on gfx1151, a llama.cpp grammar bug, and MTP.
Exploring code, technology, and gaming one post at a time.
What it took to run a 125B MoE model at 17-25 tok/s on a $3K mini PC — the BIOS carve-out, ROCm 7.2.4 on gfx1151, a llama.cpp grammar bug, and MTP.
How I got ROCm 7.2.4 working on gfx1151, built llama.cpp correctly for AMD Strix Halo, and what 67 tok/s on a mini PC actually means.
How I built and connected a working MCP server in Python — understanding the protocol, the tooling, and how it wires up to Copilot CLI over stdio.
How I built a minimal, fully understood swaybar status line with battery status, volume control, and mute detection — no plugins required.
Learn from my mistakes building a SvelteKit blog with AI assistance and deploying to Netlify. Avoid common pitfalls and configuration issues.
My journey creating a blog from scratch while embracing AI as a coding partner. Why I believe AI-assisted development is the future of programming.