Run a 125B parameter AI model on your consumer GPU with Strata

You can now run a 125-billion-parameter AI model on a standard gaming PC at up to 100 tokens per second.
Developer Niko1221 released Strata, a free, open-source engine built to run Qwen3.8-Flash-Next locally. The tool splits model data across system memory and graphics memory, letting standard Nvidia or AMD cards with 12 GB of VRAM handle the workload. On a 24 GB card like an RTX 3090, output speeds range from 100 to 140 tokens per second.
Why it matters: Smart models of this size usually require a server cluster or paid subscription. Strata brings full local capability—code editing, image reading, and tool use—to desktop hardware without sending a single byte of data to the cloud.
Here's the gist: you need at least 12 GB of VRAM, 32 GB of system RAM, and roughly 80 GB of SSD space. Windows users can run START-HERE.bat while Linux users run ./setup.sh. You can also hand the setup documentation to your AI coding assistant to configure it automatically.
Your local graphics card just got a whole lot more capable.
Sources
- Strata GitHub Repository — https://github.com/Niko1221/Strata
- Hacker News Discussion — https://news.ycombinator.com/item?id=49953495

