NoteSeptember 29, 2026 at 7:33 AM · 1 min read
Five used mining boards run Qwen3-Coder-Next at 40 tokens per second
A Reddit user says a cluster of five BC-250 boards generates Qwen3-Coder-Next at 40 tokens per second. The setup reportedly cost less than $800, but uses power inefficiently.
A LocalLLaMA Reddit user has put five used BC-250 mining boards together to run Qwen3-Coder-Next, an AI model for writing code. They report about 40 tokens per second with a 30,000-token context, slowing to around 30 tokens per second at 100,000 tokens.
The boards share about 71GB of video memory and communicate over their built-in 1-gigabit Ethernet. The builder says the setup cost less than $800, while noting that it is “wildly inefficient” with power; they may add two more boards to test another model.
Why it matters
The build shows how used mining hardware can be repurposed for local AI, though its power use may limit its practicality.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s ↗ | r/LocalLLaMA | Sep 28, 2026 | Using an asrock 12 unit case running one board as the main with the rest of them headless. |
1 sources
Last updated: September 29, 2026
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.