Oossa
Subscribe

NoteSeptember 29, 2026 at 7:33 AM · 1 min read

Five used mining boards run Qwen3-Coder-Next at 40 tokens per second

A Reddit user says a cluster of five BC-250 boards generates Qwen3-Coder-Next at 40 tokens per second. The setup reportedly cost less than $800, but uses power inefficiently.

A LocalLLaMA Reddit user has put five used BC-250 mining boards together to run Qwen3-Coder-Next, an AI model for writing code. They report about 40 tokens per second with a 30,000-token context, slowing to around 30 tokens per second at 100,000 tokens.

The boards share about 71GB of video memory and communicate over their built-in 1-gigabit Ethernet. The builder says the setup cost less than $800, while noting that it is “wildly inefficient” with power; they may add two more boards to test another model.

Why it matters

The build shows how used mining hardware can be repurposed for local AI, though its power use may limit its practicality.

Was this article useful?

Sources & references

#SourceOutletDateKey takeaway
1I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s ↗r/LocalLLaMASep 28, 2026Using an asrock 12 unit case running one board as the main with the rest of them headless.

1 sources

Last updated: September 29, 2026

Oossallms.txt.md

Share

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.