Note · 1 min read
Reddit user asks for a coding model between Qwen 3.8 27B and Flash Next
The user wants a model that balances coding ability with decode speed of 75‑100 tokens per second on an RTX 5090 system.
Oossa · About Oossa
On Sep 28 2026, a Reddit user posted in r/LocalLLaMA asking which model fits between Qwen 3.8 27B and Flash Next for coding tasks. They reported 200+ TPS with Qwen 3.8 27B and about 50 TPS with Flash Next on an RTX 5090 and 96 GB DDR5. They want a middle ground that keeps decode speed between 75‑100 TPS without sacrificing coding capability. The user plans to upgrade RAM to 128 GB but wants to avoid a major slowdown.
Why it matters
Finding a model with balanced speed and coding power helps developers iterate faster on limited hardware.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | What model sits between Qwen 3.8 27b and Flash next for coding? ↗ | Reddit r/LocalLLaMA | Sep 28, 2026 | Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabi |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.