OossaAI is evolving fast. We explain it simply.
Newsletter

Note · 1 min read

Reddit user asks for a coding model between Qwen 3.8 27B and Flash Next

The user wants a model that balances coding ability with decode speed of 75‑100 tokens per second on an RTX 5090 system.

Oossa · About Oossa

On Sep 28 2026, a Reddit user posted in r/LocalLLaMA asking which model fits between Qwen 3.8 27B and Flash Next for coding tasks. They reported 200+ TPS with Qwen 3.8 27B and about 50 TPS with Flash Next on an RTX 5090 and 96 GB DDR5. They want a middle ground that keeps decode speed between 75‑100 TPS without sacrificing coding capability. The user plans to upgrade RAM to 128 GB but wants to avoid a major slowdown.

Why it matters

Finding a model with balanced speed and coding power helps developers iterate faster on limited hardware.

Was this article useful?

Sources & references

#SourceOutletDateKey takeaway
1What model sits between Qwen 3.8 27b and Flash next for coding? ↗Reddit r/LocalLLaMASep 28, 2026Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabi

1 sources

Last updated:

Oossallms.txt.md

Share

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.