# Reddit user asks for a coding model between Qwen 3.8 27B and Flash Next

> The user wants a model that balances coding ability with decode speed of 75‑100 tokens per second on an RTX 5090 system.

Oossa · 2026-09-28 · https://oossa.com/en/reddit-user-asks-for-a-coding-model-between-qwen-3-8-27b-and-flash-next

On Sep 28 2026, a Reddit user posted in r/LocalLLaMA asking which model fits between Qwen 3.8 27B and Flash Next for coding tasks. They reported 200+ TPS with Qwen 3.8 27B and about 50 TPS with Flash Next on an RTX 5090 and 96 GB DDR5. They want a middle ground that keeps decode speed between 75‑100 TPS without sacrificing coding capability. The user plans to upgrade RAM to 128 GB but wants to avoid a major slowdown.

## The facts

- Qwen 3.8 27B: >200 tokens per second; Flash Next: ~50 TPS on RTX 5090.
- User hardware: RTX 5090 with 96 GB DDR5 (upgrade to 128 GB planned).

## Why it matters

Finding a model with balanced speed and coding power helps developers iterate faster on limited hardware.

## Sources & references

1. [What model sits between Qwen 3.8 27b and Flash next for coding?](https://www.reddit.com/r/LocalLLaMA/comments/1wsekkj/what_model_sits_between_qwen_38_27b_and_flash/) – Reddit r/LocalLLaMA, 2026-09-28

Last updated: 2026-09-28
