Note · 1 min read
Swift 1.5 cuts benchmark task times on an RTX 3090
A Reddit user reports that a modified Swift 1.5 model completed benchmark tasks about 37% faster than HyperQwen’s base setup on one RTX 3090. The results come with small differences across quality tests.
Oossa · About Oossa
A Reddit user says a Swift 1.5 model adapted for HyperQwen averaged 68.2 seconds per task, compared with 108.1 seconds for HyperQwen’s fast-quantized Qwen model. The tests ran on a single RTX 3090 with 24GB of memory, using a configured 150,000-token context.
The poster says the modified model was faster overall because it generated fewer tokens, despite slightly lower decoding speed. Across several reported tests, scores were close but not identical; for example, Swift 1.5 with INT4 output layers scored 91% on a 100-problem coding test, versus 89% for HyperQwen.
Why it matters
If the poster’s results hold up in other tests, they suggest model tuning can reduce the time and token use needed for local AI tasks on consumer hardware.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090 ↗ | Reddit r/LocalLLaMA | Sep 28, 2026 | Hi everyone :) The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.