# Swift 1.5 cuts benchmark task times on an RTX 3090

> A Reddit user reports that a modified Swift 1.5 model completed benchmark tasks about 37% faster than HyperQwen’s base setup on one RTX 3090. The results come with small differences across quality tests.

Oossa · 2026-09-28 · https://oossa.com/en/swift-1-5-cuts-benchmark-task-times-on-an-rtx-3090

A Reddit user says a Swift 1.5 model adapted for HyperQwen averaged 68.2 seconds per task, compared with 108.1 seconds for HyperQwen’s fast-quantized Qwen model. The tests ran on a single RTX 3090 with 24GB of memory, using a configured 150,000-token context.

The poster says the modified model was faster overall because it generated fewer tokens, despite slightly lower decoding speed. Across several reported tests, scores were close but not identical; for example, Swift 1.5 with INT4 output layers scored 91% on a 100-problem coding test, versus 89% for HyperQwen.

## The facts

- Reported average task time: 68.2 seconds for Swift 1.5 with INT4 heads, versus 108.1 seconds for HyperQwen fast quant.
- Tests used one RTX 3090 with 24GB of memory and a 150,000-token configured context.

## Why it matters

If the poster’s results hold up in other tests, they suggest model tuning can reduce the time and token use needed for local AI tasks on consumer hardware.

## Sources & references

1. [Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090](https://www.reddit.com/r/LocalLLaMA/comments/1wsqjku/swift_15_hyperqwen_37_less_task_completion_time/) – Reddit r/LocalLLaMA, 2026-09-28

Last updated: 2026-09-28
