# LlamAmpere v0.4 reports 95+ tokens per second on an RTX 3090

> A developer says the latest version of the Llama.cpp fork can run a 27-billion-parameter Qwen model with a 262K-token context on one RTX 3090. The reported speed held through 100,000 generated tokens.

Oossa · 2026-09-28 · https://oossa.com/en/llamampere-v0-4-reports-95-tokens-per-second-on-an-rtx-3090

Developer JakeATX has released v0.4 of LlamAmpere, a Llama.cpp fork tuned for Nvidia Ampere graphics cards. In a Reddit post, they report more than 95 tokens per second while generating 100,000 tokens with a 27-billion-parameter Qwen model on one RTX 3090.

The test used a compressed model and a context window set to 262,144 tokens. JakeATX says v0.4 was about 10% faster than the previous release and supported more than 10% more context; these are the developer’s reported results, not an independent benchmark.

## The facts

- The reported setup used one Nvidia RTX 3090 and a 262,144-token context.
- The developer says v0.4 improved speed by about 10% over the prior version.

## Why it matters

If the reported results hold up for other users, they suggest a single older graphics card can handle very long text-generation runs at useful speeds.

## Sources & references

1. [95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090](https://www.reddit.com/r/LocalLLaMA/comments/1wsmtd4/95_tps_through_100k_generated_for_qwen38_27b_262k/) – Reddit r/LocalLLaMA, 2026-09-28

Last updated: 2026-09-28
