Note · 1 min read
LlamAmpere v0.4 reports 95+ tokens per second on an RTX 3090
A developer says the latest version of the Llama.cpp fork can run a 27-billion-parameter Qwen model with a 262K-token context on one RTX 3090. The reported speed held through 100,000 generated tokens.
Oossa · About Oossa
Developer JakeATX has released v0.4 of LlamAmpere, a Llama.cpp fork tuned for Nvidia Ampere graphics cards. In a Reddit post, they report more than 95 tokens per second while generating 100,000 tokens with a 27-billion-parameter Qwen model on one RTX 3090.
The test used a compressed model and a context window set to 262,144 tokens. JakeATX says v0.4 was about 10% faster than the previous release and supported more than 10% more context; these are the developer’s reported results, not an independent benchmark.
Why it matters
If the reported results hold up for other users, they suggest a single older graphics card can handle very long text-generation runs at useful speeds.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | 95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090 ↗ | Reddit r/LocalLLaMA | Sep 28, 2026 | Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.