OossaAI is evolving fast. We explain it simply.
Newsletter

Note · 1 min read

LlamAmpere v0.4 reports 95+ tokens per second on an RTX 3090

A developer says the latest version of the Llama.cpp fork can run a 27-billion-parameter Qwen model with a 262K-token context on one RTX 3090. The reported speed held through 100,000 generated tokens.

Oossa · About Oossa

Developer JakeATX has released v0.4 of LlamAmpere, a Llama.cpp fork tuned for Nvidia Ampere graphics cards. In a Reddit post, they report more than 95 tokens per second while generating 100,000 tokens with a 27-billion-parameter Qwen model on one RTX 3090.

The test used a compressed model and a context window set to 262,144 tokens. JakeATX says v0.4 was about 10% faster than the previous release and supported more than 10% more context; these are the developer’s reported results, not an independent benchmark.

Why it matters

If the reported results hold up for other users, they suggest a single older graphics card can handle very long text-generation runs at useful speeds.

Was this article useful?

Sources & references

#SourceOutletDateKey takeaway
195+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090 ↗Reddit r/LocalLLaMASep 28, 2026Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up

1 sources

Last updated:

Oossallms.txt.md

Share

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.