# Cloudflare adds audio and video to its Clef decision model and lowers price of Clef‑flash

> Clef‑omni can handle text, images, audio and video in one API call; Clef‑flash now costs $0.038 per million tokens, and the original Clef model is faster.

Oossa · 2026-10-09 · https://oossa.com/en/cloudflare-adds-audio-and-video-to-its-clef-decision-model-and-lowers-price-of-c

Cloudflare announced on Oct 9, 2026 that it is launching Clef‑omni, a decision model that accepts audio, video, image and text inputs together. The same update also makes the Clef‑flash model cheaper than the competing Jev model and speeds up the original Clef model on Cloudflare’s Workers AI platform.

## What Clef‑omni can do

Clef‑omni is built on a Qwen3‑Omni‑30B‑A3B‑Instruct mixture‑of‑experts foundation. Unlike earlier decision models that only handled text, it can directly ingest wav or mp3 audio and mp4 or webm video along with pictures and words. A single API request can now score options from any of these media types, eliminating the need for separate transcription or image‑recognition pipelines.

In tests, text‑only queries return in about 130 ms, image inputs in about 150 ms, audio clips in a few hundred milliseconds, and a 21‑second video with sound is processed in roughly 1.5 seconds.

## Pricing and speed changes

Clef‑flash’s price has been cut from $0.09 to $0.038 per million input tokens, making it cheaper than TypeSafe’s Jev model. The hosted Clef‑flash context window is now limited to 24 k tokens (down from 64 k), though the open‑weight weights still support 256 k tokens for self‑hosting.

The original Clef model remains at $0.24 per million tokens and now runs about twice as fast for large inputs thanks to infrastructure upgrades, including a move to the SGLang serving framework.

## The facts

- Clef‑omni launched on Oct 9, 2026 and can process text, image, audio (wav/mp3) and video (mp4/webm) in one call.
- Clef‑flash price dropped to $0.038 per million input tokens, cheaper than Jev.
- Clef‑flash context window reduced to 24 k tokens for the hosted version; self‑hosted weights still support 256 k tokens.
- Median response time for a 21‑second video is about 1.5 seconds.
- Clef model speedup: 800‑token requests now median 152 ms (down from 262 ms).

## Why it matters

Developers can now build a single AI step that looks at a photo, a sound clip and a short video to decide, for example, whether a home appliance is installed correctly. The lower price of Clef‑flash makes decision models affordable for small businesses that need to scan many images or documents, while the faster Clef model reduces latency for real‑time applications.

## Sources & references

1. [Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash](https://blog.cloudflare.com/clef-faster-cheaper-multimodal/) – Cloudflare, 2026-10-09

Last updated: 2026-10-09
