Eigen RadarAI
Analysis

Cloudflare’s Clef-omni turns audio and video into structured choices

Cloudflare has made Clef-omni available on Workers AI, adding audio and video to its decision-model inputs. The model scores permitted answers in a single pass and aligns a video’s soundtrack with its frames. An independent checkpoint examination confirms the direct media path. Applications can submit sounds and images together without first turning the audio into a transcript.

Artificial Intelligence··Evening
A bright experimental setup with a camera and microphone beside a fan.

Audio and video enter the decision model directly

Cloudflare, the cloud infrastructure company, made Clef-omni available on its Workers AI model-hosting service on October 9. The new model accepts audio and video alongside text and images. It returns probabilities for answers that an application explicitly permits, without producing free-form text. Florian Zimmermeister’s examination of the released checkpoint and reference code confirms that media enter the decision path directly.[1], [2]

A video’s soundtrack is aligned with its frames. Applications can therefore supply sound and visual information in the same request instead of first transcribing speech and then asking a separate model to decide. The supported media include WAV and MP3 audio and MP4 and WebM video, which can accompany text and images in a request.[1]

The model scores permitted answers in one pass

Clef-omni is developed from Qwen3-Omni-30B-A3B-Instruct, a multimodal mixture-of-experts model. Cloudflare retains the comprehension backbone while excluding speech-generation components from the decision path. The model scores all permitted answers to defined questions in a single processing pass. Text, acoustic and visual evidence contribute to those option probabilities; the system does not generate a spoken answer or a written explanation.[1]

Equipment checks combine a photograph, sound and moving images

Cloudflare illustrates the interface with a photograph of an installed unit, a recording of it running and a video of its fan. Questions concern whether a label is visible, whether sounds are abnormal and whether the fan operates. The model uses the existing System One application interface and works with AI Gateway, Cloudflare’s service for AI applications. A Workers AI request selects the new model identifier and model selector.[1]

References

  1. News sourceCloudflareClef-omni makes structured decisions directly from audio and video↩1↩2↩3↩4
  2. News sourceflozi00Clef Omni processes audio and video into schema decisions↩