Audio and video enter the decision model directly
Cloudflare, the cloud infrastructure company, made Clef-omni available on its Workers AI model-hosting service on October 9. The new model accepts audio and video alongside text and images. It returns probabilities for answers that an application explicitly permits, without producing free-form text. Florian Zimmermeister’s examination of the released checkpoint and reference code confirms that media enter the decision path directly.[1], [2]
A video’s soundtrack is aligned with its frames. Applications can therefore supply sound and visual information in the same request instead of first transcribing speech and then asking a separate model to decide. The supported media include WAV and MP3 audio and MP4 and WebM video, which can accompany text and images in a request.[1]
