CourionAI
EN
Newsletter
← Glossary Term

Omnimodal

A model that handles text, images, audio and video together in one system rather than passing between specialised tools.

Multimodal usually means a model can handle more than one kind of input, often text plus images. Omnimodal is the marketing step beyond that: text, images, audio and video, all taken natively by a single model instead of being converted first.

The difference is not pedantic. The older approach was a relay: one tool transcribes speech, another describes the video frames, and a language model reasons about those descriptions. Every handoff loses exactly the things that carry meaning, such as hesitation in a voice, or what was on screen at the moment a sentence was spoken. A model that receives the raw audio and video keeps that material in play, which is why labs keep stressing the word native.