CourionAI
EN
Newsletter
← Glossary Model

Pixtral

Mistral's vision component, the part that lets a model look at images instead of only reading text.

Pixtral started life as Mistral’s multimodal model, meaning one that handles pictures and text together. The vision encoder from it is now reused as a building block inside other Mistral models, including the Shieldstral safety classifier.

A vision encoder is the piece that turns an image into the same kind of internal representation the model already uses for words, so the rest of the network can reason about a photo and a sentence in one go. If a Mistral model can see, Pixtral is usually the reason why.