CourionAI
EN
Newsletter
← Glossary Model

Chirp

Google's earlier family of speech-to-text models, now largely superseded by transcription built directly into Gemini.

Chirp was Google’s dedicated speech recognition model, the thing that turned audio into text behind products and cloud APIs. It was a conventional speech model in the sense that it transcribed what it heard and stopped there.

It appears in our coverage mainly as the baseline that newer models are measured against. Google’s Gemini transcription models are compared to Chirp 3 on accuracy and on how long you wait for a final transcript, and the shift from a standalone speech model to a language model doing the listening is what allows the newer ones to clean up filler words and self-corrections as they go.