Clean text gains another editor
Google's Gemini 3.5 Transcribe turns live or pre-recorded audio into formatted text while removing filler words and a speaker's own corrections. It automatically detects more than 85 languages and can label who said what, with timestamps, for up to three speakers in a recording. That can reduce the first editing pass for meeting notes, captions or a long piece of dictation.[1]
The convenience moves friction from the editing screen into the model's decision. Dropping an abandoned phrase helps with a shopping list; in an interview or feedback session, the same hesitation may reveal uncertainty or how a thought changed. A polished output saves time, while the route back to the raw audio still matters: the user needs to see what the model removed as well as the sentence it kept.[1]
What speed leaves open
Artificial Analysis measured an average word error rate of 4.0 per cent in streaming use and 2.6 per cent on recorded audio. Google also reports a 70 per cent reduction in time to the final transcript against its previous model. Those figures cover speed and mistaken words; they do not separately say how often a proper name, a language switch or a meaningful correction disappears. A lower rate therefore does not create the same level of trust for every user task.[1]
Detecting more than 85 languages describes the breadth of the model; Google's Android dictation feature being limited to selected countries and languages determines who can use it today. Speaker labelling remains experimental once the count goes above three. Seen from the street, that is the boundary: the model can draw a wide language map while access on your phone and reliability in a crowded room remain narrower.[1]
A small test for trust
For someone using the feature in a class, meeting or interview, the smallest useful test is to open the same short recording beside the polished transcript. Check proper names, live language switches and the speaker's own corrections; with more than three people, verify the speaker labels separately. A stronger outside signal would be a test of the same model version across noisy rooms and languages that counts ordinary word errors separately from cleanups that change meaning.[1]