On 1 October Microsoft released MAI-Transcribe-2-Streaming and MAI-Voice-2.1 and Flash speech models. Transcripts arrive while someone speaks. Vercel separately documents access through AI Gateway, not independent performance measurements.

Possible use: while someone dictates a service request, an application could display relevant guidance progressively. The benefit would be less waiting and a chance to correct misunderstandings earlier. Our scenario does not allow an unfinished sentence to order goods or change a reservation on its own.

What needs to be solved: a partial transcription hypothesis can change. The interface should distinguish provisional text from a confirmed record and require approval for irreversible actions. Model speed is not total call response time: the network, end-of-speech detection, reasoning and playback also matter. Service access does not imply processing without sending audio off-device.

Optimistic horizon: for an existing application, we estimate 2–4 weeks for a limited pilot using its own recordings, local names and background noise. We have not tested Czech or voice cloning; language support and the speaker's consent must be checked before deployment. This is our estimate, not a deadline for a reliable autonomous telephone operator.