Python Speech-to-Text (whisper-1)
Speech-to-Text API Guide#
Overview#
The Audio API provides two endpoints:transcriptions: transcribe audio to text (in the original language)
translations: translate audio into English text
Currently available model: whisper-1 (refer to GET /v1/models for the authoritative list).Formats: mp3, mp4, mpeg, mpga, m4a, wav, webm
1. Transcription#
2. Translation into English#
3. Word-Level Timestamps#
4. Handling Files Larger Than 25MB#
Split the file first with PyDub:Prompt Tips#
Correct the recognition of proper nouns
Control punctuation and Simplified/Traditional Chinese
Last verified/modified: 2026-09-28 (Base URL changed to api.crazyrouter.com; whisper-1 is in the catalog; examples not run live)
Modified at 2026-09-28 12:56:05