Transcribe speech
Transcribe Kazakh/Russian audio, including speakers who switch language mid-sentence. Send audio as a raw request body (naming the container with an x-filename header), as a multipart file, or pass ?url= and let the server fetch it. Billed on audio duration at 2.5 credits per second, from the same credit balance as text-to-speech.
Authorizations
Headers
Filename hint; only the extension matters, so the server knows the container. Defaults to audio.wav.
"call.wav"
Query Parameters
Public URL of the audio. Use instead of sending a body.
Body
The body is of type file.
Response
The transcript, with timing and the credits charged.
The transcript. Lowercase and unpunctuated — the model's vocabulary contains neither, so nothing is being stripped.
"сәлеметсіз бе бүгін ауа райы өте жақсы"
"seta-kk-ru-v2"
Duration of the submitted audio. This is the billed quantity.
2.84
Credits billed: ceil(audio_seconds × 2.5).
8
0.153
Real-time factor: processing_seconds / audio_seconds. Lower is faster.
0.0539
How many times faster than real time this ran.
18.5
How many pieces the audio was split into at pause boundaries.
1
Sample rate of the submitted audio, before resampling to 16 kHz.
24000