Standard & Paid Price
Standard Price
| Meter | Standard Price |
|---|---|
| Input price | $0.1260 / audio hour |
Billing details
Description
Alibaba Qwen3-ASR-Flash — speech recognition built on the Qwen3 foundation model, which automatically determines the spoken language and transcribes 11 languages; the Qwen3-ASR line documents Mandarin alongside Sichuanese, Min Nan, Wu, and Cantonese, and multiple English accents. Accepts aac, amr, avi, aiff, flac, flv, mkv, mp3, mpeg, ogg, opus, wav, webm, wma, and wmv up to 5 minutes and 10 MB per request; pcm input must be 16 kHz and other formats are resampled to 16 kHz before recognition.
Capabilities
- Tasks
- transcription
API endpoints
POST /v1/audio/transcriptions(audio-transcription)