← Models
Qwen3-ASR-1.7B
1.7B · Streaming speech-to-text
Open-source streaming ASR covering 30 languages and 22 Chinese dialects. Competitive with commercial APIs on multilingual and dialect benchmarks — listed here as a reference model.
Listed for reference. Autoloops does not currently expose a customer API for this model.
Model details
| Developed by | Qwen (Alibaba Cloud) |
| Model family | Qwen3-ASR |
| Architecture | AuT encoder + Qwen3-1.7B decoder |
| Parameters | ~1.7B (decoder) + 300M encoder |
| Languages | 30 languages + 22 Chinese dialects |
| Inference | Offline and streaming |
| Audio types | Speech, singing, songs with BGM |
| License | Apache 2.0 |
Capabilities
- Language identification and transcription across 52 languages and dialects
- Unified offline / streaming inference from a single checkpoint
- Robust under complex acoustics, accents, and challenging text patterns
- Open weights and inference toolkit on Hugging Face and GitHub
- Not currently offered as an Autoloops customer API