Qwen3-ASR × BibiGPT
On 2026-01-30 the Qwen team at Alibaba Cloud open-sourced Qwen3-ASR — a family of speech recognition models covering 52 languages and dialects, shipped in a 1.7B and a 0.6B size, alongside a non-autoregressive forced-alignment model that lines text up with audio in 11 languages. The 1.7B is state of the art among open ASR models and competitive with the strongest commercial APIs; the 0.6B trades a little accuracy for roughly 2000× throughput at a concurrency of 128. What this means in practice: accurate transcripts of dialect-heavy, music-heavy, or accented audio stopped being a paid-API-only capability. BibiGPT turns that same raw material into something you can actually use — timestamped transcripts, summaries, and searchable notes from a single link.
Key facts (90-second read)
On 2026-01-30 the Qwen team at Alibaba Cloud open-sourced Qwen3-ASR — speech recognition models in 1.7B and 0.6B sizes covering 52 languages and dialects (30 languages plus 22 Chinese dialects, and accented English), with language detection and timestamp prediction built in, plus a non-autoregressive forced-alignment model for 11 languages. The 1.7B is state of the art among open ASR models and competitive with the strongest commercial APIs; the 0.6B reaches roughly 2000× throughput at a concurrency of 128. The practical takeaway: accurate transcription of dialect-heavy, accented, and musical audio is no longer gated behind a paid API. Getting the words out is the easy half — BibiGPT covers the other half, turning any link into a timestamped transcript, a summary, and searchable notes.
Features
What is Qwen3-ASR?
Open-sourced on 2026-01-30 by the Qwen team at Alibaba Cloud, Qwen3-ASR is a family of all-in-one speech recognition models. Two sizes ship — Qwen3-ASR-1.7B and Qwen3-ASR-0.6B — both doing language identification and transcription across 52 languages and dialects, with timestamp prediction built in. A separate non-autoregressive forced-alignment model lines existing text up with audio in 11 languages.
52 languages and dialects
Coverage spans 30 languages plus 22 Chinese dialects, and English spoken with accents from multiple countries and regions. Dialect and accent coverage is where most open ASR models fall apart, and it is exactly what a lot of real recordings sound like.
Two sizes, two trade-offs
The 1.7B is state of the art among open ASR models and competitive with the strongest commercial APIs. The 0.6B keeps accuracy respectable while reaching roughly 2000× throughput at a concurrency of 128 — built for volume rather than a single perfect file.
Speech, music, and song — plus timestamps
The models handle speech, music, and sung audio, detect the language on their own, and predict timestamps. Timestamps are the difference between a wall of text and a transcript you can jump around in.
Why an open ASR release matters if you work with video
A transcript is the raw material for everything else — summaries, search, translation, clips, notes. When accurate multilingual transcription becomes open and cheap, the bottleneck moves from getting the words to doing something with them. That second half is where BibiGPT lives.
Dialect-heavy audio stops being a dead end
Recordings in regional Chinese dialects, or English spoken with a strong accent, have long been the files that come back as garbage. Broad dialect and accent coverage in an open model changes what is worth transcribing at all.
A transcript is only step one
Raw text out of any ASR model still needs to be read. BibiGPT takes a link and gives you a timestamped transcript plus an AI summary, a mind map, and notes — so a two-hour recording becomes something you can skim in minutes.
No setup for the people who just want the answer
Running an open model means weights, a GPU, and everything around them. Most people with a lecture to catch up on do not want any of that. BibiGPT keeps it to pasting a link in a browser — the same result, none of the assembly.
5 key facts (90-second read)
Headline facts from the Qwen team's 2026-01-30 open-source release of Qwen3-ASR.
- 1
Open-sourced on 2026-01-30
Alibaba Cloud's Qwen team released Qwen3-ASR as an open-source model family rather than an API-only product.
- 2
52 languages and dialects
30 languages plus 22 Chinese dialects, and English spoken with accents from multiple countries and regions — the coverage most open ASR models lack.
- 3
Two sizes: 1.7B and 0.6B
The 1.7B leads open ASR benchmarks and holds up against the strongest commercial APIs; the 0.6B trades some accuracy for roughly 2000× throughput at a concurrency of 128.
- 4
Language detection and timestamps built in
The models identify the spoken language themselves and predict timestamps, and they cover music and sung audio rather than clean speech only.
- 5
A separate 11-language alignment model
A non-autoregressive forced-alignment model lines existing text up with audio in 11 languages — the piece you need for subtitling and lyric timing.
3 typical scenarios for BibiGPT users
Where broad, accurate speech recognition changes what is worth transcribing at all.
A lecture in a regional dialect
A student records a class where the lecturer slips between Mandarin and a regional dialect. Instead of a transcript full of holes, BibiGPT returns readable timestamped text plus a summary — so revision starts from notes rather than from re-listening.
A podcast archive in several languages
A creator has years of episodes recorded with guests from different countries and accents. Each link becomes a searchable transcript and a structured summary, turning a back-catalogue nobody can search into material worth repurposing.
Subtitles that line up with the audio
An editor has a script and a finished cut but no timing. Timestamped transcription is what makes subtitles land on the right frame — and BibiGPT can produce dual-language subtitles from the same source in one pass.
Loved by creators, students & researchers
Why people use BibiGPT to turn videos into text every day.
Trusted by 50,000+ users worldwide
“I paste a link and get clean captions in seconds — it saves me hours of retyping every single week.”
Maya R.
Content Creator · Repurposes short videos
“Exporting the transcript lets me review new words at my own pace instead of pausing the video constantly.”
Daniel K.
Language Learner · Studies with real videos
“Accurate, timestamped text I can quote directly. It has quietly become part of my daily workflow.”
Priya S.
Researcher · Cites public talks
FAQ'S
Frequently Asked Questions
Ask us anything!
Popular guides
1 Bilibili AI Video Summary Tool: BibiGPT Summarizes 30+ Platforms Instantly (2026)
Best Bilibili AI video summary tool 2026? Paste a link for a free AI recap, mind map, and transcript-style takeaways on 30+ platforms — no login needed.
2 How to Install Skills in DeepSeek Harness: A Hands-On Guide to Teaching dsh to Watch Videos
Install a video-summary skill into DeepSeek Harness in one copy-paste — SKILL.md matches Claude Code. Or skip dsh: paste a Bilibili or YouTube link in the browser and get a timestamped summary.
3 How to Reverse a Video Without an App — Free, In-Browser, No Download (2026)
Reverse any video free with no app and no download, right in your browser. Step-by-step for iPhone, Android, and PC, plus how to reverse audio (2026).
Turn any video or podcast into a transcript you can search
Paste a Bilibili, YouTube, or podcast link — or upload your own file. BibiGPT produces a timestamped transcript, an AI summary, and notes you can search, translate, and send to your note app. No model setup, no GPU, no command line.