Qwen3.8-Flash Explained

As of 2026-09-01, Alibaba's Qwen team open-sourced Qwen3.8-Flash-Next on 2026-08-26: a 125B multimodal MoE that activates 6B parameters per token, plus 51B N-gram embeddings. Native context is 262K tokens, extendable to 1M with YaRN. If you need chapters from a long video, paste the link — you do not have to host 125B weights.

Open-sourced · 2026-08-26 125B MoE · 6B active Live demo below

Get the notes, skip the architecture diagram

Paste a long-form video URL — BibiGPT turns it into chapters, a transcript, and follow-up Q&A.

Add BibiGPT as a preferred source on Google See more BibiGPT in Top Stories and AI answers.

Key facts (90-second read)

As of 2026-09-01, Alibaba's Qwen team open-sourced Qwen3.8-Flash-Next on 2026-08-26: a 125B multimodal MoE that activates 6B parameters per token, plus 51B N-gram embeddings, with a native 262K context (YaRN to 1M) and training cost about one-ninth of Qwen3.7-Plus. Last updated 2026-09-01.

Features

What Qwen3.8-Flash-Next actually is

On 2026-08-26 the Qwen team open-sourced Qwen3.8-Flash-Next as an architecture preview of Qwen4. The production API name is Qwen3.8-Flash. Same generation, two surfaces: weights versus a hosted endpoint.

125B MoE, 6B active, plus 51B N-gram embeddings

Qwen states a 125-billion-parameter main model that activates 6 billion parameters per token, with an extra 51 billion N-gram embedding parameters that can sit in system RAM rather than GPU memory.

Native 262K context, YaRN to 1M

The default context window is 262,144 tokens. Qwen says YaRN can stretch that to one million tokens for long files and long conversations. A lecture transcript still needs chapters; a wide window is not a notebook.

Training cost about 1/9 of Qwen3.7-Plus

Qwen says training Qwen3.8-Flash-Next took about one-ninth the cost of Qwen3.7-Plus, with gains on coding and office tasks. That is a vendor training-cost claim, not a user price.

What this means if you work with long video

A cheaper long-context multimodal model helps API teams. It does not, by itself, give a student timestamped notes from a two-hour Bilibili upload. BibiGPT is the product for that second job.

You should not have to host 125B weights

Paste a link, get chapters, a transcript, and follow-up Q&A. BibiGPT supports Qwen3.8-Flash as a model you can use for long-form notes — you do not pick an architecture diagram.

Long context only pays off if you keep the notes

A 90-minute lecture still needs timestamps and an export. One cheap API call that you never file is wasted; a chaptered summary is a document you can search next week.

This is not Qwen 3.6 and not Qwen3-ASR

Qwen 3.6 is an earlier multimodal generation. Qwen3-ASR is the speech-recognition family. This page is the 2026-08-26 Flash-Next / Flash event. One intent, one URL.

5 key changes (90-second read)

Headline facts from Qwen's 2026-08-26 Qwen3.8-Flash / Flash-Next release.

  1. 1

    Open weights on 2026-08-26

    Qwen released Qwen3.8-Flash-Next (and an FP8 build) on Hugging Face and ModelScope the same evening, framed as an architecture preview of Qwen4.

  2. 2

    125B MoE, 6B active, 51B N-gram embeddings

    Qwen documents a 125B main model, 6B parameters activated per token, and 51B extra N-gram embeddings that can sit in system RAM instead of GPU memory.

  3. 3

    262K native context, YaRN to 1M

    Default context is 262,144 tokens. Qwen says YaRN extends that to one million tokens for long files and long threads.

  4. 4

    Training cost ~1/9 of Qwen3.7-Plus

    Qwen claims Flash-Next training cost was about one-ninth of Qwen3.7-Plus, with stronger coding and office-task scores. That is a vendor training-cost figure.

  5. 5

    API list price 1 yuan / 3 yuan per million tokens

    Reuters reported Qwen would charge 1 yuan per million input tokens and 3 yuan per million output tokens on the hosted Qwen3.8-Flash API. Confirm on QwenCloud before you budget.

3 typical scenarios for BibiGPT users

Where a cheaper long-context Qwen model matters — and where a summarizer is the actual product.

You already run Qwen yourself

Use the open weights or the Flash API for your own coding and office stack. Then still paste the finished lecture into BibiGPT so humans get chapters instead of a raw token dump.

You just need notes from a long Chinese video

A two-hour Bilibili upload does not need a Qwen4 architecture brief. Paste the URL, export the transcript, keep asking questions. That is the BibiGPT path.

You are comparing Qwen generations

Qwen 3.6 is the previous multimodal explainer. Qwen3-ASR is speech recognition. This page is Flash / Flash-Next only. One model family event, one URL.

Related BibiGPT tools

Long-video notes workflows that pair with this release.

Sources

Specs on this page come from Qwen's 2026-08-26 release notes and Reuters' same-day report.

  • Qwen open-sourced Qwen3.8-Flash-Next on 2026-08-26 as a 125B multimodal MoE with 6B active parameters, 51B N-gram embeddings, native 262K context (YaRN to 1M), and training cost about 1/9 of Qwen3.7-Plus.

    GitHub — QwenLM/Qwen3.8-Flash-Next ↗
  • Reuters reported the 2026-08-26 launch, the 262,144-token default window (expandable to 1 million), one-ninth the training cost of Qwen3.7-Plus, and API list prices of 1 yuan / 3 yuan per million input / output tokens.

    Reuters — Alibaba's Qwen launches Qwen3.8-Flash ↗
  • IT Home reported the same-day Chinese-language brief: open weights at 23:00 Beijing time on Hugging Face and ModelScope, plus the 1 yuan / 3 yuan API list prices.

    IT Home — Qwen3.8-Flash (Next) ↗

Terms used on this page

What is a mixture-of-experts (MoE) model?

A mixture-of-experts model keeps a large pool of parameters and activates only a subset for each token. Qwen3.8-Flash-Next is documented as 125B total with 6B active per token. The practical effect is more capacity without paying the full dense-model compute bill on every token.

What is a 262K context window good for in video notes?

A 262,144-token window can hold a long transcript plus instructions in one call. It does not, by itself, produce chapters, timestamps, or an export you can search next week. Those artifacts are the job of a video-notes product.

Loved by creators, students & researchers

Why people use BibiGPT to turn videos into text every day.

Trusted by 50,000+ users worldwide

★★★★★

“I paste a link and get clean captions in seconds — it saves me hours of retyping every single week.”

Maya R.

Content Creator · Repurposes short videos

★★★★★

“Exporting the transcript lets me review new words at my own pace instead of pausing the video constantly.”

Daniel K.

Language Learner · Studies with real videos

★★★★★

“Accurate, timestamped text I can quote directly. It has quietly become part of my daily workflow.”

Priya S.

Researcher · Cites public talks

Frequently Asked Questions

Ask us anything!

Popular guides

Skip the 125B checkpoint. Get the notes.

Paste a long YouTube, Bilibili, or podcast link. BibiGPT returns chapters, a transcript, and follow-up Q&A. Last updated 2026-09-01.