2026-10-06
Qwen3.8-Flash-Next, a preview of Qwen4
Qwen3.8-Flash-Next FP8 runs on 4× RTX PRO 6000 at about $1.3/h measured, or in Q4 GGUF on an RTX 5090.
Qwen3.8-Flash-Next is the preview of the Qwen4 family. The catalog has three routes, from cheapest to fastest.
The recipes
- Q4 GGUF on 1× RTX 5090 (llama.cpp/Unsloth): about $1.2/h measured. Multimodal input, text output.
- FP8 on 4× RTX PRO 6000 (vLLM): about $1.3/h measured, with an n-gram table in host RAM to speed up generation.
- FP8 on 4× H200 (vLLM, NVLink): about $22.8/h measured, for maximum throughput.
Prices are the ones measured in the catalog and follow the Vast market: the app shows the real price before every start.
For security work
It's a good base for code review and triage on customer repositories: light enough to keep running for a long time, without sending code to an external API.