Alibaba’s Qwen team has released Qwen-Image-3.0-Pro and a Standard variant on Qwen Cloud. The company says the Pro model ranks first among Chinese models and second overall on the Arena text-to-image leaderboard.
The specs are aimed at practical image work. The model handles 4.5k-token prompts, renders text at 10-pixel scale, and supports 12 languages. Pricing starts at $0.04 per image for Pro and $0.03 per image for Standard.
What the Release Includes
The headline capability is long-prompt handling. A 4.5k-token prompt is far beyond what most image models accept, which matters for complex compositions, detailed product shots, and multi-element scenes where short prompts lose context.
The 10px text rendering claim is worth attention too. Text in generated images has been a weak point across the industry, and small-scale text is where failures show first. If the model holds up at that scale, it changes what can be generated directly instead of patched in post-production.
Twelve languages of support widens the practical use, since prompts and on-image text are no longer limited to English.
How the Pricing Compares
At $0.04 per image for the Pro tier, the model sits at a price point that makes volume experimentation affordable. The $0.03 Standard tier pushes the cost lower for cases that do not need the top tier.
The Arena ranking is the strongest signal in the announcement, though it is worth noting that leaderboard positions shift as new models land. A second-place global ranking today is a snapshot, not a permanent position.
Who Should Test It
The release sits in a crowded field, but the combination is distinct: long-prompt support, small-scale text rendering, and pricing under five cents per image. Teams doing localized ad creatives, product pages, or multilingual social content are the most obvious fit, since those are the workloads where prompt length and text accuracy actually bite.
The Practical Takeaway
For teams that generate product visuals, marketing images, or localized content across languages, the price point changes the calculation on doing it in-house. At these prices, batch experimentation becomes affordable, and the long-prompt support removes a real bottleneck for complex scenes.
The practical takeaway: if you have been waiting for an image model that handles detailed prompts and renders text cleanly without a premium price tag, this release is worth testing against your own workload. Run your own prompts, check the text rendering at small scale, and compare the output against whatever you use today.
Related Reads
- [Qwen3.8: The 2.4T Parameter Open-Source Release](https://getaibest.com/qwen3-8-open-source-caught-up-to-claude/)
- [DeepSeek V4 Flash Full Release](https://getaibest.com/deepseek-v4-flash-full-release-agent/)
- [AirLLM: Running 70B Models on 4GB of VRAM](https://getaibest.com/airllm-70b-models-4gb-vram-local-ai/)