20B MMDiT text-to-image. Best-in-class text rendering, English and Chinese.
Qwen-Image is a 20B Multimodal Diffusion Transformer (MMDiT) from the Qwen team, paired with a Qwen2.5-VL-7B text encoder. Its standout capability is text rendering inside the image — English and especially Chinese typography, where it leads the open models by a wide margin. Apache-2.0, fully commercial.
true_cfg_scale, not guidance_scale. Supply a negative prompt and it runs at the pipeline default (4.0); leave it empty and guidance is off.