Qwen open-sources Image-2.1, a 7B model that generates and edits (with transparency)
Alibaba’s Qwen team released open weights for Qwen-Image-2.1 — a unified image generation and editing model with a 7B visual-generation component, native transparent (RGBA) output, up to 10 reference images, and mask/circle local edits — via Hugging Face, GitHub, and the Qwen blog.
SOFTWARE desk — Image models are becoming product dependencies, not demo toys. A single open 7B checkpoint that generates, edits, and does native transparency is the kind of tool teams actually wire into pipelines — if the outputs hold up under product testing, not just launch screenshots.
Qwen said Qwen-Image-2.1 is open-weight and unified: one lightweight checkpoint for both text-to-image generation and image editing. Open weights means the downloaded numbers that let you run the model on your own machines, not just an API you call over the internet. @Alibaba_Qwen called it “a unified model for both generation and editing, delivering top-tier quality in a lightweight package.” File that unified / open-weight / lightweight picture as Qwen’s. “Top-tier quality” is the company’s slogan, not a desk ranking. This desk did not download the checkpoint or generate a picture.
The visual generation component is 7 billion parameters — 7B — across 32 Single-Stream DiT layers, per the model card and blog. DiT means Diffusion Transformer, the image-making stack. 7B is the visual-generation piece Qwen names, not a claim this desk counted every other weight in the bundle. File those 7B / 32-layer lines as Qwen’s. This desk did not open a model file.
Native transparent image generation is a headline feature. RGBA means red, green, blue, plus alpha — the transparency channel — so the model can output a sticker-style image with no background baked in, edit a transparent layer, or pull a subject out of a regular photograph. File that RGBA / transparent-layer / subject-extract picture as Qwen’s. This desk did not extract a sticker.
Editing, still company: up to 10 reference images in one pass; local edits marked with circles, painted annotations, or a separate mask; and identity preservation for people and products. File those 10-reference / circle-paint-mask / identity lines as Qwen’s claims. This desk did not mark a circle or check a face match.
License, as the model card and LICENSE file have it: Qwen Research License. The grant is for non-commercial research or evaluation only. Commercial use needs a separate license from the company. File those license terms as the company’s statement on the card — not as legal advice and not as a claim this desk cleared a commercial seat.
Same-day ComfyUI support: @Alibaba_Qwen quoted ComfyUI saying Qwen-Image-2.1 is supported, with native 2K generation, editing from up to 10 reference images, and RGBA output. Qwen’s GitHub news line also says ComfyUI natively supports the model from Day 0, with compatible weights at Comfy-Org/Qwen-Image-2.1. File that same-day ComfyUI line as Qwen / ComfyUI. This desk did not open a workflow.
Score claims, still company: the Qwen blog posts a Qwen-Image-Bench chart comparing the model with other open-source and closed-source systems. File that chart as Qwen’s own scoreboard. This desk did not rerun a benchmark. Independent boards such as Artificial Analysis may lag a launch-day card. Do not treat a company “beats closed models” line as independently verified.
Plain English for the rest of the card: open weights = the downloaded numbers you can run yourself. 7B = 7 billion parameters in the visual-generation piece Qwen names. DiT / Diffusion Transformer = the image-making stack. RGBA = red, green, blue, plus alpha (transparency). alpha = the channel that says which pixels are see-through. reference images = photos you feed in so the model can keep a person, product, or scene. mask = a separate picture that marks the exact pixels to change. ComfyUI = a popular node-based tool for running image models locally. Qwen Research License = non-commercial research/evaluation; commercial use needs a separate company license.
PRIMARY here: the Hugging Face model card Qwen/Qwen-Image-2.1, the GitHub repo QwenLM/Qwen-Image-2.1, and the 20 Sep 2026 Qwen blog — Tier A PRIMARY company sources, the original record. @Alibaba_Qwen at 13:06:39Z is the same-day company social announce, not a substitute primary. The ComfyUI quote tweet is same-day company social plus the quoted ComfyUI post — not a second originating model card. The 20 Sep open-weight announce, the unified generation-and-editing picture, the 7B / 32 Single-Stream DiT visual-generation component, native RGBA generation and transparent-layer editing, up to 10 reference images, circle / paint / mask local edits, identity preservation for people and products, the Qwen Research License non-commercial grant, same-day ComfyUI support, and the Qwen-Image-Bench comparison chart are Qwen / Alibaba-attributed. License terms stay company statement. Bench and “top-tier” lines stay company-attributed — not independently tested here. NOT claimed: independently verified leadership over closed-source image models, a commercial license this desk reviewed, that this desk ran the weights, a stock tip, or investment advice. Distinct from the already-filed qwen38-omni-flash, chatgpt-images-2-5, and fotor-agent.
RELATED
On 20 Sep 2026 Alibaba’s Qwen team open-sourced Qwen-Image-2.1. The company PRIMARIES are the Hugging Face model card Qwen/Qwen-Image-2.1, the GitHub repo QwenLM/Qwen-Image-2.1, and the Qwen blog at https://qwen.ai/blog?id=qwen-image-2.1. @Alibaba_Qwen posted the launch the same day at 13:06:39Z. Those company pages are the filing event. These are Qwen / Alibaba words. This desk did not run the model.
