Image & Design

Best Stable Diffusion and Open Image Models

Also known as: open source image models

"Stable Diffusion models" now means more than Stability AI's own releases. It is shorthand for any open-weight image model: one whose files you can download, run on your own graphics card, and fine-tune on your own style or products. In 2026 the strongest of these come from Alibaba (Qwen-Image, Z-Image), Black Forest Labs (FLUX.2), Ideogram, HiDream, Tencent and NVIDIA, not only Stability AI.

We ranked 10 models on image quality, licence, hardware needs, ecosystem and editing. Licences matter more than most guides admit: several top models are free to download but banned from commercial use without a paid licence, and one cannot be used at all in the EU or UK. Quality figures come from the Artificial Analysis Text to Image Arena, a blind public vote, as of 25 September 2026.

Quick answer

Z-Image Turbo is the best open image model for most people in September 2026. It is Apache 2.0 (free for commercial use), fits in 16 GB of VRAM, needs only 8 steps and has broad ComfyUI support. Qwen-Image-2512 gives better quality under the same licence if you have a 24 GB card. For editing on modest hardware pick FLUX.2 [klein] 4B; for the best raw quality with open weights, Ideogram 4 and FLUX.2 [dev] lead, but both need a paid licence for commercial use. SDXL is still the model with the most add-ons.

Top picks at a glance

Scoreboard

Scores are out of 10. The overall score is the weighted average of the criteria below.

#ToolOverallImage qualityLicence & commercial useHardware needsEcosystem & toolsEditing & controlPrice fromBest for
1Z-Image Turbo
Alibaba (Tongyi-MAI)
8.78.29.89.38.57.5Free (open weights)
Free tier
Fast, commercial-friendly generation on a 16 GB graphics card
2Qwen-Image-2512
Alibaba (Qwen)
8.59.19.86.58.38.5Free (open weights)
Free tier
The best commercially usable quality on a 24 GB card
3FLUX.2 [klein]
Black Forest Labs
8.37.97.59.38.38.8Free (open weights)
Free tier
Fast generation and image editing on a gaming PC
4HiDream-O1-Image
HiDream.ai
8.38.810.07.55.88.5Free (open weights)
Free tier
Developers who want a permissive all-in-one generate-and-edit model
5SDXL 1.0
Stability AI
7.74.58.89.69.88.0Free (open weights)
Free tier
Low-VRAM PCs and heavy use of community fine-tunes and LoRAs
6FLUX.2 [dev]
Black Forest Labs
7.69.25.05.59.29.3Free (non-commercial)
Free tier
Hobbyists and researchers with big GPUs who want top editing quality
7Ideogram 4
Ideogram
7.39.45.07.56.57.0Free (non-commercial); $300/month commercial
Free tier
Typography, posters and design work where text must be right
8Stable Diffusion 3.5 Large
Stability AI
7.16.57.57.87.06.8Free (under $1M revenue)
Free tier
Small businesses that want Stability AI's current model family
9NVIDIA Cosmos3-Super-Text2Image
NVIDIA
6.89.09.52.55.06.0Free (open weights)
Free tier
Companies with data-centre GPUs that need a permissive high-quality model
10HunyuanImage 3.0
Tencent
6.08.44.52.56.08.0Free (community licence)
Free tier
Research teams outside the EU, UK and South Korea with multi-GPU servers

Expert reviews

#1 · Fast, commercial-friendly generation on a 16 GB graphics card

Z-Image Turbo

by Alibaba (Tongyi-MAI) · Open source · Free (open weights)
8.7/10

Z-Image Turbo is the model we would install first. Alibaba's Tongyi-MAI lab released it in November 2025 as a distilled (sped-up) version of its 6B-parameter Z-Image model. It needs only 8 sampling steps, and the team says it fits within 16 GB of VRAM on consumer cards and excels at photorealism and English and Chinese text.

The licence is the big advantage. It is Apache 2.0, so you can use it for client work, products and fine-tunes with no revenue cap and no territory limits. The community has adopted it fast: ComfyUI has an official tutorial, ComfyUI's repackaged files show about 7.9 million Hugging Face downloads in the past month, and Alibaba's PAI team has published a ControlNet Union model for structural control. The undistilled Z-Image base model, released in January 2026, is the better starting point for training LoRAs.

Quality is good rather than top tier: it sits at #75 on the Artificial Analysis arena, level with FLUX.2 [klein] 9B and behind Qwen-Image-2512. The model card lists a Z-Image-Edit variant, but we found no public weights for it.

Pick it if you want a fast, legal-to-sell default on a mid-range GPU. Skip it if you need the very best quality and have 24 GB or more.

Score breakdown

Image quality8.2
Licence & commercial use9.8
Hardware needs9.3
Ecosystem & tools8.5
Editing & control7.5

Key facts

Pricing
Free (open weights) (Apache 2.0 licence, so commercial use is allowed. Hosted APIs listed by Artificial Analysis from about $5 per 1,000 images.)
Free option
Yes
Platforms
Local GPU, ComfyUI, Diffusers, Hosted APIs
Size
6B parameters
Licence
Apache 2.0
VRAM
Fits 16 GB consumer cards (vendor claim)
Arena rank
#75, Elo 940 (Artificial Analysis, 25 Sep 2026)

What we like

  • Apache 2.0: commercial use with no revenue cap
  • Runs on 16 GB cards in 8 steps
  • Strong photorealism and bilingual text
  • Fast-growing ComfyUI and LoRA support

Watch out for

  • Quality behind Qwen-Image-2512, FLUX.2 [dev] and Ideogram 4
  • Z-Image-Edit weights not public yet
  • Distilled model is harder to fine-tune than the base
#2 · The best commercially usable quality on a 24 GB card

Qwen-Image-2512

by Alibaba (Qwen) · Open source · Free (open weights)
8.5/10

Qwen-Image-2512 is the highest-quality open model you can use commercially without paying anyone. It is the December 2025 update of Alibaba's 20B-parameter Qwen-Image, and it sits at #42 on the Artificial Analysis arena with an Elo of 998, just 2 points behind FLUX.2 [dev] and ahead of every other Apache or MIT model. Alibaba says the update makes people look less "AI-generated", adds finer natural detail and improves text rendering.

The cost is hardware. ComfyUI's documentation lists the fp8 file at 20.4 GB and tests it on a 24 GB RTX 4090D, where one image takes about 70 to 95 seconds; an 8-step Lightning LoRA roughly halves that. On 16 GB cards you will need smaller quantised versions.

Its sibling Qwen-Image-Edit-2511, also Apache 2.0, handles instruction-based editing, so you get a full open generate-and-edit pair.

Pick it if you have a 24 GB GPU and need images you can sell. Skip it if your card has 12 GB or less; start with Z-Image Turbo or FLUX.2 [klein] 4B.

Score breakdown

Image quality9.1
Licence & commercial use9.8
Hardware needs6.5
Ecosystem & tools8.3
Editing & control8.5

Key facts

Pricing
Free (open weights) (Apache 2.0. Hosted APIs listed by Artificial Analysis from about $20 per 1,000 images.)
Free option
Yes
Platforms
Local GPU, ComfyUI, Diffusers, Hosted APIs
Size
20B parameters
Licence
Apache 2.0
Arena rank
#42, Elo 998 (Artificial Analysis, 25 Sep 2026)
File size
20.4 GB in fp8; 40.9 GB in bf16 (Qwen-Image, per ComfyUI docs)

What we like

  • Best arena score of any Apache or MIT model
  • Strong text rendering and realism
  • Apache 2.0 editing sibling (Qwen-Image-Edit-2511)
  • Native ComfyUI support and speed-up LoRAs

Watch out for

  • 20B parameters: needs about 24 GB even in fp8
  • Slow on consumer cards without a speed-up LoRA
  • Fewer community style LoRAs than SDXL or FLUX
#3 · Fast generation and image editing on a gaming PC

FLUX.2 [klein]

by Black Forest Labs · Open source · Free (open weights)
8.3/10

FLUX.2 [klein] is Black Forest Labs' small, fast FLUX.2, released in January 2026. One model does text-to-image, image editing and multi-reference work (combining several input images), and BFL says it can generate or edit in under half a second on modern hardware. The 4B version runs on cards like an RTX 3090 or 4070 with about 13 GB of VRAM, and BFL has since shipped it on-device on ASUS ProArt laptops.

Read the licence carefully, because the two sizes differ. The 4B model is Apache 2.0, free for commercial use. The 9B model uses the FLUX Non-Commercial License, so paid work needs a licence from BFL. The 9B is clearly better: it ties Z-Image Turbo at #76 on the Artificial Analysis arena, while the 4B sits at #117.

ComfyUI supports both, and BFL offers FP8 and NVFP4 versions that it says run up to 2.7 times faster.

Pick it if you want quick edits and generation on a 12 to 16 GB card, or an Apache 2.0 model you can build into a product. Skip it if raw image quality matters most; Qwen-Image-2512 and FLUX.2 [dev] are clearly better.

Score breakdown

Image quality7.9
Licence & commercial use7.5
Hardware needs9.3
Ecosystem & tools8.3
Editing & control8.8

Key facts

Pricing
Free (open weights) (4B model: Apache 2.0, commercial use allowed. 9B model: FLUX Non-Commercial License, so commercial use needs a paid licence from Black Forest Labs. Hosted 9B from about $15 per 1,000 images (Artificial Analysis listing).)
Free option
Yes
Platforms
Local GPU, ComfyUI, Diffusers, Hosted APIs
Sizes
4B (Apache 2.0) and 9B (non-commercial)
VRAM
About 13 GB for the 4B model (vendor figure)
Arena rank
9B #76 (Elo 940); 4B #117 (Elo 864)
Released
15 January 2026

What we like

  • Generation, editing and multi-reference in one small model
  • 4B version is Apache 2.0 and needs about 13 GB
  • Very fast: under half a second per image (vendor claim)

Watch out for

  • The better 9B version is non-commercial
  • 4B quality is well below the leaders
  • Two licences in one family cause confusion
#4 · Developers who want a permissive all-in-one generate-and-edit model

HiDream-O1-Image

by HiDream.ai · Open source · Free (open weights)
8.3/10

HiDream-O1-Image, released in May 2026, takes an unusual approach. It is a single 8B-parameter transformer that works directly on pixels, with no separate VAE (image compressor) or text encoder. The same model does text-to-image, long multilingual text rendering, instruction editing and subject-driven personalisation (keeping a person or product consistent across new scenes) at up to 2,048 x 2,048. Recent updates added layout and skeleton (pose) conditioning.

The MIT licence is the most permissive in this ranking: no revenue cap, no territory limits and no attribution beyond keeping the licence notice.

Quality is strong for its size. The full model sits at #60 on the Artificial Analysis arena (Elo 979), ahead of Z-Image Turbo. HiDream said a Dev version debuted at #8 in May; leaderboards move fast.

The weakness is tooling. It ships with its own inference scripts built on the Transformers library rather than native ComfyUI nodes, and its community is small (about 6,400 Hugging Face downloads in the past month versus hundreds of thousands for Z-Image).

Pick it if you are a developer who wants one permissive model for generation and editing. Skip it if you want a click-and-go ComfyUI workflow.

Score breakdown

Image quality8.8
Licence & commercial use10.0
Hardware needs7.5
Ecosystem & tools5.8
Editing & control8.5

Key facts

Pricing
Free (open weights) (MIT licence, the most permissive here. Undistilled (50 steps) and Dev (28 steps) versions.)
Free option
Yes
Platforms
Local GPU, Transformers, Hugging Face Spaces
Size
About 8B parameters
Licence
MIT
Arena rank
#60, Elo 979 (Artificial Analysis, 25 Sep 2026)
Resolution
Up to 2,048 x 2,048

What we like

  • MIT licence: the most permissive here
  • Generation, editing, personalisation and pose control in one model
  • High quality for 8B parameters, up to 2K resolution

Watch out for

  • Small community and few LoRAs so far
  • Own scripts rather than mature ComfyUI support
  • Undistilled model needs 50 steps
#5 · Low-VRAM PCs and heavy use of community fine-tunes and LoRAs

SDXL 1.0

by Stability AI · Open source · Free (open weights)
7.7/10

SDXL is three years old and its base model now ranks near the bottom of the Artificial Analysis arena (#156, Elo 677). So why is it fifth? Because almost nobody uses the base model any more. SDXL is the foundation for a huge library of community fine-tunes, LoRAs and control add-ons, and it runs on 8 GB graphics cards that newer models cannot use.

Every major local app supports it: ComfyUI, Forge, InvokeAI and SwarmUI. Its licence, CreativeML OpenRAIL++-M, allows commercial use as long as you avoid the listed harmful uses. And it is still heavily used: the base repository had about 3.6 million Hugging Face downloads in the past month.

The limits are real. Out of the box it follows complex prompts poorly, struggles with text in images and needs more tries than 2026 models. Fine-tunes close much of that gap for specific styles, not for general prompt following.

Pick it if you have an 8 to 12 GB card, or you rely on a specific SDXL fine-tune or LoRA library. Skip it if you are starting fresh with a 16 GB card; Z-Image Turbo is better in almost every way.

Score breakdown

Image quality4.5
Licence & commercial use8.8
Hardware needs9.6
Ecosystem & tools9.8
Editing & control8.0

Key facts

Pricing
Free (open weights) (CreativeML OpenRAIL++-M licence: commercial use is allowed, with a list of banned uses. Hosted from about $9 per 1,000 images (Artificial Analysis listing).)
Free option
Yes
Platforms
Local GPU, ComfyUI, Forge, InvokeAI, Hosted APIs
Size
3.5B base model (6.6B with refiner)
VRAM
Works on 8 GB consumer GPUs (Stability AI)
Arena rank
#156, Elo 677 (base model)
Downloads
About 3.6 million in the past month on Hugging Face

What we like

  • Runs on 8 GB GPUs
  • Largest library of fine-tunes, LoRAs and control add-ons
  • Supported by every major local app
  • Commercial use allowed under OpenRAIL++-M

Watch out for

  • Base quality far behind 2026 models
  • Weak text rendering and prompt following
  • RAIL licence carries use restrictions
#6 · Hobbyists and researchers with big GPUs who want top editing quality

FLUX.2 [dev]

by Black Forest Labs · Free · Free (non-commercial)
7.6/10

FLUX.2 [dev] is the most capable open model for editing, and the second-best open model for raw quality. Released in November 2025, this 32B-parameter model does text-to-image and editing with up to 10 reference images in one checkpoint, at up to 4 megapixels. It ranks #40 on the Artificial Analysis arena (Elo 1000), and FLUX has one of the largest LoRA and tool ecosystems, with ComfyUI support from launch day.

Two things hold it back. First, hardware: NVIDIA says the full model needs about 90 GB of VRAM to load, and its FP8 version cuts that by 40%; on a 24 GB card you rely on ComfyUI offloading parts to system RAM, which is slower. Second, the licence. The FLUX [dev] Non-Commercial License covers personal, research and testing use only; using it for paid work or in a product needs a BFL licence. BFL claims no ownership of the images, but making them for a client counts as commercial use.

BFL unveiled FLUX 3 in July 2026 but had not released open FLUX 3 image weights by 25 September.

Pick it if you have a large GPU and non-commercial projects. Skip it if you need to sell the results without a licence.

Score breakdown

Image quality9.2
Licence & commercial use5.0
Hardware needs5.5
Ecosystem & tools9.2
Editing & control9.3

Key facts

Pricing
Free (non-commercial) (FLUX [dev] Non-Commercial License v2.0: free for personal, research and testing use only. Any revenue-generating or production use needs a commercial licence from Black Forest Labs (price on request). Hosted from about $12 per 1,000 images (Artificial Analysis listing).)
Free option
Yes
Platforms
Local GPU, ComfyUI, Diffusers, Hosted APIs
Size
32B parameters
Arena rank
#40, Elo 1000 (Artificial Analysis, 25 Sep 2026)
Editing
Up to 10 reference images; edits up to 4 megapixels
VRAM
About 90 GB to load fully; FP8 cuts that by 40% (NVIDIA)

What we like

  • Best open model for multi-reference editing
  • Top-tier open quality (arena Elo 1000)
  • Huge ecosystem and day-one ComfyUI support

Watch out for

  • Non-commercial licence; paid licence price not published
  • 32B parameters: heavy even in FP8
  • Slow on consumer GPUs with offloading
#7 · Typography, posters and design work where text must be right

Ideogram 4

by Ideogram · Freemium · Free (non-commercial); $300/month commercial
7.3/10

Ideogram 4 is the best open-weight image model on quality. Ideogram released it in June 2026 as its first open model: a 9.3B-parameter diffusion transformer trained from scratch, using the Qwen3-VL-8B vision-language model to understand prompts. On the Artificial Analysis arena its Quality setting ranks #35 with an Elo of 1010, the highest of any open-weights model, and Ideogram's heritage shows in text rendering: posters, logos and signs come out readable.

It also brings design controls that other open models lack: structured JSON prompts, bounding boxes to place elements, colour palettes and native 2K output with aspect ratios up to 6:1.

The licence is the catch. The free weights are for research, evaluation and personal projects only. Commercial use needs the self-serve licence at $300/month for 10,000 images, or Ideogram's own app and API. The weights are gated, and the reference script by default calls Ideogram's hosted prompt-expansion service, so a fully offline setup takes extra work. Community ComfyUI support is young.

Pick it if you want the best open quality and text, for personal use. Skip it if you need free commercial use; choose Qwen-Image-2512.

Score breakdown

Image quality9.4
Licence & commercial use5.0
Hardware needs7.5
Ecosystem & tools6.5
Editing & control7.0

Key facts

Pricing
Free (non-commercial); $300/month commercial (Weights are free under the Ideogram Non-Commercial Model Agreement for research, evaluation and personal projects. A self-serve commercial licence costs $300/month for 10,000 images a month (up to 100,000 on higher tiers) and allows self-hosting but not reselling access.)
Free option
Yes
Platforms
Local GPU (CUDA), Diffusers, Hosted APIs
Released
3 June 2026 (gated on Hugging Face)
Size
9.3B, published in nf4 and fp8 versions
Arena rank
#35, Elo 1010: the top open-weights model (Artificial Analysis)
Controls
JSON prompts, bounding-box layout and colour palettes

What we like

  • Highest arena score of any open-weights model
  • Best open text rendering and design controls
  • Quantised 9.3B model is lighter than FLUX.2 [dev]

Watch out for

  • Non-commercial weights; commercial licence $300/month
  • Gated download and new, small tool ecosystem
  • Reference setup relies on an online prompt service by default
#8 · Small businesses that want Stability AI's current model family

Stable Diffusion 3.5 Large

by Stability AI · Free · Free (under $1M revenue)
7.1/10

Stable Diffusion 3.5 is Stability AI's newest open model family, released in October 2024: Large (8.1B parameters), Large Turbo (a 4-step distilled version) and Medium (2.5B, which Stability says needs 9.9 GB of VRAM excluding text encoders). Stability designed it to be easy to fine-tune, and it runs in ComfyUI and Diffusers.

It has aged quickly. SD 3.5 Large sits at #125 on the Artificial Analysis arena (Elo 839), roughly level with FLUX.1 [dev] and well behind the 2025 and 2026 models above it. Its community never grew as large as SDXL's or FLUX's (about 99,000 Hugging Face downloads in the past month for Large, against 3.6 million for SDXL), so there are fewer LoRAs and fine-tunes.

The licence sits in the middle. The Stability AI Community License is free for commercial use only while you and your affiliates earn under $1 million a year; above that you must get an enterprise licence from Stability.

Pick it if you already have SD 3.5 fine-tunes or pipelines. Skip it if you are choosing a model today: Z-Image Turbo and Qwen-Image-2512 are better and use Apache 2.0.

Score breakdown

Image quality6.5
Licence & commercial use7.5
Hardware needs7.8
Ecosystem & tools7.0
Editing & control6.8

Key facts

Pricing
Free (under $1M revenue) (Stability AI Community License: free for non-commercial use and for commercial use by individuals and organisations with under $1 million in annual revenue. Above that you need an enterprise licence. Hosted from about $65 per 1,000 images (Artificial Analysis listing).)
Free option
Yes
Platforms
Local GPU, ComfyUI, Diffusers, Hosted APIs
Sizes
Large 8.1B; Large Turbo (4 steps); Medium 2.5B
VRAM
Medium needs 9.9 GB, excluding text encoders (Stability AI)
Arena rank
Large #125, Elo 839
Released
22 October 2024

What we like

  • Medium version runs on about 10 GB cards
  • Free commercial use under $1 million revenue
  • Designed for fine-tuning

Watch out for

  • Quality well behind 2026 open models
  • Revenue cap on commercial use
  • Smaller community than SDXL or FLUX
#9 · Companies with data-centre GPUs that need a permissive high-quality model

NVIDIA Cosmos3-Super-Text2Image

by NVIDIA · Open source · Free (open weights)
6.8/10

NVIDIA's Cosmos3-Super-Text2Image, released on 31 May 2026, is part of its Cosmos 3 family of "world models" for robotics and physical AI, but it is also a very good general image generator. With a prompt-rewriting (agentic) mode it ranks #46 on the Artificial Analysis arena with an Elo of 995, close to FLUX.2 [dev] and Qwen-Image-2512.

Its licence is excellent. OpenMDW 1.1 allows commercial use, places no restrictions on outputs and has no revenue cap or territory limit. That makes it one of the few near-frontier open models a large company can deploy without a separate deal.

But it is not for home users. At 64B parameters, NVIDIA's recommended serving setup is an 8x H100 server node, and the diffusers example was tested on GB200. There is no mainstream ComfyUI workflow, and the tooling is aimed at physical-AI developers.

Pick it if you run a data centre or rent large GPU servers and need a permissive, high-quality model. Skip it if you have a single consumer GPU; choose Qwen-Image-2512.

Score breakdown

Image quality9.0
Licence & commercial use9.5
Hardware needs2.5
Ecosystem & tools5.0
Editing & control6.0

Key facts

Pricing
Free (open weights) (OpenMDW 1.1 licence: commercial use allowed and no restrictions on outputs. Needs multi-GPU data-centre hardware.)
Free option
Yes
Platforms
Data-centre GPUs, vLLM-Omni, Diffusers
Size
64B parameters
Licence
OpenMDW 1.1 (commercial use allowed)
Arena rank
#46, Elo 995 (agentic mode)
Hardware
Serving recipe written for an 8x H100 node

What we like

  • Near-top open quality (arena Elo 995)
  • Permissive OpenMDW licence with no output restrictions
  • Backed by NVIDIA's serving stack

Watch out for

  • Needs multi-GPU data-centre hardware
  • Little community or ComfyUI support
  • Built mainly for physical-AI use cases
#10 · Research teams outside the EU, UK and South Korea with multi-GPU servers

HunyuanImage 3.0

by Tencent · Free · Free (community licence)
6.0/10

HunyuanImage 3.0 is Tencent's largest image model and, Tencent says, the largest open image-generation mixture-of-experts model: 80B parameters in 64 experts, with 13B active for each token (a mixture-of-experts model only uses part of itself for each step). The Instruct version, released in January 2026, adds prompt reasoning and image-to-image editing, and ranks #66 on the Artificial Analysis arena (Elo 963).

Two problems push it to the bottom. The hardware is data-centre class: Tencent's demo is configured for four GPUs by default. And the licence is the most restrictive here. The Tencent Hunyuan Community License does not apply in the European Union, the United Kingdom or South Korea, so you cannot use the model or its outputs there, and services with over 100 million monthly users need a separate licence.

The older HunyuanImage 2.1 (#101, Elo 886) is lighter but carries the same territory limits.

Pick it if you are a research team outside the excluded regions with large GPUs. Skip it if you are in Europe or the UK, or want a model for one PC.

Score breakdown

Image quality8.4
Licence & commercial use4.5
Hardware needs2.5
Ecosystem & tools6.0
Editing & control8.0

Key facts

Pricing
Free (community licence) (Tencent Hunyuan Community License: royalty-free, but it does not apply in the EU, UK or South Korea, and services with over 100 million monthly users must request a licence. Hosted from about $90–$100 per 1,000 images (Artificial Analysis listing).)
Free option
Yes
Platforms
Multi-GPU servers, Hosted APIs
Size
80B mixture-of-experts, 13B active per token
Arena rank
Instruct #66 (Elo 963); base #71 (Elo 945)
Territory
Licence excludes the EU, UK and South Korea
Instruct version
Released 26 January 2026, adds reasoning and image-to-image editing

What we like

  • Strong prompt reasoning and editing in the Instruct version
  • Good quality (arena Elo 963 for Instruct)
  • Royalty-free for most users in permitted regions

Watch out for

  • Licence excludes the EU, UK and South Korea
  • 80B model needs a multi-GPU server
  • Lower arena scores than smaller, permissive models

How we scored these tools

Each tool is scored 0–10 on the criteria below, using public evidence: independent benchmarks, vendor documentation and pricing pages, aggregate user ratings and reputable reviews. The overall score is the weighted average. Nobody pays to be listed. Read our full methodology.

CriterionWeightWhat we look at
Image quality30%Prompt following, realism, detail and text rendering, anchored to the Artificial Analysis Text to Image Arena Elo.
Licence & commercial use20%Whether you can use the model for paid work, in a product, and in every country, and what a commercial licence costs.
Hardware needs20%How much GPU memory (VRAM) it needs and whether it runs on a normal gaming or creator PC.
Ecosystem & tools15%Support in ComfyUI and other apps, community LoRAs and fine-tunes, quantised versions and hosted APIs.
Editing & control15%Image editing, reference images, layout or pose control and inpainting in the model or its official siblings.

Licences at a glance: who can use what commercially

Model Licence Commercial use Parameters Arena Elo
Z-Image Turbo Apache 2.0 Yes 6B 940
Qwen-Image-2512 Apache 2.0 Yes 20B 998
FLUX.2 [klein] 4B Apache 2.0 Yes 4B 864
FLUX.2 [klein] 9B FLUX Non-Commercial Paid licence needed 9B 940
HiDream-O1-Image MIT Yes 8B 979
SDXL 1.0 CreativeML OpenRAIL++-M Yes, with use restrictions 3.5B 677
FLUX.2 [dev] FLUX [dev] Non-Commercial v2.0 Paid licence needed 32B 1000
Ideogram 4 Ideogram Non-Commercial $300/month self-serve licence 9.3B 1010
SD 3.5 Large Stability AI Community Yes, under $1M annual revenue 8.1B 839
Cosmos3-Super OpenMDW 1.1 Yes 64B 995
HunyuanImage 3.0 Tencent Hunyuan Community Not in the EU, UK or South Korea 80B (13B active) 945

Elo figures are from the Artificial Analysis Text to Image Arena on 25 September 2026. "Non-commercial" licences usually cover the model, not just the images: under the FLUX [dev] licence, for example, generating images for a paying client is a commercial use of the model even though BFL claims no ownership of the pictures. This is a summary, not legal advice; read the licence before you build a business on a model.

What hardware you need

The amount of VRAM (memory on your graphics card) decides which models you can run. Rough guide, from vendor and ComfyUI figures:

  • 8 GB: SDXL and its fine-tunes (Stability AI says SDXL works on 8 GB cards).
  • About 10–13 GB: SD 3.5 Medium (9.9 GB excluding text encoders) and FLUX.2 [klein] 4B (about 13 GB).
  • 16 GB: Z-Image Turbo, which Alibaba says fits 16 GB consumer cards.
  • 24 GB: Qwen-Image in fp8 (a 20.4 GB file, tested by ComfyUI on a 24 GB RTX 4090D), and FLUX.2 [dev] in FP8 with offloading to system RAM.
  • Data-centre GPUs: HunyuanImage 3.0 (80B) and Cosmos3-Super (64B, served on an 8x H100 node).

Quantised versions (fp8, nf4, GGUF) shrink models to fit smaller cards at a small quality cost. NVIDIA says its FP8 FLUX.2 cuts VRAM needs by 40%.

See our best GPUs for AI for card recommendations, or rent one from our best GPU cloud providers.

How to run open image models: ComfyUI, Forge and hosted options

  • ComfyUI is the standard. It is a node-based app (you connect boxes into a workflow), GPL-3.0 licensed, with about 135,000 GitHub stars, and it often supports major new models on launch day, as it did for FLUX.2. Its documentation includes ready-made workflows for Qwen-Image, FLUX.2 and Z-Image. The learning curve is steeper than a simple form, but templates help.
  • Forge is a simpler form-based interface derived from AUTOMATIC1111. The original Forge repository has not been updated since July 2025; community forks such as Forge Classic keep it current. It remains popular for SDXL.
  • InvokeAI (Apache 2.0) is built around a canvas for inpainting and editing, and SwarmUI (MIT) puts a simple front end on top of ComfyUI.
  • Hosted APIs: if you lack a GPU, providers such as fal and Replicate run these models per image. Artificial Analysis lists prices from about $5 per 1,000 images for Z-Image Turbo to $100 for the largest models.
  • Hugging Face and Civitai are where you download base models, LoRAs and fine-tunes. Check each file's licence: a LoRA inherits the base model's terms.

If you just want great images without setup, a closed tool may suit you better: see our best AI image generators and best free AI image generators.

How good are open models compared with closed ones?

Still clearly behind the best, but closing. On 25 September 2026 the top of the Artificial Analysis Text to Image Arena was OpenAI's GPT Image 2.5 Sunburst (max) at an Elo of 1196, while the best open-weights model, Ideogram 4 Quality, scored 1010 at #35. The best Apache-licensed model, Qwen-Image-2512, scored 998.

What open models give you instead is control: no per-image fees once you own the hardware, no content filters you did not choose, fine-tuning on your own face, products or art style, and privacy, because prompts and images never leave your machine. For product shots, consistent characters and brand styles, a fine-tuned open model often beats a better general model.

How we ranked these models

We scored each model from 0 to 10 on five criteria: image quality (30%), licence and commercial use (20%), hardware needs (20%), ecosystem and tools (15%) and editing and control (15%). The overall score is the weighted average. Image quality scores are anchored to each model's Elo on the Artificial Analysis Text to Image Arena, a blind public preference vote, on 25 September 2026; for model families we used the variant most people run.

We used public sources only: model cards and licence files on Hugging Face and GitHub, vendor announcements, Stability AI and NVIDIA blogs, ComfyUI documentation and GitHub repository data. Download counts are Hugging Face's past-month figures. We did not run our own image tests. Licence summaries are simplified and are not legal advice.

Expert tips
  1. Check the licence before you train a LoRA for client work. A LoRA built on FLUX.2 [dev] or FLUX.2 [klein] 9B inherits their non-commercial terms, while one built on Z-Image or Qwen-Image stays Apache 2.0.
  2. Train LoRAs on the undistilled base model (Z-Image rather than Z-Image Turbo), then use them with the fast distilled version. Distilled models are tuned for speed and fine-tune less predictably.
  3. Start from the official ComfyUI template for a new model instead of a random workflow from social media. Templates pin the right text encoder and VAE files, which is where most first-run errors come from.
  4. On a 16 GB card, try fp8 or GGUF quantised files before buying new hardware. They often run models that need 24 GB at full precision with only a small quality loss.
  5. Keep a folder of 10 test prompts (a face, a hand, a sign with text, a product shot) and run every new model on the same seeds. You will learn more in ten minutes than from any leaderboard.

Jargon explained

Open weights
The trained model files are published so you can download and run them yourself. It does not always mean open source: the licence may still ban commercial use.
VRAM
The memory on your graphics card. The whole model (or most of it) must fit there to run at a reasonable speed.
LoRA
A small add-on file that teaches a base model a new style, character or product without retraining the whole model.
Quantisation (fp8, nf4, GGUF)
Storing a model's numbers with less precision so it takes less memory and runs on smaller GPUs, usually with a slight drop in quality.
Distilled model
A faster version of a model trained to produce similar images in far fewer steps, for example 8 instead of 50.
Elo (arena score)
A rating calculated from many blind head-to-head votes between models. A higher Elo means people preferred that model's images more often.

Frequently asked questions

What is the best Stable Diffusion model in 2026?

For most people, Z-Image Turbo: it is Apache 2.0 (free for commercial use), fits 16 GB of VRAM and needs only 8 steps. If you have a 24 GB card, Qwen-Image-2512 gives better quality under the same licence. Among Stability AI's own models, SDXL has the biggest ecosystem and SD 3.5 is the newest.

Is Stable Diffusion still worth using?

Stability AI's own models have fallen behind: SD 3.5 Large ranks #125 and SDXL #156 on the Artificial Analysis arena. SDXL is still worth it on 8 GB cards and for its huge library of fine-tunes and LoRAs. For new projects, newer open models such as Z-Image Turbo and Qwen-Image-2512 are better.

Which open image models can I use commercially for free?

Z-Image Turbo, Qwen-Image-2512 and FLUX.2 [klein] 4B (Apache 2.0), HiDream-O1-Image (MIT), Cosmos3-Super (OpenMDW) and SDXL (OpenRAIL++-M, with use restrictions). SD 3.5 is free for commercial use only under $1 million in annual revenue. FLUX.2 [dev], FLUX.2 [klein] 9B and Ideogram 4 need a paid licence, and HunyuanImage cannot be used in the EU, UK or South Korea.

Can I sell images made with FLUX.2 [dev]?

Not without a commercial licence. The FLUX [dev] Non-Commercial License allows personal, research and testing use only, and says revenue-generating use is not a non-commercial purpose. BFL claims no ownership of outputs, but generating images for paid work is commercial use of the model. FLUX.2 [klein] 4B and the older FLUX.1 [schnell] are Apache 2.0 alternatives.

How much VRAM do I need to run open image models?

8 GB runs SDXL. About 13 GB runs FLUX.2 [klein] 4B, 16 GB runs Z-Image Turbo, and 24 GB runs Qwen-Image in fp8 or FLUX.2 [dev] with offloading. The largest models, HunyuanImage 3.0 and Cosmos3-Super, need multi-GPU data-centre servers.

Should I use ComfyUI or Forge?

ComfyUI if you want the newest models and full control: it often supports major releases on launch day and has official workflows for Qwen-Image, FLUX.2 and Z-Image. Forge is simpler for SDXL-era workflows, but the original repository has not been updated since July 2025, so use an active fork.

Are open image models as good as Midjourney or GPT Image?

Not yet on raw quality. The top closed model on Artificial Analysis scores an Elo of 1196, against 1010 for the best open-weights model. Open models win on cost at volume, privacy and fine-tuning. Compare them with closed tools in our best AI image generators ranking.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Text to Image Arena leaderboard (Artificial Analysis)
  2. Z-Image-Turbo model card (Hugging Face)
  3. Z-Image base model (Hugging Face)
  4. Z-Image-Turbo ControlNet Union (Hugging Face)
  5. ComfyUI Z-Image-Turbo tutorial (ComfyUI)
  6. Qwen-Image-2512 model card (Hugging Face)
  7. Qwen-Image-Edit-2511 (Hugging Face)
  8. ComfyUI Qwen-Image tutorial (ComfyUI)
  9. FLUX.2 [klein]: towards interactive visual intelligence (Black Forest Labs)
  10. FLUX.2 announcement (Black Forest Labs)
  11. Black Forest Labs blog (FLUX 3 posts) (Black Forest Labs)
  12. FLUX [dev] Non-Commercial License v2.0 (Black Forest Labs (GitHub))
  13. FLUX.2 [dev] model card (Hugging Face)
  14. FLUX.2 image models optimised for NVIDIA RTX GPUs (NVIDIA)
  15. HiDream-O1-Image model card (Hugging Face)
  16. Ideogram 4 GitHub repository (Ideogram (GitHub))
  17. Ideogram 4 collection (Hugging Face)
  18. Ideogram licensing (Ideogram)
  19. Introducing Stable Diffusion 3.5 (Stability AI)
  20. Stability AI Community License (Stability AI)
  21. SDXL 1.0 announcement (Stability AI)
  22. SDXL base 1.0 model card (Hugging Face)
  23. Cosmos3-Super-Text2Image model card (Hugging Face)
  24. OpenMDW 1.1 licence (OpenMDW)
  25. HunyuanImage 3.0 model card (Hugging Face)
  26. Tencent Hunyuan Community License (Tencent (Hugging Face))
  27. ComfyUI repository (GitHub)
  28. Stable Diffusion WebUI Forge (GitHub)
  29. Forge Classic fork (GitHub)
  30. InvokeAI (GitHub)
  31. SwarmUI (GitHub)
  32. Civitai (Civitai)