The Problem With One Model Doing Everything
Stable Diffusion XL 1.0 made a design bet that sounds almost too simple: split the denoising job between two specialized models instead of training one network to handle every step of the diffusion process. Stability AI calls this an "ensemble of experts" pipeline for latent diffusion, and it's the architectural core that separates SDXL from its predecessors.
Here's the reasoning. A diffusion model learns to progressively remove noise from a latent representation until a clean image emerges — the same fundamental process outlined in our explainer How AI Image Models Turn Random Noise Into a Finished Picture. Early denoising steps deal with broad composition — shapes, layout, global color. Later steps handle fine texture, edge sharpness, small details. Training a single model to excel at both is wasteful. You get a jack-of-all-trades network that's mediocre at each phase.
SDXL sidesteps this by pairing a base model that handles the heavy early denoising with a separate refiner model purpose-built for the final steps. You can use either model alone, or chain them for better output.
The Base Model: Where Generation Starts
The SDXL base model is a diffusion-based text-to-image generative model. It operates in latent space — the same general approach as the original Latent Diffusion Models paper — but it's significantly larger at 3 billion parameters with F32 tensor precision.
What makes it interesting architecturally isn't just scale. It uses two fixed, pretrained text encoders: OpenCLIP-ViT/G and CLIP-ViT/L. These aren't fine-tuned during generation; they're frozen. The model was developed by Stability AI and ships under the CreativeML Open RAIL++-M license.
The base model generates what the model card calls "(noisy) latents" — a compressed representation of the target image that still contains significant noise. In standalone mode, the base model runs the full denoising schedule itself and produces a complete image. Plenty of users stop here. The model card notes that "the base model can be used as a standalone module," and for many workflows that's sufficient.
But the architecture explicitly anticipates handoff.
The Two-Stage Pipeline and SDEdit
When you want the full quality, SDXL hands the noisy latents to the refiner. The official mechanism is SDEdit, the technique from the 2021 paper "Guided Image Synthesis and Editing with Stochastic Differential Equations" (arXiv 2108.01073), sometimes called "img2img."
Here's what happens in sequence. First, the base model generates latents at the desired output size. Then the refiner, described in the model card as "a specialized high-resolution model", applies SDEdit to those latents using the same text prompt. The refiner picks up where the base model left off and handles the final denoising steps.
The model card is honest about the cost: "This technique is slightly slower than the first one, as it requires more function evaluations." You're paying compute for quality. That tradeoff is baked into the design.
The user-preference evaluation chart referenced in the model card shows SDXL with refinement outperforming both SDXL 0.9 and Stable Diffusion 1.5 and 2.1. So the two-stage approach isn't vanity, it measurably changes how humans rate the outputs.
The Refiner: Specialized for Late-Stage Denoising
The refiner model lives in its own Hugging Face repository (stabilityai/stable-diffusion-xl-refiner-1.0) and registers on the hub as an Image-to-Image task with the pipeline class StableDiffusionXLImg2ImgPipeline. That naming tells you exactly what it does: it takes an image (or latent) as input and produces a refined version.
In Diffusers, you load it the same way you'd load any pipeline:
pipe = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-refiner-1.0",
dtype=torch.bfloat16,
device_map="cuda"
)
The refiner also supports standalone img2img use, you can feed it any input image and a prompt. The model card example shows "Turn this cat into a dog" as a prompt. But its intended role in the SDXL ensemble is as the second stage, catching the fine-grained denoising that the base model was never optimized to handle.
Both models share the same license, the same underlying latent space, and the same text encoder stack. They're trained to be compatible.
Running It: Diffusers, Optimum, and CPU Offloading
The most direct path to inference is via the Diffusers library:
pip install -U diffusers transformers accelerate
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
device_map="cuda"
)
image = pipe("Astronaut in a jungle, cold color palette, muted colors, detailed, 8k").images[0]
VRAM-constrained? The model card documents pipe.enable_model_cpu_offload() as a drop-in replacement for .to("cuda"). That offloads model components to system RAM between inference steps, trading speed for memory headroom.
For production or edge deployment, Optimum provides Stable Diffusion pipelines compatible with both OpenVINO and ONNX Runtime. You swap StableDiffusionXLPipeline for OVStableDiffusionXLPipeline (Intel) or ORTStableDiffusionXLPipeline (ONNX Runtime) and the rest of the code stays the same. Both support on-the-fly conversion from PyTorch by setting export=True.
The official source code lives at github.com/Stability-AI/generative-models. Clipdrop provides free hosted SDXL inference.
What It Can't Do
The model card includes a limitations section that reads like a confession list, and credit to Stability AI for publishing it plainly:
- The model does not achieve perfect photorealism.
- It cannot render legible text.
- It struggles with compositionality, prompts like "A red cube on top of a blue sphere" break it.
- Faces and people may not generate properly.
- The autoencoding part of the model is lossy.
That last point matters more than people realize. The VAE that compresses the image into latent space and back loses information. The refiner compensates somewhat by handling final denoising at higher effective resolution, but the pipeline never fully recovers what the encoder discarded.
The card also flags bias directly: "While the capabilities of image generation models are impressive, they can also reinforce or exacerbate social biases." That's not a legal disclaimer, it's an engineering acknowledgment that training data shapes outputs, and ensemble architecture doesn't solve that.
Adoption and Ecosystem Scale
The numbers on Hugging Face tell you how much this architecture caught on. The official stabilityai/stable-diffusion-xl-base-1.0 repo shows over 4 million downloads per month, 8,270 likes, and a model tree containing 9,723 adapters, 1,219 finetunes, 7 merges, and 24 quantizations. The refiner has 2,070 likes and its own ecosystem of derived models.
That adapter count is the real story. It means thousands of people took the base architecture and fine-tuned it for specific styles, subjects, and domains, the same efficient fine-tuning and distillation techniques examined in Streamlining Generative AI: Knowledge Distillation and Efficient Fine-Tuning with SD-Tiny and LoRA. The ensemble design held up under that kind of community pressure.
It's also worth zooming out: SDXL is one pillar of a wider release strategy, detailed in Stability AI's Open Model Portfolio Spans More Than Image Generation.
The Design Lesson
SDXL's two-stage split isn't revolutionary in concept, Mixture-of-Experts has been around, and cascaded diffusion predates it. But applying the principle specifically to the denoising timeline, rather than to the model capacity, was a clean insight. You're not routing between experts on input features. You're routing on when in the generation process a model operates.
The base model knows composition. The refiner knows texture. Neither is great at both. Together, they beat anything Stable Diffusion had shipped before, according to the human preference data Stability published.
It's a modest architectural choice that produced a meaningful quality jump. Sometimes that's all you need.