Skip to main content
Browse Models

stabilityai

ControlNet SDXL Canny

Released

2023-08-29

Family

Stable Diffusion XL

Type

ControlNet Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · 774 MB · sai_xl_canny_256lora.safetensors

Model Report

Overview

ControlNet SDXL Canny is a generative AI model within the ControlNet 1.1 family, designed to introduce precise structural guidance to image synthesis using Canny edge maps in conjunction with Stable Diffusion. Leveraging edge-detected representations, ControlNet SDXL Canny allows for constrained yet creative image generation, ensuring outputs remain faithful to the shape and contours specified by the provided edge image. The model's implementation, architecture, and datasets have been developed to promote image quality, robustness, and adaptability across a breadth of artistic and technical applications.

ControlNet 1.1 model naming convention diagram

Figure 1. Diagram illustrating the Standard ControlNet Naming Rules (SCNNRs), including the naming structure for ControlNet 1.1 Canny and related models.

Model Architecture and Training Procedure

ControlNet SDXL Canny maintains architectural parity with the original ControlNet 1.0 framework, as outlined in the official documentation. The model integrates with the Stable Diffusion image synthesis pipeline, introducing an additional pathway that accepts structural control data—in this instance, the output of the Canny edge detection algorithm—alongside standard text prompts. This edge map is processed through an encoder, the output of which is global average pooled before being merged at intermediate layers within the U-Net backbone of Stable Diffusion.

The Canny model was further refined from its predecessor through continued training, benefitting from an extended session on high-performance hardware—specifically, 72 hours utilizing eight NVIDIA A100 80GB GPUs at a batch size of 256. This computational investment aimed to enhance model robustness and visual quality. During training, Canny edge maps were generated with variable thresholds and paired with high-quality, prompt-annotated images. To improve generalization, data augmentations such as random horizontal flipping were employed. Several known issues in the earlier dataset, including duplication of certain image modalities and corrupted data, were carefully addressed, resulting in a cleaner and more representative data distribution.

Technical Features and Control Capabilities

The defining feature of ControlNet SDXL Canny is its ability to harness edge information for tightly constrained image synthesis. By conditioning the generative process on a Canny edge map, the model ensures that the resulting images adhere closely to the key contours and shapes defined by the user. This facilitates a variety of applications, such as reconstructing images from scanned line art, transforming technical blueprints into photorealistic renders, or enabling precise style transfer while preserving essential form.

Canny edge-controlled AI image synthesis sample

Figure 2. Demonstration of Canny edge-guided generation: The left shows a Canny edge map, while the right displays four stylized outputs synthesized by the model from the detected edge contours. (Prompt unavailable.)

The model’s structure enables multi-condition input, allowing for the use of several ControlNets simultaneously. This feature broadens the creative possibilities by supporting parallel guidance signals, such as blending edge constraints with segmentation maps, pose estimation, or other modalities. Operational details for using multi-ControlNet setups are elaborated in the ControlNet documentation.

Model Performance and Use Cases

ControlNet SDXL Canny was designed to provide enhancements in robustness and subjective perceptual quality. Although no quantitative benchmarks have been formally published, developer insights include observations of its fidelity and reliability during image generation. Practical applications extend across a spectrum of domains, including image reconstruction, stylization with strict geometric adherence, blueprint interpretation, and technical visualization.

Multiple ControlNet outputs using Canny and Shuffle

Figure 3. Generated outputs showing the effect of combined Canny and Shuffle ControlNets on image realism and stylistic diversity in a photorealistic setting. (Prompt unavailable.)

In practice, users can employ the model for tasks requiring high-fidelity adherence to structure, with the outputs often serving production environments in visual design, digital artistry, and technical illustration.

Interface demonstrating Canny edge-guided generation

Figure 4

Integration, Installation, and Related Models

ControlNet SDXL Canny is provided as part of the broader ControlNet 1.1 model suite, each model following the Standard ControlNet Naming Rules (SCNNRs). The Canny model specifically requires the relevant model and configuration files—typically named control_v11p_sd15_canny.pth and control_v11p_sd15_canny.yaml—alongside a compatible version of Stable Diffusion (such as Stable Diffusion 1.5). Official integration and multi-model interoperability are supported through the sd-webui-controlnet extension, which permits the orchestration of multiple conditioning models within user workflows.

Within the 1.1 release, a total of 14 models were introduced, addressing a wide range of guidance modalities. These include depth maps (ControlNet Depth), normal maps (ControlNet Normal), straight lines, scribbles, soft edges, semantic segmentation, pose estimation, line art, anime-specific line art, content shuffling, inpainting, and tile-based upscaling or enhancement. Each model is tailored to leverage a specific annotation or structured control input, while sharing core architectural consistency for ease of use and extensibility.

Limitations and Considerations

The developers note that ControlNet SDXL Canny, like other ControlNet models, is not directly provided as an extension for all Stable Diffusion user interfaces; users are advised to consult the sd-webui-controlnet repository for guidance on proper integration. Certain related models in the family (such as Shuffle, InstructPix2Pix, and Tile) are considered experimental and may require further curation or exhibit varying output stability.

The architecture’s reliance on precise edge map inputs makes it sensitive to the quality and relevance of those inputs. The model's ability to generalize beyond the contours defined by the edge map is deliberately constrained, which is desirable for some use cases but may limit flexibility in others. Requirements such as specific base model checkpoints or compatibility with certain prompt structures should be considered prior to deployment.

External Links and Resources

About Stable Diffusion XL: The SDXL family of AI models represents a significant technological advancement in text-to-image generation, featuring a substantial increase in parameters—3.5 billion for the base model and 6.6 billion for the ensemble—alongside a dual-stage architecture that includes a base model and a refiner, enabling the creation of high-resolution (1024x1024) images with enhanced detail, color accuracy, and style versatility.

More in the Stable Diffusion XL Family

stabilityai /

Stable Diffusion XL

A text-to-image diffusion model with 3.5 billion parameters utilizing a two-stage generation pipeline for enhanced image quality and prompt adherence.
stabilityai /

SDXL Turbo

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, only requiring 2-6 steps instead of 20-60.
ByteDance /

SDXL Lightning

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, with multiple variants for different numbers of steps (1-8) and a more permissive license than SDXL Turbo.
dataautogpt3 /

OpenDalle

SDXL model focused on prompt adherence and semantic understanding, with a stated aim of achieving Dalle-3 level understanding of prompts.
Yamer /

Yamer's Realistic

SDXL fine-tune capable of generating realistic images of people and landscapes.
albedobond /

AlbedoBase XL

SDXL fine-tune resulting from merging 300+ top community models, strong performance across a wide range of image types.
KandooAI /

Juggernaut XL

A popular versatile model finetuned on SDXL by KandooAI and RunDiffusion. This version features improved prompt adherence due to an innovative GPT-4 captioning system built by LEOSAM.
SG_161222 /

Realistic Vision XL

SDXL fine-tune optimizing for generating photorealistic people, animals, and landscapes. In addition to training, this model is a merge of over 10 other SDXL models aimed at realism.
ALIENHAZE /

New Reality XL

Merge of multiple SDXL checkpoints and LoRAs, with impressive breadth and consistency in generating photorealistic images.
razzz /

Realism Engine SDXL

SDXL checkpoint optimized for photorealistic generation of a wide range of humans.
CagliostroLab /

Animagine XL

SDXL fine-tune with very strong performance in genearating anime images.
SoCalGuitarist /

Nightvision XL

Photography-focused SDXL checkpoint with strong prompt adherence and versatile output.
Lykon /

Dreamshaper XL

Fine-Tuned on SDXL, this is a general purpose model designed for photos, art, anime, and manga. This Lightning version can generate quality images in few steps.
PurpleSmartAI /

Pony Diffusion V6 XL

A fine-tuned Stable Diffusion XL model trained on 2.6 million images for generating cartoon, anime, and anthropomorphic character art.
diffusers /

ControlNet SDXL Diffusers Canny

SDXL ControlNet model for edge detection.
diffusers /

ControlNet SDXL Diffusers Depth

SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Depth

Smaller SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Recolor

SDXL ControlNet model for recoloring images.
h94 /

ControlNet SDXL IP Adapter

SDXL ControlNet model for conditioning on an image prompt.
thibaud /

ControlNet SDXL Open Pose

SDXL ControlNet model for copying human poses.

More from stabilityai

stabilityai /

Stable Video Diffusion

A latent diffusion model that generates short video clips up to 25 frames from single images using temporal convolution and attention layers.
stabilityai /

Stable Video Diffusion XT

A video generation model that creates coherent sequences from static images or text prompts using latent diffusion architecture.
stabilityai /

Stable Video Diffusion XT 1.1

Video generation model that transforms single images into 25-frame sequences at 1024x576 resolution with controllable motion and camera parameters.
stabilityai /

Stable Video 3D

Stable Video 3D (SV3D) is a generative model based on Stable Video Diffusion that takes in a still image of an object as a conditioning frame, and generates an orbital video of that object.
stabilityai /

Stable Video 4D

A generative video-to-video diffusion model that synthesizes temporally and spatially consistent multi-view video sequences from single input videos.
stabilityai /

Stable Fast 3D

Generates textured 3D meshes with material properties from single images in approximately 0.5 seconds using a transformer-based architecture.
stabilityai /

Stable Diffusion 2

A text-to-image diffusion model offering 768×768 resolution generation with depth conditioning, inpainting capabilities, and 4x upscaling functionality.
stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
stabilityai /

Stable Diffusion 3.5 Large

This 8-billion parameter model brings improvements to SD3 in customizability, efficiency, and diversity. This Large variant was designed for professional use cases at 1 megapixel resolution.
stabilityai /

Stable Diffusion 3.5 Turbo

This latest iteration in the Stable diffusion family brings improvements over SD3 in customizability, efficient performance, and diverse outputs. The Turbo variant is distilled from the Large version to generate images in just 4 steps.
stabilityai /

Stable Cascade Stage A

A vector quantized GAN encoder that compresses 1024×1024 images to 256×256 discrete tokens as part of a three-stage hierarchical text-to-image pipeline.
stabilityai /

Stable Cascade Stage B

Intermediate latent super-resolution module that upscales compressed text-conditional representations from Stage C to enable high-fidelity image generation.
stabilityai /

Stable Cascade Stage C

Text-to-image diffusion model using three-stage cascaded architecture with 42:1 spatial compression for efficient high-resolution synthesis.
stabilityai /

Stable Audio Open 1.0

Open-source text-to-audio synthesis model with 1.21 billion parameters, trained exclusively on Creative Commons data using latent diffusion architecture.