Skip to main content
Browse Models

stabilityai

ControlNet SDXL Recolor

Released

2023-08-29

Family

Stable Diffusion XL

Type

ControlNet Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · 774 MB · sai_xl_recolor_256lora.safetensors

Model Report

Overview

ControlNet SDXL Recolor is a generative AI model designed for image colorization, functioning as a Control-LoRA extension for the Stable Diffusion XL (SDXL) architecture. This model specializes in adding color to grayscale photographs and hand-drawn sketches, supporting production applications within the broader ControlNet 1.1 model suite. Developed with a focus on efficiency and versatility, ControlNet SDXL Recolor enables users to transform monochrome media into colored outputs. The model implements parameter-efficient fine-tuning strategies, making it suitable for various computational environments and integration with other ControlNet models.

Collage showing colorized versions of a black and white portrait and sketch

Figure 1. Sample outputs of ControlNet SDXL Recolor: colorized versions of both black and white photographs and sketches demonstrate the model's dual functionality for restoration and creative coloring tasks.

Model Architecture and Technology

ControlNet SDXL Recolor is built upon the ControlNet 1.1 architecture, inheriting the modular framework designed for guided image synthesis. Control-LoRA models utilize low-rank adaptation (LoRA), which introduces parameter-efficient fine-tuning by injecting small trainable matrices into the original architecture. This approach reduces the file size compared to full ControlNet checkpoints, decreasing typical model storage from 4.7GB to as little as 738MB or 377MB, depending on the selected rank, while maintaining performance characteristics consistent with the original model.

At its core, ControlNet overlays conditioning networks onto the image generation pipeline of SDXL, allowing for fine-grained control based on input hints such as edges, depth maps, or, in the case of Recolor, grayscale tonal distribution and sketch outlines. A crucial implementation detail involves the use of global average pooling to aggregate features before merging them into the SDXL UNet layers. This enables efficient modulation of the generation process, ensuring colorization is applied only on the conditional side of the classifier-free guidance (Cfg) scale. These technical improvements foster the model's integration with SDXL, resulting in context-aware colorization from varied forms of monochrome input. Details of the architecture and LoRA methodology are available in the ControlNet 1.1 documentation and Hugging Face release notes.

Colorization Capabilities and Use Cases

ControlNet SDXL Recolor is configured for two principal colorization tasks. First, in “Recolor” mode, it restores and enhances black and white photographs, simulating naturalistic or creatively stylized color palettes. Second, its “Sketch” mode targets white-on-black drawings—either hand-drawn or generated with edge detection methods such as the pidi edge model—providing a framework for artists to apply color to line art.

Both modes employ learned mappings from tonal cues to plausible color distributions, leveraging the SDXL backbone’s expressivity. The model can be deployed for photo restoration, archival enhancement, creative illustration, and preprocessing in digital workflows where coloring is desired. Its compatibility with other ControlNets further expands its utility, facilitating combined operations such as joint colorization and inpainting, or mixing with depth or posture controls for complex image manipulations. General model usage and multi-ControlNet workflows are supported as described in the ControlNet project documentation and the Automatic1111 plugin integration.

Recolor workflow in ComfyUI showing grayscale input and colorized output

Figure 2. The ComfyUI interface demonstrates ControlNet SDXL Recolor in action: a grayscale portrait is transformed into a realistic, colorized image through a visual processing workflow. The prompt used is a grayscale headshot of a woman, illustrating both the input and the resulting output.

Training Data and Methodology

Details about the precise datasets used to train ControlNet SDXL Recolor have not been fully disclosed. However, the broader ControlNet 1.1 family emphasizes enhanced data quality and augmentation techniques. Enhancements over earlier versions include the removal of duplicate grayscale human images, reduction of low-quality or compressed images, and correction of prompts for better paired training. Augmentation strategies—such as random image flipping—help generalize model performance across diverse photographic and sketch styles. These refinements contribute to reduced artifacts in colorization outcomes, as discussed in the official ControlNet 1.1 updates.

The overall training process involves conditioning the model to learn plausible mappings between input grayscale or line art images and their corresponding colorized forms. The Control-LoRA training procedure focuses on minimizing additional parameter footprint while maintaining fidelity in generated results.

Integration and Compatibility

ControlNet SDXL Recolor is engineered for integration with SDXL workflows, benefiting from compatibility with various user interfaces such as ComfyUI and StableSwarmUI. The model also participates in the “Multi-ControlNet” ecosystem, wherein additive controls from multiple ControlNets (including community models and custom LoRAs) can be arbitrarily combined, especially in production pipelines managed through the Automatic1111 plugin.

Model deployment typically involves referencing the appropriate Control-LoRA checkpoint (by rank and mode), ensuring that associated configuration files activate recommended features such as global average pooling. Implementation instructions, compatibility details, and advanced setup recommendations are regularly maintained in the ControlNet project repositories.

Model Family and Related Models

ControlNet SDXL Recolor belongs to a suite of ControlNet 1.1 models, each designed to condition SDXL-based image generation on different input modalities. Notable models in the family include ControlNet SDXL Canny for edge-based guidance, ControlNet SDXL Depth for spatial awareness, ControlNet SDXL Openpose for pose estimation, and ControlNet SDXL IP-Adapter for image prompt adaptation. This modular approach enables coverage of image transformation tasks, with each Control-LoRA providing efficient, focused parameter updates for its target function.

The full set of ControlNet 1.1 models, including both production-ready and experimental controls, is described in the project documentation. The suite’s heterogeneous architecture and shared integration strategy facilitate advanced, multi-stage workflows in digital art, photography, and research contexts.

Limitations and Known Issues

While ControlNet SDXL Recolor and its companion models offer broad utility, they inherit several known limitations from the ControlNet 1.1 framework. Some models in the family are classified as experimental and may require manual selection of optimal outputs. Functionality such as multi-ControlNet composition or tiled upscaling is primarily supported via specific plugins and user interfaces, with standalone deployments often lacking full feature parity. Dependency on the Automatic1111 plugin is recommended for utilizing advanced orchestration features. Additionally, some related models—such as Anime Lineart—may require supplementary files not provided directly by the project.

External Resources

For further information and technical details, the following resources are available:

About Stable Diffusion XL: The SDXL family of AI models represents a significant technological advancement in text-to-image generation, featuring a substantial increase in parameters—3.5 billion for the base model and 6.6 billion for the ensemble—alongside a dual-stage architecture that includes a base model and a refiner, enabling the creation of high-resolution (1024x1024) images with enhanced detail, color accuracy, and style versatility.

More in the Stable Diffusion XL Family

stabilityai /

Stable Diffusion XL

A text-to-image diffusion model with 3.5 billion parameters utilizing a two-stage generation pipeline for enhanced image quality and prompt adherence.
stabilityai /

SDXL Turbo

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, only requiring 2-6 steps instead of 20-60.
ByteDance /

SDXL Lightning

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, with multiple variants for different numbers of steps (1-8) and a more permissive license than SDXL Turbo.
dataautogpt3 /

OpenDalle

SDXL model focused on prompt adherence and semantic understanding, with a stated aim of achieving Dalle-3 level understanding of prompts.
Yamer /

Yamer's Realistic

SDXL fine-tune capable of generating realistic images of people and landscapes.
albedobond /

AlbedoBase XL

SDXL fine-tune resulting from merging 300+ top community models, strong performance across a wide range of image types.
KandooAI /

Juggernaut XL

A popular versatile model finetuned on SDXL by KandooAI and RunDiffusion. This version features improved prompt adherence due to an innovative GPT-4 captioning system built by LEOSAM.
SG_161222 /

Realistic Vision XL

SDXL fine-tune optimizing for generating photorealistic people, animals, and landscapes. In addition to training, this model is a merge of over 10 other SDXL models aimed at realism.
ALIENHAZE /

New Reality XL

Merge of multiple SDXL checkpoints and LoRAs, with impressive breadth and consistency in generating photorealistic images.
razzz /

Realism Engine SDXL

SDXL checkpoint optimized for photorealistic generation of a wide range of humans.
CagliostroLab /

Animagine XL

SDXL fine-tune with very strong performance in genearating anime images.
SoCalGuitarist /

Nightvision XL

Photography-focused SDXL checkpoint with strong prompt adherence and versatile output.
Lykon /

Dreamshaper XL

Fine-Tuned on SDXL, this is a general purpose model designed for photos, art, anime, and manga. This Lightning version can generate quality images in few steps.
PurpleSmartAI /

Pony Diffusion V6 XL

A fine-tuned Stable Diffusion XL model trained on 2.6 million images for generating cartoon, anime, and anthropomorphic character art.
diffusers /

ControlNet SDXL Diffusers Canny

SDXL ControlNet model for edge detection.
stabilityai /

ControlNet SDXL Canny

Smaller SDXL ControlNet model for edge detection.
diffusers /

ControlNet SDXL Diffusers Depth

SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Depth

Smaller SDXL ControlNet model for depth generation.
h94 /

ControlNet SDXL IP Adapter

SDXL ControlNet model for conditioning on an image prompt.
thibaud /

ControlNet SDXL Open Pose

SDXL ControlNet model for copying human poses.

More from stabilityai

stabilityai /

Stable Video Diffusion

A latent diffusion model that generates short video clips up to 25 frames from single images using temporal convolution and attention layers.
stabilityai /

Stable Video Diffusion XT

A video generation model that creates coherent sequences from static images or text prompts using latent diffusion architecture.
stabilityai /

Stable Video Diffusion XT 1.1

Video generation model that transforms single images into 25-frame sequences at 1024x576 resolution with controllable motion and camera parameters.
stabilityai /

Stable Video 3D

Stable Video 3D (SV3D) is a generative model based on Stable Video Diffusion that takes in a still image of an object as a conditioning frame, and generates an orbital video of that object.
stabilityai /

Stable Video 4D

A generative video-to-video diffusion model that synthesizes temporally and spatially consistent multi-view video sequences from single input videos.
stabilityai /

Stable Fast 3D

Generates textured 3D meshes with material properties from single images in approximately 0.5 seconds using a transformer-based architecture.
stabilityai /

Stable Diffusion 2

A text-to-image diffusion model offering 768×768 resolution generation with depth conditioning, inpainting capabilities, and 4x upscaling functionality.
stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
stabilityai /

Stable Diffusion 3.5 Large

This 8-billion parameter model brings improvements to SD3 in customizability, efficiency, and diversity. This Large variant was designed for professional use cases at 1 megapixel resolution.
stabilityai /

Stable Diffusion 3.5 Turbo

This latest iteration in the Stable diffusion family brings improvements over SD3 in customizability, efficient performance, and diverse outputs. The Turbo variant is distilled from the Large version to generate images in just 4 steps.
stabilityai /

Stable Cascade Stage A

A vector quantized GAN encoder that compresses 1024×1024 images to 256×256 discrete tokens as part of a three-stage hierarchical text-to-image pipeline.
stabilityai /

Stable Cascade Stage B

Intermediate latent super-resolution module that upscales compressed text-conditional representations from Stage C to enable high-fidelity image generation.
stabilityai /

Stable Cascade Stage C

Text-to-image diffusion model using three-stage cascaded architecture with 42:1 spatial compression for efficient high-resolution synthesis.
stabilityai /

Stable Audio Open 1.0

Open-source text-to-audio synthesis model with 1.21 billion parameters, trained exclusively on Creative Commons data using latent diffusion architecture.