Skip to main content
Browse Models

stabilityai

ControlNet SDXL Depth

Released

2023-08-29

Family

Stable Diffusion XL

Type

ControlNet Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · 774 MB · sai_xl_depth_256lora.safetensors

Model Report

Overview

The depth-specific variant of ControlNet 1.1 for Stable Diffusion 1.5, formally known as control_v11f1p_sd15_depth, is a generative artificial intelligence model that enables the targeted guidance of Stable Diffusion 1.5 using depth maps. As part of the ControlNet 1.1 release, it preserves the architecture of previous ControlNet models while introducing improvements in robustness and quality that facilitate more accurate image synthesis conditioned on depth estimations. By integrating depth information, this model extends the capabilities of diffusion models to produce images that adhere closely to the spatial constraints and three-dimensional cues embedded in input depth maps.

Sample outputs from Control-LoRA's MiDaS and ClipDrop Depth workflow

Figure 1. Visualization of depth map-based image generation using a grayscale depth map and resulting portrait outputs. Prompt: Women’s portrait synthesis with grayscale depth control.

Model Architecture and Operation

This model is grounded in the ControlNet 1.1 architecture, closely following the network design of ControlNet 1.0. The model functions as a conditional control system layered upon the Stable Diffusion framework, providing additional input channels that align the generative process with external structural cues, particularly those encoded in depth maps.

Depth maps, which may originate from monocular depth estimation algorithms or real-world 3D renderers, are preprocessed and entered into the model. The depth-aware conditioning enables the diffusion process to respect the foreground-background relationships, occlusion order, and relative distances expressed in the map. This is achieved without architectural divergence from earlier ControlNet versions, ensuring model stability and reproducibility as the developers deliberately deferred major changes until future releases.

ControlNet's codebase is primarily implemented in Python, facilitating modifiability and integration within established diffusion pipelines.

Training Data and Methodology

The depth-specific variant of ControlNet 1.1 was trained using a multi-source dataset that amalgamates depth maps produced by MiDaS, Leres, and Zoe Depth methods. Training incorporated data augmentation strategies, including random left-right flipping, and utilized depth maps across multiple input resolutions (256, 384, and 512 pixels) to bolster generalization to different scale and source variations.

Key improvements over prior models address several data quality challenges. Corrections to the training set eliminated duplicated grayscale images and low-quality samples, while also refining the correspondence between images and textual prompts. These enhancements yielded a model that is not tuned to any singular depth estimation method, resulting in robust performance even when driven by depth maps from novel sources or varying preprocessing pipelines.

To further harness resource efficiency, parallel developments such as Control-LoRA's MiDaS and ClipDrop Depth models have demonstrated that low-rank adaptation techniques can produce compact models capable of similar depth-based guidance, using training inputs like MiDaS dpt_beit_large_512 and finetuning with ClipDrop's Portrait Depth Estimation.

Applications and Use Cases

This model principally serves to guide the image generation capabilities of Stable Diffusion 1.5 according to the constraints imposed by input depth maps. The resulting system is able to generate images consistent with the geometric arrangement and spatial relationships of objects specified by the depth input. This is particularly useful in tasks requiring fidelity to three-dimensional scene layout, such as the generation of photorealistic portraits, interior renderings, and arbitrary scene synthesis from sketched or programmatically derived depth cues.

Depth-based control is one modality among several in the broader ControlNet 1.1 family of models, which encompasses additional control types including edge, normal map, scribble, line art, soft edge, segmentation, pose (OpenPose), inpainting, and stylization controls.

Comparison within the Model Family

The depth-specific ControlNet 1.1 model is one of a suite of specialized models designed to provide conditional image generation through a diversity of input modalities. Other prominent models within the same release manage control via Canny edge detection, Bae's normal map estimation, scribble and line art inputs, semantic segmentation, and OpenPose outputs.

Efforts towards more resource-efficient deployment are realized in the form of Control-LoRA variants, which employ low-rank adaptation to reduce model size from approximately 4.7GB to as little as 377MB. This enables depth, edge, and stylization controls to be accessible on consumer hardware with a reduced computational footprint. These LoRA-based models maintain the core functionalities of their larger ControlNet counterparts, offering similar user control while facilitating broader accessibility.

Limitations and Considerations

While this model generally provides robust depth conditioning, certain modalities in the ControlNet family—such as the experimental instruct-based ip2p and the stylization-oriented shuffle—may require user discretion and iterative refinement for optimal results. Additionally, some specialized models like the anime-focused line art variant (control_v11p_sd15s2_lineart_anime.pth) are contingent on specific model checkpoints and do not support all operational modes present in other controls.

The license governing this model and its family of models has not been explicitly detailed in the available documentation.

External Resources

For further technical insights, code, and downloads:

About Stable Diffusion XL: The SDXL family of AI models represents a significant technological advancement in text-to-image generation, featuring a substantial increase in parameters—3.5 billion for the base model and 6.6 billion for the ensemble—alongside a dual-stage architecture that includes a base model and a refiner, enabling the creation of high-resolution (1024x1024) images with enhanced detail, color accuracy, and style versatility.

More in the Stable Diffusion XL Family

stabilityai /

Stable Diffusion XL

A text-to-image diffusion model with 3.5 billion parameters utilizing a two-stage generation pipeline for enhanced image quality and prompt adherence.
stabilityai /

SDXL Turbo

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, only requiring 2-6 steps instead of 20-60.
ByteDance /

SDXL Lightning

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, with multiple variants for different numbers of steps (1-8) and a more permissive license than SDXL Turbo.
dataautogpt3 /

OpenDalle

SDXL model focused on prompt adherence and semantic understanding, with a stated aim of achieving Dalle-3 level understanding of prompts.
Yamer /

Yamer's Realistic

SDXL fine-tune capable of generating realistic images of people and landscapes.
albedobond /

AlbedoBase XL

SDXL fine-tune resulting from merging 300+ top community models, strong performance across a wide range of image types.
KandooAI /

Juggernaut XL

A popular versatile model finetuned on SDXL by KandooAI and RunDiffusion. This version features improved prompt adherence due to an innovative GPT-4 captioning system built by LEOSAM.
SG_161222 /

Realistic Vision XL

SDXL fine-tune optimizing for generating photorealistic people, animals, and landscapes. In addition to training, this model is a merge of over 10 other SDXL models aimed at realism.
ALIENHAZE /

New Reality XL

Merge of multiple SDXL checkpoints and LoRAs, with impressive breadth and consistency in generating photorealistic images.
razzz /

Realism Engine SDXL

SDXL checkpoint optimized for photorealistic generation of a wide range of humans.
CagliostroLab /

Animagine XL

SDXL fine-tune with very strong performance in genearating anime images.
SoCalGuitarist /

Nightvision XL

Photography-focused SDXL checkpoint with strong prompt adherence and versatile output.
Lykon /

Dreamshaper XL

Fine-Tuned on SDXL, this is a general purpose model designed for photos, art, anime, and manga. This Lightning version can generate quality images in few steps.
PurpleSmartAI /

Pony Diffusion V6 XL

A fine-tuned Stable Diffusion XL model trained on 2.6 million images for generating cartoon, anime, and anthropomorphic character art.
diffusers /

ControlNet SDXL Diffusers Canny

SDXL ControlNet model for edge detection.
stabilityai /

ControlNet SDXL Canny

Smaller SDXL ControlNet model for edge detection.
diffusers /

ControlNet SDXL Diffusers Depth

SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Recolor

SDXL ControlNet model for recoloring images.
h94 /

ControlNet SDXL IP Adapter

SDXL ControlNet model for conditioning on an image prompt.
thibaud /

ControlNet SDXL Open Pose

SDXL ControlNet model for copying human poses.

More from stabilityai

stabilityai /

Stable Video Diffusion

A latent diffusion model that generates short video clips up to 25 frames from single images using temporal convolution and attention layers.
stabilityai /

Stable Video Diffusion XT

A video generation model that creates coherent sequences from static images or text prompts using latent diffusion architecture.
stabilityai /

Stable Video Diffusion XT 1.1

Video generation model that transforms single images into 25-frame sequences at 1024x576 resolution with controllable motion and camera parameters.
stabilityai /

Stable Video 3D

Stable Video 3D (SV3D) is a generative model based on Stable Video Diffusion that takes in a still image of an object as a conditioning frame, and generates an orbital video of that object.
stabilityai /

Stable Video 4D

A generative video-to-video diffusion model that synthesizes temporally and spatially consistent multi-view video sequences from single input videos.
stabilityai /

Stable Fast 3D

Generates textured 3D meshes with material properties from single images in approximately 0.5 seconds using a transformer-based architecture.
stabilityai /

Stable Diffusion 2

A text-to-image diffusion model offering 768×768 resolution generation with depth conditioning, inpainting capabilities, and 4x upscaling functionality.
stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
stabilityai /

Stable Diffusion 3.5 Large

This 8-billion parameter model brings improvements to SD3 in customizability, efficiency, and diversity. This Large variant was designed for professional use cases at 1 megapixel resolution.
stabilityai /

Stable Diffusion 3.5 Turbo

This latest iteration in the Stable diffusion family brings improvements over SD3 in customizability, efficient performance, and diverse outputs. The Turbo variant is distilled from the Large version to generate images in just 4 steps.
stabilityai /

Stable Cascade Stage A

A vector quantized GAN encoder that compresses 1024×1024 images to 256×256 discrete tokens as part of a three-stage hierarchical text-to-image pipeline.
stabilityai /

Stable Cascade Stage B

Intermediate latent super-resolution module that upscales compressed text-conditional representations from Stage C to enable high-fidelity image generation.
stabilityai /

Stable Cascade Stage C

Text-to-image diffusion model using three-stage cascaded architecture with 42:1 spatial compression for efficient high-resolution synthesis.
stabilityai /

Stable Audio Open 1.0

Open-source text-to-audio synthesis model with 1.21 billion parameters, trained exclusively on Creative Commons data using latent diffusion architecture.