Skip to main content
Browse Models

lllyasviel

ControlNet SD 1.5 Inpaint

Released

2023-04-13

Family

Stable Diffusion 1

Type

ControlNet Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · control_v11p_sd15_inpaint.pth

Model Report

Overview

ControlNet SD 1.5 Inpaint is a specialized deep learning model designed for controlled inpainting within the broader ControlNet 1.1 model suite, leveraging Stable Diffusion 1.5 as its generative backbone. It introduces masking-based guidance for image completion, enabling the generation of realistic content within specified regions of an image. The model is trained to be applicable not only for static image inpainting but also for video-related tasks, where handling occlusions and maintaining temporal coherence are crucial.

Diagram explaining ControlNet model naming conventions (SCNNRs), highlighting parts of an example model filename.

Figure 1. Diagram illustrating the Standard ControlNet Naming Rules (SCNNRs) for ControlNet 1.1 models.

Technical Capabilities

ControlNet SD 1.5 Inpaint is optimized for intelligent completion of masked regions within images, using user-provided prompts and surrounding context as guidance. The model's architecture is identical to that of earlier ControlNet versions, ensuring compatibility and stability across the 1.x series. Its training regimen involves a mixture of randomly generated masks and optical flow occlusion masks, promoting robustness in both general inpainting and video processing scenarios.

Batch output grid showing ControlNet 1.1 Inpaint model's generated completions for masked input image with prompt 'a handsome man'.

Figure 2. Example batch outputs from ControlNet 1.1 Inpaint for the prompt 'a handsome man', where the masked region of the man's head is rendered with various plausible completions.

A crucial element of the model's functionality is mask-based training. Approximately half of the training utilizes arbitrary random masks, while the remainder employs occlusion masks derived from optical flow data. This dual approach ensures the model can both fill in missing image content and adapt to motion-induced occlusions in video frames. As a result, the model supports not only static image restoration but also video inpainting and optical flow warping applications.

Model Architecture and Training Methods

ControlNet SD 1.5 Inpaint is built upon the ControlNet 1.1 architecture, which itself is closely aligned with the earlier 1.0 release. This consistent architecture enables seamless integration with Stable Diffusion 1.5, utilizing the v1-5-pruned.ckpt checkpoint as the generative foundation.

The training process incorporates two main mask strategies. In the first, the model is exposed to images with randomly positioned and shaped masks, fostering general-purpose inpainting adaptability. In the second, masks are informed by simulated occlusions based on optical flow, a technique commonly used in video analysis to track pixel movements. This blend trains the model to handle both static and dynamic occlusions, facilitating plausible image completions even within complex visual environments.

All annotator submodels required for processing control information (such as HED or OpenPose) can be automatically retrieved or manually installed from the annotators repository, streamlining deployment and experimentation.

Applications and Use Cases

The primary application of ControlNet SD 1.5 Inpaint is guided image completion, where masked areas of an image are filled in contextually, based on adjacent visual information and textual prompts. This capability is widely leveraged in artistic editing, restoration of incomplete or damaged photographs, object removal, and scene modification.

Additionally, the model's exposure to optical flow occlusion masks extends its utility to video-related domains. It can support temporally consistent frame interpolation, video stabilization, and the seamless filling of regions that become exposed due to object motion or camera movement. The controlled approach allows for iterative refinement and fine-tuning of masked content to achieve visually coherent results across multiple frames or images.

Related Models and Extensions

ControlNet 1.1 encompasses a diverse set of models tailored for distinct conditional input types, governed by the Standard ControlNet Naming Rules (SCNNRs). These include models for edge detection, depth estimation, normal maps, line art, pose approximation, semantic segmentation, and more, each optimized with domain-specific dataset augmentations or architectures.

Within this family, the "production-ready" models—for instance, those for depth (control_v11f1p_sd15_depth), canny edges, and normal maps—prioritize robustness and accuracy for practical scenarios. Experimental models explore novel forms of control, such as compositional shuffling or instruction-based editing, as exemplified by the Instruct Pix2Pix variant.

A notable development in the ecosystem is the introduction of Control-LoRAs, a parameter-efficient fine-tuning technique reducing resource demands while maintaining substantial control fidelity. These variants facilitate consumer-level deployment by dramatically decreasing model size without heavily compromising capability.

Limitations

The documentation for ControlNet SD 1.5 Inpaint does not specify unique model-specific limitations, beyond general caveats applicable to machine learning systems—such as potential artifacts in complex masking scenarios or inconsistencies in highly ambiguous regions. As with other models in the suite, successful operation may depend on correct file management and version compatibility with the base Stable Diffusion model and associated annotators.

Licensing and Availability

While explicit licensing details for ControlNet SD 1.5 Inpaint are not detailed in the main documentation, the model and related resources are widely distributed for academic, research, and non-commercial experimentation. Researchers and developers are advised to review individual repositories and documentation for any updated licensing statements.

External Resources

About Stable Diffusion 1: Stable Diffusion is an open-source text-to-image generative AI model that transforms textual adminDescriptions into corresponding images. Technologically, it employs a latent diffusion model architecture, enhancing computational efficiency by performing diffusion processes in a compressed latent space, which enables high-quality image generation with reduced resource requirements.

More in the Stable Diffusion 1 Family

stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
prompthero /

OpenJourney v4

SD 1.5 fine-tuned on 124k+ additional images generated with Midjourney v4, leading to results that resemble this other closed-source image generation model.
Photographer /

Photon

Photon aims to generate photorealistic and visually appealing images effortlessly.
KandooAI /

Juggernaut

Popular SD 1.5 fine-tune with capability to produce detailed images of a versatile breadth of subjects.
wavymulder /

Analog Diffusion

SD 1.5 fine-tuned on a diverse set of analog images, yielding a vintage photographic look.
Lykon /

Dreamshaper

SD 1.5 fine-tune with strong art generation ability and a broad generalist capabilities.
SG_161222 /

Realistic Vision

SD 1.5 fine-tune specialized in creating photorealistic portraits of humans.
Meina /

Meina Mix

Model resulting for merging 7 different anime-focused SD 1.5 checkpoints.
epinikion /

epiCRealism

Popular SD 1.5 fine-tune with high competence in translating simple text prompts into realistic images of people.
Lykon /

Absolute Reality

One of the top SD 1.5 variant for generating life-like images of people and objects.
Cyberdelia /

Cyber Realistic

Versatile photorealistic SD 1.5 fine-tune capable of generating a wide range of convincing photographic images.
Merjic /

MajicMIX Realistic

Popular SD 1.5 photorealism fine-tune with training data weighted on people of asian descent.
epinikion /

epiCPhotoGasm

A Stable Diffusion 1.5-based checkpoint model designed for photorealistic image generation with simplified prompting and demographic diversity.
lllyasviel /

ControlNet SD 1.5 Canny

SD 1.5 ControlNet model to replicate the composion of a source image using edge-detection.
lllyasviel /

ControlNet SD 1.5 IP2P

SD 1.5 ControlNet trained with pixel-to-pixel instruction.
lllyasviel /

ControlNet SD 1.5 Depth

SD 1.5 ControlNet model to replicate the depth of a source image.
lllyasviel /

ControlNet SD 1.5 MLSD

SD 1.5 ControlNet model to detect straight-lines, useful for architecture and man-made objects.
lllyasviel /

ControlNet SD 1.5 Normal

SD 1.5 ControlNet model to replicate the depth of a source image, with additional surface details and geometry.
lllyasviel /

ControlNet SD 1.5 Open Pose

SD 1.5 ControlNet model for copying human poses.
lllyasviel /

ControlNet SD 1.5 Scribble

SD 1.5 ControlNet model for converting sketches to images.
lllyasviel /

ControlNet SD 1.5 Segmentation

SD 1.5 ControlNet model for detecting and segmenting distinct parts of images to use in the generation.
lllyasviel /

ControlNet SD 1.5 Soft Edge

SD 1.5 ControlNet model to detect soft-edges, especially useful for recoloring and stylizing.
lllyasviel /

ControlNet SD 1.5 Line Art

SD 1.5 ControlNet model trained with line art generation.
lllyasviel /

ControlNet SD 1.5 Lineart Anime

SD 1.5 ControlNet model trained with anime line art generation.
lllyasviel /

ControlNet SD 1.5 Shuffle

SD 1.5 ControlNet model trained with image shuffling.
lllyasviel /

ControlNet SD 1.5 Tile

SD 1.5 ControlNet model trained with image tiling.
tencent /

ControlNet 1.5 IP Adapter

SD 1.5 ControlNet model for conditioning on an image prompt.
tencent /

ControlNet 1.5 QR Code

SD 1.5 ControlNet model for generating stylized QR codes.