Skip to main content
Browse Models

lllyasviel

ControlNet SD 1.5 Shuffle

Released

2023-04-13

Family

Stable Diffusion 1

Type

ControlNet Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · control_v11e_sd15_shuffle.pth

Model Report

Overview

ControlNet SD 1.5 Shuffle is a generative artificial intelligence model belonging to the ControlNet 1.1 family, designed to enhance the controllability and flexibility of Stable Diffusion 1.5 via guided image content shuffling. This model employs an approach for reorganizing image content, thereby enabling transformations in image style and structure under prompt-based guidance. As an experimental addition to the ControlNet 1.1 suite, ControlNet SD 1.5 Shuffle introduces mechanisms for direct manipulation of image composition, facilitating image-to-image tasks while leveraging the generative capabilities of Stable Diffusion.

ControlNet 1.1 Naming Convention Diagram

Figure 1. Diagram illustrating the naming conventions for ControlNet models, clarifying each element of model filenames in the ControlNet 1.1 release.

Model Architecture and Design

ControlNet SD 1.5 Shuffle maintains architectural consistency with ControlNet 1.0, utilizing the same core neural network structure. This deliberate architectural continuity is expected to persist through at least version 1.5 of ControlNet models, simplifying integration and compatibility across different model variants. Shuffle is characterized as a "pure ControlNet," indicating that it operates independently of external computer vision modules such as CLIP, relying solely on its internal mechanisms for image analysis and modification.

A distinguishing architectural feature of the Shuffle model is the inclusion of a global average pooling layer between the encoder outputs and the Stable Diffusion U-Net layers. This layer ensures that global image statistics inform the generative process, fostering coherent reorganization during content shuffling. The implementation is managed through a global average pooling configuration entry in the model’s YAML configuration file. In use, ControlNet SD 1.5 Shuffle must be applied to the conditional branch of classifier-free guidance, a technique widely adopted for fine-tuned diffusion model control.

Training Data and Methodology

While specific training datasets for ControlNet SD 1.5 Shuffle have not been publicly disclosed, the model is described as having been "trained to reorganize images" using a technique referred to as random flow shuffling. This process focuses on disrupting and rearranging image content, thereby equipping the model to direct Stable Diffusion in reconstructing and recomposing images according to textual prompts.

The broader ControlNet 1.1 family benefited from targeted improvements to training datasets, such as reducing duplication, removing low-quality samples, and refining paired prompts. These interventions contribute to improved model robustness and diversity of outputs, as evidenced across the release suite.

Functional Capabilities and Use Cases

ControlNet SD 1.5 Shuffle is designed to perform image transformation, including content recomposition, restyling, and the introduction of structural variations in output images. One of its abilities is to operate effectively even when the input image is not pre-shuffled, demonstrating an inherent capability for interpreting and reorganizing original image content.

The model’s core use cases include style transfer, where an input image is rearranged or stylized based on a provided prompt, and content recomposition, where the model reconstructs shuffled or original imagery according to textual or multimodal instructions. Its utility is further enhanced when used in tandem with other ControlNet models, supporting multi-conditional workflows for visual manipulation tasks.

ControlNet Shuffle: Cityscape Reorganization

Figure 2. Output of ControlNet SD 1.5 Shuffle reorganizing an urban night scene with the prompt 'hong kong' (seed 12345). The model transforms and stylizes the cityscape via content shuffling.

ControlNet Shuffle: Armor Style Transfer

Figure 3. Style change result from ControlNet SD 1.5 Shuffle with the prompt 'iron man' (seed 12345). The model takes the input figure and outputs diverse, Iron Man-inspired armor designs.

ControlNet Shuffle: Spider-Man Style Transformation

Figure 4. Model output for the prompt 'spider man' (seed 12345). The input is transformed into variations of armored Spider-Man-like characters, demonstrating flexible content recomposition.

Position within the ControlNet 1.1 Family

ControlNet 1.1 comprises 14 models, with ControlNet SD 1.5 Shuffle categorized as one of three experimental variants. It is distributed alongside both established and experimental methodologies, including models for depth inference, edge detection (e.g., Canny, MLSD), pose estimation, semantic segmentation, and other style-relevant transformations. All models in the family share the foundational ControlNet architecture but vary in their control mechanisms and training data optimizations.

Notably, experimental models such as Instruct Pix2Pix and Tile, launched in parallel with Shuffle, explore paradigms for controlled image generation, such as instruction-based image transformation and tile-based high-resolution synthesis, respectively. The broader ControlNet suite is intended to facilitate research and development in guided image generation, testing the boundaries of multimodal conditioning and content-level control.

Limitations and Experimental Status

ControlNet SD 1.5 Shuffle is explicitly marked as experimental within the ControlNet 1.1 release. While early communications described Shuffle as a primary method for image stylization—especially in contrast to CLIP-based approaches—further development signals openness to supporting additional stylization techniques in the future. Therefore, the model's long-term direction and canonical role within the family remain subject to ongoing evaluation and potential revision. There is no documented information regarding licensing within official repository materials; practitioners are advised to consult the distribution platform for up-to-date licensing terms.

Release and Development Timeline

ControlNet SD 1.5 Shuffle was introduced as part of the broader ControlNet 1.1 release, which featured expanded training protocols and the initiation of public beta testing within the Automatic1111 (A1111) ecosystem. This version introduced experimental integrations and improvements over prior iterations, notably in dataset curation and conditional guidance mechanics.

External Resources

For additional information, technical documentation, datasets, and related discussion, consult the following resources:

About Stable Diffusion 1: Stable Diffusion is an open-source text-to-image generative AI model that transforms textual adminDescriptions into corresponding images. Technologically, it employs a latent diffusion model architecture, enhancing computational efficiency by performing diffusion processes in a compressed latent space, which enables high-quality image generation with reduced resource requirements.

More in the Stable Diffusion 1 Family

stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
prompthero /

OpenJourney v4

SD 1.5 fine-tuned on 124k+ additional images generated with Midjourney v4, leading to results that resemble this other closed-source image generation model.
Photographer /

Photon

Photon aims to generate photorealistic and visually appealing images effortlessly.
KandooAI /

Juggernaut

Popular SD 1.5 fine-tune with capability to produce detailed images of a versatile breadth of subjects.
wavymulder /

Analog Diffusion

SD 1.5 fine-tuned on a diverse set of analog images, yielding a vintage photographic look.
Lykon /

Dreamshaper

SD 1.5 fine-tune with strong art generation ability and a broad generalist capabilities.
SG_161222 /

Realistic Vision

SD 1.5 fine-tune specialized in creating photorealistic portraits of humans.
Meina /

Meina Mix

Model resulting for merging 7 different anime-focused SD 1.5 checkpoints.
epinikion /

epiCRealism

Popular SD 1.5 fine-tune with high competence in translating simple text prompts into realistic images of people.
Lykon /

Absolute Reality

One of the top SD 1.5 variant for generating life-like images of people and objects.
Cyberdelia /

Cyber Realistic

Versatile photorealistic SD 1.5 fine-tune capable of generating a wide range of convincing photographic images.
Merjic /

MajicMIX Realistic

Popular SD 1.5 photorealism fine-tune with training data weighted on people of asian descent.
epinikion /

epiCPhotoGasm

A Stable Diffusion 1.5-based checkpoint model designed for photorealistic image generation with simplified prompting and demographic diversity.
lllyasviel /

ControlNet SD 1.5 Canny

SD 1.5 ControlNet model to replicate the composion of a source image using edge-detection.
lllyasviel /

ControlNet SD 1.5 IP2P

SD 1.5 ControlNet trained with pixel-to-pixel instruction.
lllyasviel /

ControlNet SD 1.5 Depth

SD 1.5 ControlNet model to replicate the depth of a source image.
lllyasviel /

ControlNet SD 1.5 MLSD

SD 1.5 ControlNet model to detect straight-lines, useful for architecture and man-made objects.
lllyasviel /

ControlNet SD 1.5 Normal

SD 1.5 ControlNet model to replicate the depth of a source image, with additional surface details and geometry.
lllyasviel /

ControlNet SD 1.5 Open Pose

SD 1.5 ControlNet model for copying human poses.
lllyasviel /

ControlNet SD 1.5 Scribble

SD 1.5 ControlNet model for converting sketches to images.
lllyasviel /

ControlNet SD 1.5 Segmentation

SD 1.5 ControlNet model for detecting and segmenting distinct parts of images to use in the generation.
lllyasviel /

ControlNet SD 1.5 Soft Edge

SD 1.5 ControlNet model to detect soft-edges, especially useful for recoloring and stylizing.
lllyasviel /

ControlNet SD 1.5 Inpaint

SD 1.5 ControlNet model trained with image inpainting.
lllyasviel /

ControlNet SD 1.5 Line Art

SD 1.5 ControlNet model trained with line art generation.
lllyasviel /

ControlNet SD 1.5 Lineart Anime

SD 1.5 ControlNet model trained with anime line art generation.
lllyasviel /

ControlNet SD 1.5 Tile

SD 1.5 ControlNet model trained with image tiling.
tencent /

ControlNet 1.5 IP Adapter

SD 1.5 ControlNet model for conditioning on an image prompt.
tencent /

ControlNet 1.5 QR Code

SD 1.5 ControlNet model for generating stylized QR codes.