Skip to main content
Browse Models

stabilityai

Stable Diffusion 3.5 Large

Released

2024-10-22

Family

Stable Diffusion 3

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · 9.6 GB · sd3.5_large.safetensors

Model Report

Overview

Stable Diffusion 3.5 Large is an advanced text-to-image generative model developed by Stability AI, released in October 2024 as part of the broader Stable Diffusion 3.5 family. Leveraging the Multimodal Diffusion Transformer (MMDiT) architecture, the model converts natural language prompts into detailed, high-resolution images across a variety of styles and subjects. This model incorporates architectural and training innovations to improve prompt adherence, output diversity, and typographic fidelity, supporting both professional and research-oriented applications, as detailed in the Stability AI announcement.

Compilation grid showcasing Stable Diffusion 3.5 Large capabilities

Figure 1. A collage of diverse AI-generated images, illustrating Stable Diffusion 3.5 Large's ability to produce a variety of subjects and artistic styles based on different text prompts.

Model Architecture and Technical Innovations

Stable Diffusion 3.5 Large builds upon the Multimodal Diffusion Transformer (MMDiT) framework, employing diffusion-based generative mechanisms in conjunction with transformer networks, as described in the MMDiT research paper. A key innovation is the integration of Query-Key Normalization (QK-Normalization) within transformer blocks, a feature that enhances training stability and simplifies the fine-tuning process, according to the Stability AI announcement. This architectural refinement allows for greater variety in outputs from identical prompts with different random seeds, promoting a broader expressive range and stylistic diversity.

The model uses three fixed, pretrained text encoders: OpenCLIP-ViT/G, CLIP-ViT/L, and T5-xxl. These encoders facilitate a maximum context length of up to 256 tokens, depending on the training phase, as noted in the Hugging Face model card. Their inclusion supports complex prompt interpretation and improved text-to-image alignment.

Versatile Styles: Model outputs in etched illustration, photography, and anime digital painting

Figure 2. Stable Diffusion 3.5 Large demonstrates versatile style generation, shown here in etched illustration, photorealistic, and anime digital painting outputs.

Training Data and Performance

Stable Diffusion 3.5 Large was trained on a wide spectrum of datasets, encompassing both large-scale synthetic and rigorously filtered publicly available data. The resulting model contains 8.1 billion parameters and is optimized for image synthesis at resolutions up to 1 megapixel, as reported by Stability AI.

Empirical analysis by Stability AI indicates strong performance in prompt adherence and image quality among diffusion-based text-to-image systems, showing capabilities comparable to models of larger parameter counts, according to a Stability AI blog post. Performance benchmarks indicate consistency in visual quality and representation across various prompts and subject matter.

Bar chart of model prompt adherence and aesthetic quality

Figure 3. Comparison chart of prompt adherence and aesthetic quality (Elo Score) for Stable Diffusion 3.5 Large and peer models.

Distinct Features and Output Diversity

Stable Diffusion 3.5 Large is designed for flexibility and stylistic breadth. The model produces images encompassing a range of genres, including photorealism, illustration, 3D renderings, line art, and digital painting, as stated in the Hugging Face model card. Its outputs are diverse not only in aesthetic but also in depiction, capable of synthesizing visual representations of varied skin tones, features, and subjects with minimal prompting requirements.

Portraits demonstrating diversity of outputs

Figure 4. Triptych of AI-generated portraits, highlighting the model's ability to produce diverse and realistic representations of people without elaborate prompts.

The model's fine-tuning and customization capabilities enable adaptation to specific creative or professional workflows. Its architecture encourages varied and unpredictable outputs in response to less constrained prompts, broadening its application scope across design, art, education, and research uses, as discussed in a Stability AI blog post.

AI-generated images for three distinct prompts

Figure 5. Example prompt-to-image generations, showcasing the model's nuanced interpretation of detailed textual prompts.

Model Variants and Family

The Stable Diffusion 3.5 family comprises Stable Diffusion 3.5 Large, Stable Diffusion 3.5 Large Turbo, and Stable Diffusion 3.5 Medium, as announced by Stability AI. Stable Diffusion 3.5 Large Turbo is a distilled version designed for rapid image synthesis, performing in as few as four inference steps. Stable Diffusion 3.5 Medium, with 2.5 billion parameters, is designed for broad compatibility. Each variant is tailored for distinct use cases, balancing speed, quality, and resource requirements.

Limitations, Safety, and Licensing

Although Stable Diffusion 3.5 Large incorporates multiple safety mitigations, comprehensive removal of all potential harmful content cannot be assured, as stated on the Stability AI safety page. The model was not developed for factual representation of real-world individuals or events, and outcomes may vary in consistency, especially for vague or underspecified prompts, according to the Hugging Face model card. Developers and researchers are encouraged to conduct independent evaluations and implement supplementary safety measures where necessary.

Stable Diffusion 3.5 Large is released under the Stability Community License, which allows free research and non-commercial use. Commercial usage is free for entities with less than $1 million annual revenue; larger organizations require a separate enterprise license, as outlined in the Stability AI license FAQ. Users retain ownership rights over generated media.

Photorealistic model output demonstrating skills on prompt interpretation

Figure 6. A photorealistic AI-generated image created by Stable Diffusion 3.5 Large, exemplifying its capacity for natural visual depiction from text prompts.

Applications and Use Cases

Intended applications include image generation for design, illustration, and art creation, as well as integration in creative and educational tools, as noted in the Hugging Face model card. The model is also suitable for research into generative AI methods and their societal, ethical, or technical boundaries. Usage must adhere to the Acceptable Use Policy established by Stability AI.

Helpful External Links

About Stable Diffusion 3: The Stable Diffusion 3 series introduces a novel Multimodal Diffusion Transformer (MMDiT) architecture, enhancing text comprehension and image generation capabilities. This advancement enables the models to produce high-quality images with improved prompt adherence and diverse outputs, all while maintaining efficiency suitable for consumer hardware.

More from stabilityai

stabilityai /

Stable Video Diffusion

A latent diffusion model that generates short video clips up to 25 frames from single images using temporal convolution and attention layers.
stabilityai /

Stable Video Diffusion XT

A video generation model that creates coherent sequences from static images or text prompts using latent diffusion architecture.
stabilityai /

Stable Video Diffusion XT 1.1

Video generation model that transforms single images into 25-frame sequences at 1024x576 resolution with controllable motion and camera parameters.
stabilityai /

Stable Video 3D

Stable Video 3D (SV3D) is a generative model based on Stable Video Diffusion that takes in a still image of an object as a conditioning frame, and generates an orbital video of that object.
stabilityai /

Stable Video 4D

A generative video-to-video diffusion model that synthesizes temporally and spatially consistent multi-view video sequences from single input videos.
stabilityai /

Stable Fast 3D

Generates textured 3D meshes with material properties from single images in approximately 0.5 seconds using a transformer-based architecture.
stabilityai /

Stable Diffusion 2

A text-to-image diffusion model offering 768×768 resolution generation with depth conditioning, inpainting capabilities, and 4x upscaling functionality.
stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
stabilityai /

Stable Diffusion XL

A text-to-image diffusion model with 3.5 billion parameters utilizing a two-stage generation pipeline for enhanced image quality and prompt adherence.
stabilityai /

SDXL Turbo

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, only requiring 2-6 steps instead of 20-60.
stabilityai /

ControlNet SDXL Canny

Smaller SDXL ControlNet model for edge detection.
stabilityai /

ControlNet SDXL Depth

Smaller SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Recolor

SDXL ControlNet model for recoloring images.
stabilityai /

Stable Cascade Stage A

A vector quantized GAN encoder that compresses 1024×1024 images to 256×256 discrete tokens as part of a three-stage hierarchical text-to-image pipeline.
stabilityai /

Stable Cascade Stage B

Intermediate latent super-resolution module that upscales compressed text-conditional representations from Stage C to enable high-fidelity image generation.
stabilityai /

Stable Cascade Stage C

Text-to-image diffusion model using three-stage cascaded architecture with 42:1 spatial compression for efficient high-resolution synthesis.
stabilityai /

Stable Audio Open 1.0

Open-source text-to-audio synthesis model with 1.21 billion parameters, trained exclusively on Creative Commons data using latent diffusion architecture.