Skip to main content
Browse Models

CagliostroLab

Animagine XL

Released

2024-03-24

Family

Stable Diffusion XL

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · 6.3 GB · animagine-xl.fp16.safetensors

Model Report

Overview

Animagine XL is an open-source series of anime-themed text-to-image generative models created by Cagliostro Research Lab in collaboration with SeaArt.ai. Fine-tuned from Stable Diffusion XL, Animagine XL specializes in producing high-resolution, detailed anime-style illustrations from descriptive textual prompts. It is designed to enhance the synthesis of character art, improve prompt interpretation, and more faithfully render complex anatomical features such as hands, which present particular challenges for generative AI models.

Introductory showcase video illustrating the output quality and diversity of Animagine XL V3.1. · Source

Model Features and Prompting Strategy

Animagine XL employs a variety of features and architectural refinements tailored for the anime art domain. The model integrates advanced prompt parsing strategies—using a structured tag ordering inspired by NovelAI tag ordering documentation—to achieve consistent and accurate character synthesis. Recommended prompts typically begin with the number and gender of characters, followed by character name, series, and additional descriptive tags. This method enables more precise interpretation of user intentions and facilitates the rendering of both iconic and original anime characters.

The model supports a diverse set of tags affecting output characteristics, including quality modifiers (such as "masterpiece" or "good quality"), rating tags (for content control, such as "safe" or "sensitive"), and art era modifiers (such as "newest" or "oldest"). From version 3.1 onwards, aesthetic evaluation tags, derived from a dedicated Vision Transformer (ViT) classifier trained on anime art, can guide outputs towards higher visual appeal. Multi-aspect ratio generation is also supported, covering square, portrait, and landscape formats at various resolutions.

Training Data and Technical Foundations

The architecture of Animagine XL is grounded in diffusion-based generative modeling, utilizing the Stable Diffusion XL base and further fine-tuned through proprietary methods. The model employs a specialized VAE, madebyollin/sdxl-vae-fp16-fix, to improve encoding and decoding of high-resolution images.

Training for Animagine XL 3.0 was conducted on approximately 1.2 million images during the initial feature alignment stage, using additional curated image subsets for subsequent refinement and aesthetic tuning, for a total of roughly 2.1 million images across versions 2.0 and 3.0. Training processes incorporated custom scripts adapted from kohya-ss/sd-scripts and leveraged advanced label association techniques to optimize tag learning. Hyperparameters were tuned in multiple training stages, adjusting learning rates and batch sizes to balance stability, convergence, and expressiveness.

Aesthetic evaluation tags were established using the aesthetic-shadow-v2 ViT classifier to score and prioritize visually appealing outputs in the training pipeline, resulting in more refined generations and consistent character appeal.

Applications and Evaluations

Animagine XL primarily caters to anime artists, illustrators, and enthusiasts seeking to create character art, fan art, or original concept pieces from descriptive text. The model demonstrates proficiency in generating recognizable anime characters, often requiring only prompt-based specification rather than supplementary fine-tuning via LoRA techniques. Its improvements in anatomical rendering, especially of hands, address previously noted deficits within AI art generation.

Anime-style girl waving with five distinct fingers

Figure 1. Sample output illustrating improved hand anatomy; prompt: 'smiling girl in school uniform waving'.

Anime girl with raised index finger, showing distinct gesture

Figure 2. Model output highlighting fine finger gesture rendering; prompt specifies animated girl with expressive pose.

Quantitatively, Animagine XL 3.1 has received high ratings on Civitai. Empirical analysis demonstrates that selection of prompt structure and CFG (Classifier-Free Guidance) scale parameter materially affects output clarity and fidelity.

Comparison grid of different CFG scale settings

Figure 3. Grid comparison illustrating the effect of different CFG scale values on generation sharpness and coherence. Lower CFG produces blurrier images; higher values yield sharper, more detailed results.

Limitations and Considerations

Animagine XL is optimized for anime aesthetics rather than photorealism, and its design makes it less suitable for tasks outside the anime domain. While improvements have been made, occasional anatomical inconsistencies can still arise, particularly with complex hand poses. Character generation is most effective when prompts use Danbooru-style structured tags; natural language prompts may yield less reliable results. The prevalence of high-quality, mature-rated images in the training data can sometimes result in incidental generation of sensitive material unless properly constrained with negative prompts and content-specific tags.

The dataset, while extensive, does not exhaustively cover the full breadth of anime character design, which may limit representation of obscure or newly introduced characters without additional fine-tuning. The training process for Animagine XL 3.0 also encountered challenges in distributed gradient synchronization, resulting in partial updates during multi-GPU training, though these were noted as areas for future optimization.

Anime girl making peace sign, model output showing improvement in gesture rendering

Figure 4. Sample depicting enhanced hand gesture synthesis, addressing a known challenge in anime-style text-to-image generation.

Model Versions and Licensing

Animagine XL has seen several developmental milestones. Version 2.0 laid the groundwork for aesthetic optimization, while version 3.0 expanded the training set and introduced advanced tag handling. The latest release, Animagine XL 3.1, further refines model performance with improved aesthetic tagging and enhanced prompt control.

The model and its weights are made available under the Fair AI Public License 1.0-SD (FAIPL-1.0-SD), which defines conditions for its use and modification. This license compels redistribution of source code for network-accessible modifications and mandates release of derivative works under compatible licensing conditions.

Sample Outputs

Anime-style depiction of Asuka Langley Soryu, output from Animagine XL

Figure 5. Generated illustration: Asuka Langley Soryu on an ornate throne. Prompt: 'Asuka Langley Soryu, sitting regally on a throne, dramatic, detailed, masterpiece'.

Anime-style maid character in motion, holding knives

Figure 6. Dynamic output showing detailed character and motion effects; prompt: 'determined maid, wielding knives in action, high detail'.

Anime character inspired by Izuku Midoriya surrounded by energy

Figure 7. Vibrant illustration inspired by My Hero Academia; prompt: 'Izuku Midoriya, dynamic pose, emitting energy, anime style, detailed'.

Anime-style portrait of young individual in a hoodie

Figure 8. Anime portrait with subtle lighting and realistic clothing details; prompt: 'young person, hoodie, glasses, soft lighting'.

About Stable Diffusion XL: The SDXL family of AI models represents a significant technological advancement in text-to-image generation, featuring a substantial increase in parameters—3.5 billion for the base model and 6.6 billion for the ensemble—alongside a dual-stage architecture that includes a base model and a refiner, enabling the creation of high-resolution (1024x1024) images with enhanced detail, color accuracy, and style versatility.

More in the Stable Diffusion XL Family

stabilityai /

Stable Diffusion XL

A text-to-image diffusion model with 3.5 billion parameters utilizing a two-stage generation pipeline for enhanced image quality and prompt adherence.
stabilityai /

SDXL Turbo

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, only requiring 2-6 steps instead of 20-60.
ByteDance /

SDXL Lightning

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, with multiple variants for different numbers of steps (1-8) and a more permissive license than SDXL Turbo.
dataautogpt3 /

OpenDalle

SDXL model focused on prompt adherence and semantic understanding, with a stated aim of achieving Dalle-3 level understanding of prompts.
Yamer /

Yamer's Realistic

SDXL fine-tune capable of generating realistic images of people and landscapes.
albedobond /

AlbedoBase XL

SDXL fine-tune resulting from merging 300+ top community models, strong performance across a wide range of image types.
KandooAI /

Juggernaut XL

A popular versatile model finetuned on SDXL by KandooAI and RunDiffusion. This version features improved prompt adherence due to an innovative GPT-4 captioning system built by LEOSAM.
SG_161222 /

Realistic Vision XL

SDXL fine-tune optimizing for generating photorealistic people, animals, and landscapes. In addition to training, this model is a merge of over 10 other SDXL models aimed at realism.
ALIENHAZE /

New Reality XL

Merge of multiple SDXL checkpoints and LoRAs, with impressive breadth and consistency in generating photorealistic images.
razzz /

Realism Engine SDXL

SDXL checkpoint optimized for photorealistic generation of a wide range of humans.
SoCalGuitarist /

Nightvision XL

Photography-focused SDXL checkpoint with strong prompt adherence and versatile output.
Lykon /

Dreamshaper XL

Fine-Tuned on SDXL, this is a general purpose model designed for photos, art, anime, and manga. This Lightning version can generate quality images in few steps.
PurpleSmartAI /

Pony Diffusion V6 XL

A fine-tuned Stable Diffusion XL model trained on 2.6 million images for generating cartoon, anime, and anthropomorphic character art.
diffusers /

ControlNet SDXL Diffusers Canny

SDXL ControlNet model for edge detection.
stabilityai /

ControlNet SDXL Canny

Smaller SDXL ControlNet model for edge detection.
diffusers /

ControlNet SDXL Diffusers Depth

SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Depth

Smaller SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Recolor

SDXL ControlNet model for recoloring images.
h94 /

ControlNet SDXL IP Adapter

SDXL ControlNet model for conditioning on an image prompt.
thibaud /

ControlNet SDXL Open Pose

SDXL ControlNet model for copying human poses.