CagliostroLab
Animagine XL
Released
2024-03-24
Family
Stable Diffusion XL
Type
Fine-Tuned Model
Downloads
Model Report
Overview
Animagine XL is an open-source series of anime-themed text-to-image generative models created by Cagliostro Research Lab in collaboration with SeaArt.ai. Fine-tuned from Stable Diffusion XL, Animagine XL specializes in producing high-resolution, detailed anime-style illustrations from descriptive textual prompts. It is designed to enhance the synthesis of character art, improve prompt interpretation, and more faithfully render complex anatomical features such as hands, which present particular challenges for generative AI models.
Model Features and Prompting Strategy
Animagine XL employs a variety of features and architectural refinements tailored for the anime art domain. The model integrates advanced prompt parsing strategies—using a structured tag ordering inspired by NovelAI tag ordering documentation—to achieve consistent and accurate character synthesis. Recommended prompts typically begin with the number and gender of characters, followed by character name, series, and additional descriptive tags. This method enables more precise interpretation of user intentions and facilitates the rendering of both iconic and original anime characters.
The model supports a diverse set of tags affecting output characteristics, including quality modifiers (such as "masterpiece" or "good quality"), rating tags (for content control, such as "safe" or "sensitive"), and art era modifiers (such as "newest" or "oldest"). From version 3.1 onwards, aesthetic evaluation tags, derived from a dedicated Vision Transformer (ViT) classifier trained on anime art, can guide outputs towards higher visual appeal. Multi-aspect ratio generation is also supported, covering square, portrait, and landscape formats at various resolutions.
Training Data and Technical Foundations
The architecture of Animagine XL is grounded in diffusion-based generative modeling, utilizing the Stable Diffusion XL base and further fine-tuned through proprietary methods. The model employs a specialized VAE, madebyollin/sdxl-vae-fp16-fix, to improve encoding and decoding of high-resolution images.
Training for Animagine XL 3.0 was conducted on approximately 1.2 million images during the initial feature alignment stage, using additional curated image subsets for subsequent refinement and aesthetic tuning, for a total of roughly 2.1 million images across versions 2.0 and 3.0. Training processes incorporated custom scripts adapted from kohya-ss/sd-scripts and leveraged advanced label association techniques to optimize tag learning. Hyperparameters were tuned in multiple training stages, adjusting learning rates and batch sizes to balance stability, convergence, and expressiveness.
Aesthetic evaluation tags were established using the aesthetic-shadow-v2 ViT classifier to score and prioritize visually appealing outputs in the training pipeline, resulting in more refined generations and consistent character appeal.
Applications and Evaluations
Animagine XL primarily caters to anime artists, illustrators, and enthusiasts seeking to create character art, fan art, or original concept pieces from descriptive text. The model demonstrates proficiency in generating recognizable anime characters, often requiring only prompt-based specification rather than supplementary fine-tuning via LoRA techniques. Its improvements in anatomical rendering, especially of hands, address previously noted deficits within AI art generation.

Figure 1. Sample output illustrating improved hand anatomy; prompt: 'smiling girl in school uniform waving'.

Figure 2. Model output highlighting fine finger gesture rendering; prompt specifies animated girl with expressive pose.
Quantitatively, Animagine XL 3.1 has received high ratings on Civitai. Empirical analysis demonstrates that selection of prompt structure and CFG (Classifier-Free Guidance) scale parameter materially affects output clarity and fidelity.

Figure 3. Grid comparison illustrating the effect of different CFG scale values on generation sharpness and coherence. Lower CFG produces blurrier images; higher values yield sharper, more detailed results.
Limitations and Considerations
Animagine XL is optimized for anime aesthetics rather than photorealism, and its design makes it less suitable for tasks outside the anime domain. While improvements have been made, occasional anatomical inconsistencies can still arise, particularly with complex hand poses. Character generation is most effective when prompts use Danbooru-style structured tags; natural language prompts may yield less reliable results. The prevalence of high-quality, mature-rated images in the training data can sometimes result in incidental generation of sensitive material unless properly constrained with negative prompts and content-specific tags.
The dataset, while extensive, does not exhaustively cover the full breadth of anime character design, which may limit representation of obscure or newly introduced characters without additional fine-tuning. The training process for Animagine XL 3.0 also encountered challenges in distributed gradient synchronization, resulting in partial updates during multi-GPU training, though these were noted as areas for future optimization.

Figure 4. Sample depicting enhanced hand gesture synthesis, addressing a known challenge in anime-style text-to-image generation.
Model Versions and Licensing
Animagine XL has seen several developmental milestones. Version 2.0 laid the groundwork for aesthetic optimization, while version 3.0 expanded the training set and introduced advanced tag handling. The latest release, Animagine XL 3.1, further refines model performance with improved aesthetic tagging and enhanced prompt control.
The model and its weights are made available under the Fair AI Public License 1.0-SD (FAIPL-1.0-SD), which defines conditions for its use and modification. This license compels redistribution of source code for network-accessible modifications and mandates release of derivative works under compatible licensing conditions.
Sample Outputs

Figure 5. Generated illustration: Asuka Langley Soryu on an ornate throne. Prompt: 'Asuka Langley Soryu, sitting regally on a throne, dramatic, detailed, masterpiece'.

Figure 6. Dynamic output showing detailed character and motion effects; prompt: 'determined maid, wielding knives in action, high detail'.

Figure 7. Vibrant illustration inspired by My Hero Academia; prompt: 'Izuku Midoriya, dynamic pose, emitting energy, anime style, detailed'.

Figure 8. Anime portrait with subtle lighting and realistic clothing details; prompt: 'young person, hoodie, glasses, soft lighting'.
Further Resources
- Animagine XL V3.1 on Civitai
- Cagliostro Research Lab – Hugging Face
- Animagine XL 3.0 on Hugging Face
- Animagine XL 2.0 on Hugging Face
- Fair AI Public License 1.0-SD details
- NovelAI Image Tag Ordering Reference
- Aesthetic-Shadow-v2 Model (ViT aesthetic classifier)
- Official Cagliostro Lab Website
- SeaArt.ai (collaborator)
- Project Community Discord
- Cagliostro Lab SD Scripts
More in the Stable Diffusion XL Family
Stable Diffusion XL
SDXL Turbo
SDXL Lightning
OpenDalle
Yamer's Realistic
AlbedoBase XL
Juggernaut XL
Realistic Vision XL
New Reality XL
Realism Engine SDXL
Nightvision XL
Dreamshaper XL
Pony Diffusion V6 XL
ControlNet SDXL Diffusers Canny
ControlNet SDXL Canny
ControlNet SDXL Diffusers Depth
ControlNet SDXL Depth
ControlNet SDXL Recolor
ControlNet SDXL IP Adapter
ControlNet SDXL Open Pose
Compatible Apps

ComfyUI
A node-based workflow builder for advanced image and video generation, ideal for custom pipelines, fine control, and power users.
Web UI · API
Image Generation · Video Generation

Stable Diffusion WebUI Forge
A faster, more experimental Stable Diffusion WebUI variant focused on improved resource use, quicker inference, and modern model support.
Web UI · API
Image Generation

Stable Diffusion Web UI
A full-featured Stable Diffusion interface with deep controls for prompting, inpainting, extensions, and advanced image workflows.
Web UI · API
Image Generation · Video Generation

Fooocus
A beginner-friendly image generator focused on strong defaults, with built-in inpainting, outpainting, upscaling, and image prompting.
Web UI · API
Image Generation · Beginner Friendly

Kohya's GUI
Train LoRAs and fine-tunes for Stable Diffusion and FLUX with a popular GUI for Kohya-based training workflows.
Web UI · API
Fine-Tuning · Image Generation