Skip to main content
Browse Models

Photographer

Photon

Released

2023-06-05

Family

Stable Diffusion 1

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · photon.safetensors

Model Report

Overview

Photon is a generative artificial intelligence model designed for the creation and enhancement of photorealistic images. Developed as a checkpoint merge based on the Stable Diffusion 1.5 framework, Photon integrates various training adaptations and model refinements to deliver outputs reflecting its training objectives with minimal user prompting. The model is noted for its image refinement abilities, versatility in generative tasks, and suitability for further tuning using low-rank adaptation methods.

Model Architecture and Development

Photon's architecture is rooted in the Stable Diffusion 1.5 latent diffusion model, a widely adopted open-source text-to-image generation system. The Photon checkpoint is classified as a checkpoint merge, having drawn from an assortment of prior model versions and fine-tuned LORA modules, each tailored to specific visual attributes and subject matter.

During its development, the model’s creator—operating under the pseudonym "Photographer"—employed a chaotic process, first merging earlier trained models, then iteratively training LORA adapters on AI-generated, photorealistic datasets. These LORA modules were integrated back into the model using dynamically weighted blending strategies to address specific representational shortcomings, particularly in the depiction of hands. While some resolution was achieved, limitations in anatomical accuracy persisted in initial releases.

Training Methods and Data

The core refinements in Photon are driven by low-rank adaptation (LORA) methods that facilitate efficient retraining and merging across thematic domains. Much of the training involved curating and leveraging AI-generated photorealistic images rather than employing large-scale, human-annotated collections.

The project's stated ambition was to scale the training dataset to include between 5,000 and 50,000 high-quality, AI-crafted photorealistic samples, with the ultimate goal of further automating the refinement and blending process. This approach emphasizes adaptability, making Photon particularly effective for developing new custom LORA modules tailored to specific stylistic or semantic requirements.

Technical Capabilities and Features

Photon’s principal function is the generation of photorealistic imagery. Its outputs exhibit characteristics of photorealism, though users note a distinction between near-photorealism and photographic fidelity. The model functions as a refiner, capable of transforming certain types of visually unrefined images into outputs consistent with photorealistic styles. This refinement occurs with minimal reliance on complex or verbose prompting.

Photon is also recognized for its robust performance in pseudo image-to-image (IMGtoIMG) tasks. In these contexts, the model consistently produces realistic outputs when low denoising settings are selected, while high redrawing intensities may introduce a stylized, two-dimensional effect inconsistent with photorealism.

Another notable feature of Photon is its compatibility with further LORA-based training, making it suitable for users seeking to customize or extend its capabilities for niche visual effects or subject domains.

Applications and Use Cases

Photon serves several primary uses in the field of generative AI. It is frequently utilized to produce photorealistic images based on textual descriptions for content creation. The model’s image-refining capability is applied in post-processing pipelines, where AI-generated imagery can be transformed into compositions reflecting photorealistic characteristics without extensive manual intervention.

Additionally, Photon is widely employed as a foundation for LORA training, enabling targeted fine-tuning by researchers and artists. Its capabilities in image-to-image workflows allow for the realistic transformation or enhancement of existing images, contributing to its applicability within digital media and content creation domains.

Performance, Limitations, and Community Reception

Since its publication on June 5, 2023, the model has been used for image generation. The model file, in fp16 pruned format, occupies 1.99 GB, reflecting its modeling scale.

Photon exhibits versatility, responsiveness to prompt variations, and effective refinement performance. However, certain limitations have been documented. Users have reported persistent challenges with accurate hand generation, despite ongoing efforts to address these anatomical issues through iterative LORA mixing. When employing high redrawing (denoising) settings in the image-to-image pipeline, results may skew towards a flat, two-dimensional aesthetic, eroding photorealistic qualities.

While the license for the model aligns with the CreativeML Open RAIL-M standard, an addendum provides additional guidance on redistribution and responsible use.

Legal Information and Licensing

Photon is released under the CreativeML Open RAIL-M License, which mandates responsible and ethical application of the model and its derivatives. Users are encouraged to consult the accompanying license addendum for detailed stipulations on permissible and restricted uses.

Helpful Links

About Stable Diffusion 1: Stable Diffusion is an open-source text-to-image generative AI model that transforms textual adminDescriptions into corresponding images. Technologically, it employs a latent diffusion model architecture, enhancing computational efficiency by performing diffusion processes in a compressed latent space, which enables high-quality image generation with reduced resource requirements.

More in the Stable Diffusion 1 Family

stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
prompthero /

OpenJourney v4

SD 1.5 fine-tuned on 124k+ additional images generated with Midjourney v4, leading to results that resemble this other closed-source image generation model.
KandooAI /

Juggernaut

Popular SD 1.5 fine-tune with capability to produce detailed images of a versatile breadth of subjects.
wavymulder /

Analog Diffusion

SD 1.5 fine-tuned on a diverse set of analog images, yielding a vintage photographic look.
Lykon /

Dreamshaper

SD 1.5 fine-tune with strong art generation ability and a broad generalist capabilities.
SG_161222 /

Realistic Vision

SD 1.5 fine-tune specialized in creating photorealistic portraits of humans.
Meina /

Meina Mix

Model resulting for merging 7 different anime-focused SD 1.5 checkpoints.
epinikion /

epiCRealism

Popular SD 1.5 fine-tune with high competence in translating simple text prompts into realistic images of people.
Lykon /

Absolute Reality

One of the top SD 1.5 variant for generating life-like images of people and objects.
Cyberdelia /

Cyber Realistic

Versatile photorealistic SD 1.5 fine-tune capable of generating a wide range of convincing photographic images.
Merjic /

MajicMIX Realistic

Popular SD 1.5 photorealism fine-tune with training data weighted on people of asian descent.
epinikion /

epiCPhotoGasm

A Stable Diffusion 1.5-based checkpoint model designed for photorealistic image generation with simplified prompting and demographic diversity.
lllyasviel /

ControlNet SD 1.5 Canny

SD 1.5 ControlNet model to replicate the composion of a source image using edge-detection.
lllyasviel /

ControlNet SD 1.5 IP2P

SD 1.5 ControlNet trained with pixel-to-pixel instruction.
lllyasviel /

ControlNet SD 1.5 Depth

SD 1.5 ControlNet model to replicate the depth of a source image.
lllyasviel /

ControlNet SD 1.5 MLSD

SD 1.5 ControlNet model to detect straight-lines, useful for architecture and man-made objects.
lllyasviel /

ControlNet SD 1.5 Normal

SD 1.5 ControlNet model to replicate the depth of a source image, with additional surface details and geometry.
lllyasviel /

ControlNet SD 1.5 Open Pose

SD 1.5 ControlNet model for copying human poses.
lllyasviel /

ControlNet SD 1.5 Scribble

SD 1.5 ControlNet model for converting sketches to images.
lllyasviel /

ControlNet SD 1.5 Segmentation

SD 1.5 ControlNet model for detecting and segmenting distinct parts of images to use in the generation.
lllyasviel /

ControlNet SD 1.5 Soft Edge

SD 1.5 ControlNet model to detect soft-edges, especially useful for recoloring and stylizing.
lllyasviel /

ControlNet SD 1.5 Inpaint

SD 1.5 ControlNet model trained with image inpainting.
lllyasviel /

ControlNet SD 1.5 Line Art

SD 1.5 ControlNet model trained with line art generation.
lllyasviel /

ControlNet SD 1.5 Lineart Anime

SD 1.5 ControlNet model trained with anime line art generation.
lllyasviel /

ControlNet SD 1.5 Shuffle

SD 1.5 ControlNet model trained with image shuffling.
lllyasviel /

ControlNet SD 1.5 Tile

SD 1.5 ControlNet model trained with image tiling.
tencent /

ControlNet 1.5 IP Adapter

SD 1.5 ControlNet model for conditioning on an image prompt.
tencent /

ControlNet 1.5 QR Code

SD 1.5 ControlNet model for generating stylized QR codes.