Skip to main content
Browse Models

stabilityai

Stable Video 3D

Released

2024-03-18

Family

Stable Video Diffusion

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Model Checkpoint

FP16 · 9.2 GB · sv3d_u.safetensors

Model Checkpoint

FP16 · 9.2 GB · sv3d_p.safetensors

Model Report

Overview

Stable Video 3D (SV3D) is a generative artificial intelligence model developed by Stability AI that specializes in creating orbital videos from a single static image. Building upon the Stable Video Diffusion (SVD) Image-to-Video architecture, SV3D is designed to synthesize short videos that depict an object as if a camera is smoothly circling around it, thereby simulating a 3D view from 2D data. The model enables the transformation of a still image into a multi-perspective visual sequence, which has applications in areas such as visualization, virtual environments, and digital content creation.

An orbital video generated from a still image using SV3D, showing a 360-degree view of a yellow-orange rubber duck.

Figure 1. SV3D output: A 360-degree orbital video generated from a single still image, illustrating the model's core capability.

Technical Capabilities

SV3D is engineered to generate short orbital videos, typically consisting of 21 frames at a resolution of 576x576 pixels, by conditioning on a single input image of an object. The model introduces two primary variants with distinct control properties: SV3D_u, which autonomously determines the camera path to create an orbital video around the subject without external guidance, and SV3D_p, which accepts specified camera trajectories, affording users enhanced control over the viewpoint sequencing and path of the virtual camera.

Both variants are capable of producing coherent object-centric videos that simulate a smooth rotation around the input subject, contributing to enhanced visualizations and interactive digital experiences. The system is fine-tuned from the base SVD Image-to-Video model, establishing continuity and improvement in Stability AI's video diffusion capabilities.

Model Architecture

The architecture of SV3D is based on the diffusion model approach introduced by Stable Video Diffusion, which utilizes probabilistic sampling to incrementally refine noisy video predictions into realistic video frames. SV3D adapts this architecture to operate with both a single image as input and, for SV3D_p, an explicit camera path specification. This adaptation allows the model to learn complex 3D-aware video representations, simulating novel views of the same object consistent with a camera's motion.

Details of the model's architecture, including network configurations and conditioning mechanisms, are described in the official SV3D technical report and the associated project page, ensuring scientific transparency and reproducibility.

Datasets and Training Procedure

The SV3D model is trained using renderings derived from the Objaverse 1.0 dataset, a large-scale collection of 3D models with diverse objects. Stability AI employed an enhanced rendering process to recreate the characteristics and distribution of real-world images, improving the model's generalization ability. The training data was further refined to maximize fidelity and diversity while adhering to the CC-BY license, facilitating responsible and open research.

This training methodology enables SV3D's capability to generate multi-angle video sequences from a single visual reference.

Applications and Use Cases

SV3D is tailored for applications requiring the creation of short, object-centric orbital videos from static images. Primary use cases involve generating 3D-like representations for virtual and augmented reality content, product visualizations in digital marketplaces, and creative content development where multi-angle object views are desirable. The model's output can facilitate rapid prototyping and visualization in contexts where acquiring complete 3D scans or multiple photographs would be impractical.

Limitations and Ethical Considerations

SV3D's outputs are the result of generative processes conditioned on static images and, optionally, a specified camera path. As such, the model is not designed to produce factual or verifiable representations of specific real-world objects, events, or persons. Its use is limited to creative and illustrative purposes and must align with the Stability AI Acceptable Use Policy. Additionally, factual accuracy with respect to real-world identities or scenarios is not guaranteed due to the synthetic nature of the output.

SV3D is released under the Stability AI Community License, which governs its use for research and non-commercial activities. Separate licensing is required for commercial applications, as detailed in the commercial license information.

Further Resources

For more comprehensive technical documentation, reference implementations, and project updates, the following resources are available:

About Stable Video Diffusion: Stable Video Diffusion is a family of AI models that generate short videos from a single image by leveraging latent diffusion techniques to produce high-quality, temporally consistent frames. This approach enables efficient video generation with customizable frame rates, surpassing leading closed models in user preference studies.

More from stabilityai

stabilityai /

Stable Diffusion 2

A text-to-image diffusion model offering 768×768 resolution generation with depth conditioning, inpainting capabilities, and 4x upscaling functionality.
stabilityai /

Stable Diffusion 1.1

A latent text-to-image diffusion model trained on LAION datasets that generates 512×512 images from natural language prompts using compressed latent space processing.
stabilityai /

Stable Diffusion 1.5

Text-to-image diffusion model trained on LAION dataset subset, generating 512x512 images from natural language prompts using latent space processing.
stabilityai /

Stable Diffusion XL

A text-to-image diffusion model with 3.5 billion parameters utilizing a two-stage generation pipeline for enhanced image quality and prompt adherence.
stabilityai /

SDXL Turbo

SDXL variant using the Adversarial Diffusion Distillation (ADD) technique to generate images very quickly, only requiring 2-6 steps instead of 20-60.
stabilityai /

ControlNet SDXL Canny

Smaller SDXL ControlNet model for edge detection.
stabilityai /

ControlNet SDXL Depth

Smaller SDXL ControlNet model for depth generation.
stabilityai /

ControlNet SDXL Recolor

SDXL ControlNet model for recoloring images.
stabilityai /

Stable Diffusion 3.5 Large

This 8-billion parameter model brings improvements to SD3 in customizability, efficiency, and diversity. This Large variant was designed for professional use cases at 1 megapixel resolution.
stabilityai /

Stable Diffusion 3.5 Turbo

This latest iteration in the Stable diffusion family brings improvements over SD3 in customizability, efficient performance, and diverse outputs. The Turbo variant is distilled from the Large version to generate images in just 4 steps.
stabilityai /

Stable Cascade Stage A

A vector quantized GAN encoder that compresses 1024×1024 images to 256×256 discrete tokens as part of a three-stage hierarchical text-to-image pipeline.
stabilityai /

Stable Cascade Stage B

Intermediate latent super-resolution module that upscales compressed text-conditional representations from Stage C to enable high-fidelity image generation.
stabilityai /

Stable Cascade Stage C

Text-to-image diffusion model using three-stage cascaded architecture with 42:1 spatial compression for efficient high-resolution synthesis.
stabilityai /

Stable Audio Open 1.0

Open-source text-to-audio synthesis model with 1.21 billion parameters, trained exclusively on Creative Commons data using latent diffusion architecture.