SkyReels-V3 Open-Sources Multimodal AI Video Model, Generates Realistic Video from Single Image

The Kunlun TianGong Skywork AI team has open-sourced SkyReels-V3, a multimodal video generation model capable of creating realistic videos from a single image and text instructions. This release follows previous iterations focused on AI short drama creation (V1) and infinite-length movie generation (V2).
SkyReels-V3 aims to address key challenges in AI video production by integrating three core capabilities within a single architecture. The model's open-source code is available on GitHub, with a research paper detailing its architecture published on arXiv.
Integrated AI Video Capabilities
SkyReels-V3 consolidates functionalities that previously required multiple specialized models. Its primary features include:
Reference Image to Video: This function generates multi-subject videos from one to four reference images and text instructions. It maintains consistent character and product appearance throughout the generated video, preventing random changes in the main subject.
Video Extension: The model can extend a 5-second video clip to 30 seconds, incorporating professional transition effects such as cut-ins, cut-outs, multi-angle switching, shot/reverse shot, and cut-aways. This feature aims to create continuous footage without visual discontinuities.
Audio-Driven Virtual Avatar: By combining a portrait image with an audio segment, SkyReels-V3 can generate minute-long videos where the avatar's lip movements precisely synchronize with the audio. It also supports natural facial expressions and subtle head movements, applicable to real people, cartoon characters, and animal images.

Three panels illustrating SkyReels-V3's features: image-to-video, video extension, and audio-driven avatar.
Enhanced Consistency and Quality
For the "Reference Image to Video" capability, SkyReels-V3 achieves a reference consistency score of 0.6698 and a visual quality score of 0.8119. These metrics, according to the developers, surpass those of mainstream commercial models. The model treats reference images as an "identity contract," ensuring the main character or product remains consistent throughout the video. This allows for the generation of high-fidelity product advertisements or scenarios involving specific individuals, such as Elon Musk.
The model’s video extension function utilizes "unified multi-segment positional encoding" and "robust spatio-temporal modeling." This enables the AI to understand the temporal logic and spatial relationships within video content, resulting in smooth extensions without the typical spatio-temporal distortions seen in other AI-generated videos. It supports 720p resolution, single-shot extensions up to 30 seconds, and various aspect ratios.
In audio-driven virtual avatar generation, SkyReels-V3 recorded an audio-visual synchronization score of 8.18 and a visual quality score of 4.60. These scores are comparable to or exceed those of industry-leading models like OmniHuman 1.5. The model supports 720p, 24fps high-definition video output with phoneme-level audio synchronization and minute-long video generation in a single pass, maintaining identity consistency and continuous action. It also supports multi-person scenarios, allowing characters to naturally switch between speaking and listening states in dialogue.

Human eye reflecting digital data, symbolizing precision and quality in AI video generation.
Open-Source Approach and Industry Impact
The SkyReels research team developed a 200-set test benchmark covering film, e-commerce, and advertising scenarios to evaluate the model's performance. The benchmark included diverse reference image types such as people, animals, objects, and backgrounds.
SkyReels-V3 is fully open-source, with its code available on GitHub. This allows individuals and enterprises to download, deploy locally, and customize the model without incurring API call fees or concerns about data privacy. The developers state that this open-source approach aims to integrate advanced AI video capabilities into existing workflows and foster an ecosystem for further development.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.