daVinci-MagiHuman
Open-source AI model that turns a portrait photo and text/audio into a lip-synced talking video in one step.
| What is it | Open-source AI model that turns a portrait photo and text/audio into a lip-synced talking video in one step. |
|---|---|
| Pricing | Paid — from $10/mo |
| Free tier | No |
| Platform | Web Application |
| Best for | Creating AI-powered video avatars for presentations, Generating multilingual talking-head content for social media |
| Domain registered | 2026 |
Data updated April 1, 2026
What does daVinci-MagiHuman do?
DaVinci MagiHuman is an open-source AI model that creates talking-head videos from a single portrait photo. You provide a clear image of a face and either a text script or an audio file. The model then generates a video where the person in the photo appears to speak, with lip movements that are synchronized to the provided audio. It does this in one unified process, producing both the video frames and the matching audio together, instead of using separate systems for speech and animation.
The tool is a 15-billion parameter model developed by Sand.ai and the GAIR Lab at Shanghai Jiao Tong University. Its main technical distinction is its single-stream Transformer architecture, which denoises video and audio tokens simultaneously. This approach aims for better lip-sync quality and efficiency. The model is released under the Apache 2.0 license, meaning the weights are free to download, inspect, and use for commercial purposes. You can run it via a hosted web demo, or self-host it by downloading the checkpoints from Hugging Face or cloning the GitHub repository.
DaVinci MagiHuman is useful for developers and researchers experimenting with generative video AI, as well as creators who need to produce simple, lip-synced avatar videos without complex editing. It's a practical option for generating explainer videos, personalized messages, or prototype content where having an open-source, self-hostable model is a priority over cinematic quality. For best results, use a clear, front-facing portrait with good lighting.
Key features
What makes it stand outWho is daVinci-MagiHuman for?
Who benefits most from this toolPricing
Basic
- 1,000 credits
- Approx 16 x daVinci-MagiHuman video generations/month (120 credits each)
- Approx 9 x daVinci-MagiHuman HD video generations/month (200 credits each)
- Standard Speed
- Email Support
- Access to daVinci-MagiHuman Video Generation Models
- Guided Prompt Builder & Presets
Pro
- 2,000 credits
- Approx 33 x daVinci-MagiHuman video generations/month (120 credits each)
- Approx 19 x daVinci-MagiHuman HD video generations/month (200 credits each)
- Standard Speed
- Email Support
- Access to daVinci-MagiHuman Video Generation Models
- Guided Prompt Builder & Presets
- Priority Processing
- Priority Support
- Batch Background Removal (Beta)
- Growing Template Library
- Future Access to New daVinci-MagiHuman AI Models
Max
- 5,000 credits
- Approx 49 x daVinci-MagiHuman video generations/month (120 credits each)
- Approx 29 x daVinci-MagiHuman HD video generations/month (200 credits each)
- Standard Speed
- Email Support
- Access to daVinci-MagiHuman Video Generation Models
- Guided Prompt Builder & Presets
- Priority Processing
- Priority Support
- Batch Background Removal (Beta)
- Growing Template Library
- Future Access to New daVinci-MagiHuman AI Models
- Highest Priority
- Dedicated Support
- Future Multi-Model Comparison Workspace
Trust & presence
Gallery
Click any image to enlargeAlternatives in Video Generator
AI tool that animates photos and videos to make them talk or sing with realistic lip sync.
Upload a portrait photo, add audio or text, and generate a realistic talking video with perfect lip sync.
Upload a photo and audio to create a perfectly lip-synced talking video with natural facial expressions in minutes.
AI video generator that makes any person or character lip-sync to your text or audio input.
Upload a video and audio to generate a lifelike talking video with perfectly synced lip movements in minutes.
AI video generator — turn photos into talking avatars with lip-sync, voice cloning, and multiple animation styles
AI video generator — turn a single image or audio file into a lifelike talking avatar video with lip sync.
AI lip-sync generator — upload a video/image and audio to create realistic talking avatars with perfect synchronization
Similar tools
Upload a photo or video, add audio, and generate a talking avatar with synced lip movements in seconds.
AI tool that animates photos and videos to make faces talk with perfectly synced lip movements.