---
title: 'Create AI Music Videos with Perfect Lip-Sync and Consistency: A Proven Method'
source: 'https://youtube.com/watch?v=N1Uw1VvI8pU'
video_id: 'N1Uw1VvI8pU'
date: 2026-07-24
duration_sec: 1024
channel: 'Arca Artificial by Lordwind Enrique '
---

# Create AI Music Videos with Perfect Lip-Sync and Consistency: A Proven Method

> Source: [Create AI Music Videos with Perfect Lip-Sync and Consistency: A Proven Method](https://youtube.com/watch?v=N1Uw1VvI8pU)

## Summary

This video presents a method for creating music videos entirely with AI, focusing on solving two key problems: character consistency across scenes and accurate lip-syncing. The creator demonstrates a workflow using tools like Midjourney, Suno, and SID 2.0, and emphasizes the importance of character sheets and careful audio segmentation.

### Key Points

- **Two Main Problems in AI Music Videos** [00:03] — Consistency (keeping characters, clothing, and faces the same across takes) and correct lip-syncing (mouth movements matching the song) are the critical challenges.
- **Creating Characters in Midjourney** [01:12] — The creator used Midjourney to design two characters: a singer and an actress. The singer was designed before the song was created.
- **Generating the Song with Suno** [01:28] — Lyrics were input into Suno along with a desired style, and the AI generated multiple songs. One was selected for the video.
- **Character Sheet for Consistency** [02:12] — A 'cheat character' sheet was created for the singer, showing different angles, close-ups, wardrobe breakdown, color palette, and props to provide context for AI consistency.
- **Detailed Character Sheet Prompt** [03:50] — The prompt included visual style, composition, age, height, build, and other data to generate a coherent character sheet.
- **Creating Wardrobe Variations** [04:48] — Based on the lyrics, different outfits were designed for different scenes. A simpler character sheet with four panels was generated for each outfit.
- **Separating Scenes into Two Types** [05:50] — Scenes were divided into those without lip-syncing (acting) and those with lip-syncing (singing). The acting scenes were created first using SID 2.0.
- **Key Rule for Lip-Syncing: Subject Looks at Camera** [07:17] — For successful lip-syncing, the protagonist must always look at the camera, even in wider shots or different environments.
- **Marking Lyrics for Audio Segmentation** [08:31] — The lyrics were marked to identify which phrases correspond to which parts of the song. The audio was exported as MP3 segments for precise timing.
- **Avoid Fast Pacing for Lip-Sync** [10:28] — To achieve accurate lip-sync, the singing pace should be normal, not extremely fast like in some rap songs, as AI struggles with rapid speech.
- **Prompt for Lip-Sync Video Generation** [10:43] — The prompt included 'no music natural sound only', the start frame image, duration (11 seconds), and the lyrics in quotation marks. The result was an imperfect lip-sync with audio.
- **Fixing Continuity Between Segments** [12:24] — When extending a video, the AI sometimes restarts from the initial frame. The fix is to take a screenshot of the final frame, use it as the start frame for the next segment.
- **Final Result and Key Insight** [15:12] — The lip-syncing was perfect, consistency was maintained, and clothing changed according to character sheets. The creator notes that using a real person (like Arcadi) works better than using one's own face because AI doesn't know your natural gestures.

### Conclusion

By using character sheets, careful audio segmentation, and specific prompts, you can achieve consistent characters and accurate lip-syncing in AI-generated music videos. The method works for both AI-generated characters and real people, though real people may yield more natural results.

## Transcript

n't know why. Someone wrote your name [music][singing] before mine. When you try to create a music video 100% with artificial intelligence, you run into two problems.  The first
is consistency, trying to ensure that the character or characters in the video do not change from take to take, that the clothing is the same, that the face is the same.  And this is a very critical problem.  And the second is that since it's
a music video, the lip-syncing or lip-syncing needs to be correct, the mouth shouldn't make strange movements, and it needs to stay in time with the song. That's why I've prepared this video so you have the
solution.  Also below you'll have a timelapse where they'll be listed by chapters.  If you don't want to see a part, you can skip ahead and see the part you want. If this is your first time visiting the channel, I'm Lorwin Enrique, captain
of the Ark, a channel specializing in images or videos with AI.  Let's begin. images or videos with AI.  Let's begin. [scream] I
only thing I did outside of it was create my characters.  In my video there are two characters, one was the singer and the other was an actress who was going to be like the singer's partner, and I did it in M ​​Journey.  Once I had chosen my
singer, because I wanted to imagine him before creating the song, although I already had some knowledge of what I wanted, I went to Suno.  In one, what I did was put the lyrics in the lyrics section , and in the style section I put the style I
, and in the style section I put the style I wanted for my song, and it created wanted for my song, and it created songs for me. according to what I asked for, and out of all of them, I chose only one.
Now that I have the song, and I have my singer, let's work on consistency.  How do we work on consistency?  In this case, the musician Arcadi, who is the main actor or main person of my video, what I
did was create a cheat character for him.  The cheat character is something I've already been explaining on the channel in other videos. In this case, I'll explain it to you again.  The chic character means that it is based on an image and starts to create
on an image and starts to create different angles: front, three-quarter side, back, three-quarter rear.   It will also show you a section where it will be like a close-up to capture the facial features.  I also
broke down a bit the possible psychological profile of the person.  If we put it on her at the prom, you'll get a breakdown of what clothes she has at that moment.  He can also provide props or objects, since he's a singer, because I
'm putting him in the prom, he brings out items that can represent this singer.  Finally, a cinematic portrait brings out your color palette, and also brings out samples of materials.  Why are we bringing this up?
Because in the end we need to have the greatest consistency or give the greatest context of consistency to artificial intelligence.  The promo I used for this, as you can see, is create a character sheet, a premium character based on the
person in the reference image.  In the node I'm placing, notice that if I remove this from here and select, you see that only this image is selected.  I'm setting this
image as the main one, I take out the sheet, the character is called ar, [clears throat] I character design sheet, studio, high-budget animation, similar to a visual production bible.  And here I 'm adding everything I want: the
visual style, the overall composition, the title, I'm adding the age, the height, the build, I'm adding all the data so that this character sheet comes out in the best and most coherent way for my work.
and most coherent way for my work. After he had it, I created Maria, I created a girl with a very marked style, super Latina and as you can see, she uses the same character sheet.  I'm using the same prom tag as well; in
this case, I adapt it to it and it gives me a character sheet that, as you can see, is similar to the other one.  The structure changed a little, but it shows her from the front, from behind, from the side, close-ups of her face, a cinematic portrait,
the color palettes, the breakdown of the wardrobe and the props she could use because she is also associated with music. The second step I took after generating my character sheet was to create a
wardrobe for him.  In this case, my song lasts 2 minutes and 8 seconds.  What did I do before I could do the wardrobe thing?  I started to read the lyrics and saw what might happen in the verse.  What could happen in the chorus?  What could happen?  Choir.  And
so I went to find out more or less, to change her wardrobe, you know?  Because it depends on the situation.  I dress her in different outfits and this enriches my video.  Once I do this, I come to the image generator, I
add another pro and it will give me another kind of character sheet, but a little simpler because I'm setting it to only give me four panels, one front, one back, and others where the face is facing forward and to the side.   So
that?  so I can have that character sheet for that specific clothing.  And I did it with different types of outfits that I was going to use in the video, or at least most of them. When you're ready to put the scenes together, you
can separate them into sections.  In my case, I separated it into two different ones.  How is this possible?  The first one where the singer wasn't going to lip-sync, but they were just going to perform , okay?  The music was going to be in the background and they were going to do
where he simply grabs the microphone, stands up and the music starts playing, like also here where he was acting, where he was grabbing the microphone, she was walking where they meet and then what I
do is a single video, in my case I did it in a single video to mix the three images.  I'm doing all this with Sidans 2.0.  In this case, PROM can't tell you, "Look, it's something you can use because it will
give you a result since it's adapted to what I wanted as a sequence or as a scene, rather. As I told you, I separated it into two. This was the first part, and the second was where he was going to lip-sync, where he was going to sing, where he was going to
connect with the lyrics, and I needed him to keep the timing of the song and for there to be no distortion in his lips. And in the end, it's separated into the same parts as the first part, where there are scenes where he acts, where
the story of the song unfolds. Here there is lip-syncing and the final scene. But let's move on to the second part you want to learn, which is want to learn, which is lip-syncing. The first thing to know is that he, or your
protagonist, is always looking at the camera. It doesn't matter if he's in a car, it does n't matter if he's playing basketball, but he has to be the main focus of the scene. That's what's important, that he's the main focus of the scene. For
example, here the shot is a little wider, but he's there and  When he goes to sing, he's going to lip-sync. Even in this scene where the camera is watching him, he 's in a different environment, a different place, but he's going to look at the camera
and the lip-syncer is going to say, "Ah, this is the person." And I'm telling you this because in this scene I wanted him to be singing as if he were speaking, but he didn't do it. In fact, the song was playing and he wasn't lip-syncing because it's
difficult for him, at least for now. So, as you can see, almost all the scenes I filmed were of him in the studio. And this had a reason. When he was with her, it was bright, but when he was singing, it
was dark. And with this, I wanted to create a contrast between how he felt when he was with her and how he felt when he was alone in the studio. In fact, the lyrics have a message and all that, and that's why I did it this way. Second point, in
my lyrics I had already marked what he was going to say. In fact, you can see it's in bold. The chorus was going to be complete. In verse two  I was going to lengthened it a bit later. And why is it important to know what I was going to
say? Because in an editing program, no matter which one you use—Capcot, Da no matter which one you use—Capcot, Da Vinci, Premiere, whatever—you go to the track, to your song track, and you look for where
he says a phrase. So, you mark—in my case, I mark here, and I can give you an example—I mark here, and I already know that this is the part where he says in the lyrics of my song, "I woke up wanting you."  I don't know why someone
wrote your name before my set. So, if it's in this part, you tell me, I'll come, I'll export it, but I'm not going to export it as an H.264 video, I'm going to go and export it
I'm going to go and export it as an MP3, simply as an MP3.  As you can see, I already have the audio here for the first part, second part, third part, fourth part,
fifth part, and seventh part.  This part and the chorus, which was the third part, I divided into two so that it wouldn't take up more than so many seconds, so that the video would be accurate.  This work is super important because you'll have the
time range you want in your video. And I'm back here, notice that it says part one and this first video will last 11 seconds.  The second part lasts 7 seconds.  The third part will last 9 seconds, which is the first part of the
chorus, and the second part lasts 15 seconds.  Thirdly, if you want to create an imperfect lip-sync, try to make sure that the part you want him to do—the part where he's seen speaking or singing—is at a normal pace,
not an Eminent song where he speaks at 400 words per second. This will result in an inaccurate lip-sync, as it's so fast that the AI ​​won't know how to handle it.  And in the end the promo I used was this one, adding "no music
natural sound only".  I mean, not without music, it has to be a natural sound, it has to be a cinematic scene.  I'll show you the image, I'll connect it so you know what the start frame is.  Even though I 'm placing the image as a
reference, I'm placing it in the prom, which is the initial frame.  Here I am telling you the time it will last, which is 11 seconds, because that is the length of my entire audio, because I need Sidá to know that this is what he is
going to create, he is not going to create 3 seconds more or 3 seconds less.  And I tell him that the singer is going to sing, and I put the part of the lyrics that I want him to sing, in quotation marks.  As you can see, I woke up wanting you and it ends
on set.  And here I am putting it on him.  I woke up wanting you and it ends on set.  and I set it to do it in an imperfect lip-sync with the audio and
then I add certain data. As always, I'll leave the prompts on the blog so you can copy and use them.  I'm leaving it in a general way so you can modify it to your liking.  Here I've put that it
modify it to your liking.  Here I've put that it 's 11 seconds if they give 2.0 and 169. And the result it gives is. I woke up wanting you.  [singing][music] I do n't know why someone wrote your name
Where did I have a problem?  Here, because he created this
because I didn't want to make it 15 seconds because it would be too cut off. So, for audio and rhythm reasons, I only chose 9 seconds, and that's why I divided it into two.  But when I wanted to make a stand video generator, look what it
did, [music] as if it started again from the initial frame and not from how it ended, something that
initial frame and not from how it ended, something that works perfectly.  And in this case, I had put it like, look, it starts from this video, it's an extension, but the audio is the third
part two which lasts 15 seconds and put the same thing.  What can you do in this case to fix your video? You look for the final frame, you take a screenshot of it, which is what I did, I put it here, I generated it in better
put it here, I generated it in better quality and that's where I started and made the quality and that's where I started and made the second video.  [music]
Notice that just one generation, just one generation.  Because?  Because the promo, connecting it this way and using this promo, works to make the lips, and in this way you get a job like this.
this is real from a prom.  [singing][music] I woke up wanting you.  I don't know why. [music] Someone wrote your name [singing] before my thirst.   They
told me Amalaya, here I am.  [music] Look at me . I delivered something [song] that [music] I didn't even understand.  He didn't write it, and my heart obeyed.  [music] If he says I love you, baby, I already love you.  I don't know
if this fire is mine, he lent it to me, but either way, I don't know what I feel, but the Lord already wrote it.  Say love I canight
that [music] I don't know
because I'd like you to go see it on Instagram if you like it and if you want to see it all and that way you support me on that network too and I can only pretty impeccable piece of work.  The lip-syncing was 100% perfect, the consistency was
perfect, and the clothes throughout the video kept changing according to the character sheets I sent at the beginning.  That's why it's important to do them.  And I know you'll face?  Can I do this with my songs?  You can do it the same way, the
same process.  There is no difference between doing it with an AI-generated image, or with your face.  What is the real difference?  It's that the AI ​​doesn't generate you as always tell it.  And I also say this in consultations, one by one, when they
myself."  I say, "The thing is, when you want to stand out, it's because you know who you are, you know how you act, you know what gestures you make, and AI doesn't know that if you make a face movement that you do n't normally make, then you say, "This is
weird."  On the other hand, when you create with a person, like Arcadi for example, who is 100% committed, it does n't matter what he believes, because for me it will be perfect.  I'm not associating it with something I already know.  So this is something
important to keep in mind.  The lips are achieved in this way.  So here's enjoyed this video.  If so, I hope you'll subscribe, hit the bell icon, share this video, and spread the word
also want to invite you to join the Telegram community, which is 100% free.  The link will be in the description .  As you know, this is the artificial long where we share the pleasure of creating quality images and videos
.  If you liked this video, leave a comment.  Before you go, I want to recommend this video here so you can learn a little about visual identity created with artificial intelligence, what works, what doesn't
and here's my face, so you can subscribe much subscribe much faster.  See you in the next video.
