---
title: 'How to Make AI Videos That Really Look Like You (Full Guide)'
source: 'https://youtube.com/watch?v=st7NRTxeKC8'
video_id: 'st7NRTxeKC8'
date: 2026-08-12
duration_sec: 793
channel: 'Youri van Hofwegen'
---

# How to Make AI Videos That Really Look Like You (Full Guide)

> Source: [How to Make AI Videos That Really Look Like You (Full Guide)](https://youtube.com/watch?v=st7NRTxeKC8)

## Summary

This video provides a comprehensive workflow for creating AI-generated videos that feature a consistent, realistic version of the creator's face. The process leverages OpenArt's GPT Image 2 model for character creation and reference images, SeaArt 2 for video generation, and CapCut for final editing, with a focus on maintaining facial consistency across all shots.

### Key Points

- **Challenge of AI Video Consistency** [00:02] — The video begins by acknowledging the difficulty of making AI videos that look like the creator, setting the stage for a workflow that ensures facial consistency.
- **Creating a Character in OpenArt** [00:27] — The first step is to create a character in OpenArt using the 'start from an image' option. Upload multiple photos of yourself with a clean background, neutral lighting, and no glasses or accessories to ensure the AI captures facial details accurately.
- **Generating Reference Images with GPT Image 2** [01:19] — Use GPT Image 2 for photorealistic reference images. In the reference section, select the character and use the '@' symbol in the prompt to anchor facial details. Choose a 16:9 ratio for a cinematic feel and select quality (low, medium, high) based on needs.
- **Creating a Second Character** [03:04] — For a second character, you can skip the separate character building step and instead use the word 'photorealistic' in the prompt. This small addition significantly pushes the model toward realism, making the output look like it came from a real camera.
- **Building a Matching Location with Claude** [04:10] — To create a location that matches the characters, use Claude to generate a prompt. Upload the reference images and describe the desired location (e.g., a Japanese dojo), asking Claude to keep the same visual style. This ensures the location's lighting, color palette, and mood match the characters.
- **Animating with SeaArt 2** [06:23] — In the video section, select SeaArt 2 as the AI model. Drop in the character and location reference images. Use the '@' symbol in the prompt to reference each character and location. Set duration to 15 seconds, resolution to 1080p, and aspect ratio to 16:9.
- **Using Video Reference for Continuity** [08:17] — For subsequent scenes, add a video reference of the previous scene. This ensures the movement, character positions, and mood stay continuous. Keep the character and location images as well.
- **Structuring Prompts for SeaArt** [08:45] — Structure prompts for SeaArt by splitting them into multiple shots (e.g., 'shot one, shot two, shot three'). Describe each character's actions using the '@' symbol, and include a separate audio section at the bottom for sounds like breathing, footsteps, and katana clashes.
- **Editing in CapCut** [11:45] — The final step is to stitch the generated clips together in CapCut. Drag all scenes onto the timeline, trim or add transitions as needed, and export the final video.

### Conclusion

The workflow successfully creates AI videos with a consistent, realistic face across all shots, a feat that previous AI models have failed to achieve. The combination of GPT Image 2 for photorealism and SeaArt 2 for animation, along with careful prompt structuring, yields professional-looking results.

## Transcript

but making videos that really look like you is hard. That's why today I'm going used to put myself into a high action consistency. From building one consistent version of yourself to
generating the cinematic scenes all the way to the final edit. So by the end of this video, you'll know exactly how to make AI videos with your own face that stay locked in across every shot. So the first step in the entire workflow is
like you. For this, I'm going to be using OpenArt as they just added the new GPT image model. And on top of that, I also get access to SeaArt's 2. So this means the entire workflow is going to take place inside the same platform. So
once we're inside, let's head over to the character section and click create the first option called start from an image. And here you want to include as many images of yourself as possible. So what I did was take a few photos of my
thing you need to make sure of is that your background is clean and the lighting is neutral. So the AI can pick up all the details from your face. And you should also avoid wearing glasses or any other facial accessories you don't
want your character to have inside the video. So for this example, I'm going to upload three different images of myself and name the character Yuri. Next, you leave it blank for now. So now I'll click create and wait a few seconds.
create the first image. I'll start by generating a reference image for the video with my character holding a katana in full samurai gear. And once that's animate. For this, let's head over to the image section. And the first thing
Like I mentioned at the start, I'm going with the new GPT image 2 because right now is the best option for photorealistic images. And honestly, if you've used Nano Banana Pro before, this model feels way better. So once that's
selected, go to the reference section and click on characters. So we can find And as you can see, it now appears right now knows exactly which character we want to use and it'll pull that
So, here's the prompt I'm going to use there's one thing you need to be really careful with here. Make sure you use the at symbol before the character's name inside the prompt. This tells the model
exactly who to include and it anchors all the facial details to that specific guessing what you meant and the moment consistency we just built. For the ratio, I'm going with 16 by 9 because I
want this to feel like a cinematic scene. Then, for the quality, this model gives you three options: low, medium, and high. If you just want to test ideas for the cheapest price, you can go with low. Medium gives you a solid balance
between speed and quality. But since I want the best possible reference image for the number of images, you can generate up to four at once, but I'm going to keep it at one for now. And now let's click generate. Here's what we get
surprising. If we zoom in, you can clearly see that it looks exactly like intact while changing the entire environment and also the outfit. So, at this point, I'm going to save it and move on to the second reference image.
to create the second character. And for this one, I'm going to take a different approach. GPT image 2 is so strong when it comes to photorealism that you can actually skip the separate character building step and still get a great
the trick I've been using for this is to add the word photorealistic directly inside the prompt. It sounds like a really small change, but the difference alone pushes the model much harder toward realism and the output starts
looking like it came from a real camera instead of an AI generator. And the funny part is that I actually found this by accident. I was testing different prompts side by side when I started to notice that the only thing separating a
was just this word. But let's now describe the second character in full detail directly inside the prompt and let GPT image build it from scratch. And here's what we get back. It's honestly crazy how clean this looks, especially
without any character reference at all. If we zoom into the face, even the skin texture feels like a real photo. So, now both characters are ready, but the video still needs an actual location for our action. So, let's go ahead and build
style of this location has to match the two characters we just created. Otherwise, the whole scene starts to feel disconnected, like the characters belong in the same world. But here's the good part. Instead of writing the whole
prompt from scratch and trying to guess which words will keep everything do this. And for that, we're actually going to step outside of OpenArt for a second and head over to Claude. Once we're inside, the first thing I do is
upload the two reference images I just generated. The first one is my image, tell Claude exactly what kind of location I want, which in this case is a traditional Japanese dojo. But the most
important part is this. I also tell Claude to keep the exact same visual uploaded. And here's what Claude gives me back. It reads all the visual cues from those two reference images and turns them into a full location prompt
that already matches the cinematic Japanese style we've built so far. So, lighting, the color palette, and the mood myself, Claude pulls all of that uploaded. And it gives me a prompt that already knows exactly what tone the
location should hit. So, now I'm going to copy that prompt, head back into section. This time, I'm removing the character reference from before, because for the location, we don't need any character attached to it. For the ratio,
I'm still going with 16 by 9 because I want it to match the cinematic frame of quality, I'm going with high again, because this is the image every video now let's click generate. And here's what we get back. The polished floor
reflects a warm sunlight coming through the shoji screens on one side of the room, which instantly gives the space the cinematic depth we need for our even see the dust particles floating through the light. And that's exactly
brain into feeling like this was a real photo instead of something generated by AI. And the last thing I want to point out is the composition. Because instead of filling the frame with random objects, the dojo stays very minimal.
gives it that clean, professional, cinematic feel you usually see in real action films. So with these reference images ready, let's bring them into seed ants and animate the entire scene. For this, we need to head over to the video
reference. This is where we choose the AI model. Open art gives us a few really But for this kind of action scene, I'm going with seed ants 2. Now the first references. So I'm dropping in the two character images we just built, Yuri and
the gunman, along with the dojo location image. And what this means is that every detail we locked in across those three references now gets carried directly prompt I'm using for the first scene.
generations, I'm using the at symbol inside the prompt to reference each character and the location. In this way, seed ants knows exactly who goes where and who is supposed to do what. Now for the duration, I'm going with 15 seconds,
generate in one go. Then for the resolution, I'm setting it to 1080p. And for the aspect ratio, I'm staying with 16 by 9 so it matches the frame of every reference image we've created so far. So now let's click generate. And here's
now let's click generate. And here's what we get back.
I expected. Even though the door slides open by itself at the beginning, which is something seed ants 2 still struggles with. But, it's still the exact same warm light coming through the shoji screens makes the whole scene feel
cinematic. But, what really stands out here is the close-up on the face, how powerful the combination of GPT image 2 and SeaDance actually is. Just character's face is to mine. It looks amazing on the zoom. At that point, it
basically looks like a 4K photo of my face, where you can even make out every single strand of hair in the eyebrows. So, let's now move on and generate the is a little bit different, as we're adding one more thing. And that's a
video reference of the first scene we just generated. SeaDance uses it to last one ended. So, this means that the second scene doesn't start from zero. It scene one. So, the movement, the character positions, and the overall
mood all stay continuous. Besides the video reference, I'm keeping both the character and the location image. So, now I can move on to the prompt section. something really important. And that's how to structure a prompt effectively
for SeaDance, so you get the best results from the first try. If you just throw everything into one huge paragraph, the model gets confused about what should happen, and the output is going to be quite bad. So, here's the
structure I use for every single one of these prompts. The first thing I do is split the prompt into multiple shots. For this one, I'm using three shots, and I label them shot one, shot two, and shot three inside the prompt. This tells
SeaDance exactly when to switch to a new angle inside the scene. And it's one of without it, the AI has to guess them itself. Then inside each shot, I describe what each character is doing using the at symbol. And if I want a
close-up or a slow-motion moment, I place it right there at the start. Then the last piece is the audio. I always give SeaDance a completely separate audio section at the bottom of the prompt, where I I every sound I want
inside the scene. Things like breathing, footsteps, and even the katana clashes. ants treats the audio separately from the visuals. So, by doing this trick, the sound actually matches what's happening on screen instead of feeling
randomly thrown on top. So, here's the full prompt I'm using for scene two. The settings stay exactly the same as before, so I'll click generate. And before, so I'll click generate. And here's what we get back.
the multi-shot. The cuts between each shot are fast and professionally timed, more dynamic and intense. But, what really stands out is the emotional range in the face during the fight. Just look at my character at the start. He seems
to be so calm. Then, the moment the fight begins, his expression shifts into something much more serious. And once the katanas start clashing, he actually becomes furious. And the crazy part is that every single one of those facial
details stays locked in through the whole sequence, which is something no AI to pull off when trying to put me inside a video. So, let's use the same process for the third scene. For this, I'm adding scene two as the video reference
and then keeping everything as it is. Here is the full prompt I'm using for this clip. And again, I've structured it the same way I did for scene two, with section at the bottom. So, let's see what we get back after we generate the
what we get back after we generate the video.
lot. The thing that stands out the most is the slow-motion moment. Just look at how focused my character is on finishing the enemy. This whole scene feels like action movie. But, what really ties everything together is the lighting. The
screens gives the shot real depth and shadows, making the entire space feel cinematic. Each scene looks amazing on its own, but right now they're three separate clips, and we need a way to stitch them all together into one video.
CapCut. So let's jump inside and drag all three scenes straight onto the trim anything you don't like or add quick transitions between the clips. But in my case, everything already flows pretty well from scene to scene, so I'll
just hit export. Now let's see the final result.
face stays exactly like mine from the first frame to the last, and on top of smooth. You can see the weight of each strike and the way the enemy reacts, pull off. Honestly, this is something I've been waiting a long time to see.
Every AI model I've tested before has failed at getting these things right, but this is the first workflow that's giving me results I'm actually proud of. videos that actually look like you, sign up to OpenArt using the link in the
description below. Thanks for watching, and see you in the next one.
