Why Your AI Videos Feel Disconnected
45sTaps into a common pain point for AI creators, promising a solution that hooks viewers immediately.
โถ Play Clip"The title promises a professional course, and the video delivers a solid, actionable workflow with clear steps and examples, though it's more of a tutorial than an 'advanced course'."
This video presents a professional workflow for creating cinematic AI-generated films, emphasizing the importance of planning and consistency. The host demonstrates a step-by-step process using AI tools like Claude, OpenArt, and Seedance 2, from story development to final editing, to produce a cohesive short film.
Most AI videos feel disconnected because they lack a plan. The professional workflow involves planning every shot before generating, ensuring each shot has a purpose and contributes to a coherent film.
The host uses Claude to shape a rough idea into a structured story. For example, a 90-second ski film with a twist: the mission is actually a birthday surprise. This approach keeps the story simple and manageable for AI generation.
To maintain consistency, the host creates reference images for the character and locations using OpenArt's image generation (GPT Image 2). This ensures the character looks the same in every shot and the environments remain consistent.
The host uses Seedance 2, built into OpenArt, for video generation. Key techniques include using reference images, keeping dialogue as audio-only, and describing multiple camera angles in a single prompt to create smooth transitions.
To avoid hard cuts, the host uses the previous video as a reference for the next shot, allowing the action to flow continuously. This is demonstrated when transitioning from the descent to the ski park.
The final shot brings in the crowd, but they are kept in a wide shot to avoid rendering detailed faces. The host uses the empty base reference and adds the crowd in the prompt, ensuring the model doesn't struggle with close-up details.
The final step involves importing the four videos into CapCut, arranging them in order, and adding background music. Because of the consistent references, minimal cutting is needed, and the music helps unify the film.
The video demonstrates that with careful planning, consistent asset creation, and strategic use of AI tools, anyone can create professional-looking AI films that feel like a single cohesive piece rather than a collection of disjointed shots.
What is the first step in the professional AI film workflow?
Locking down the story before generating any shots.
00:44
Why is it important to create reference images for characters and locations?
To ensure consistency across all scenes, so the character looks the same and environments remain coherent.
02:51
What is the benefit of using a character sheet split into two panels?
It gives the model more angles of the character to work from, ensuring the character looks the same from every perspective.
03:32
How can you avoid hard cuts between shots?
Use the previous video as a reference for the next shot, so the action flows continuously.
09:38
Why are ski tricks and crowds kept in wide shots?
Because they are hard for AI to render up close; wide shots make them small elements in the distance, so the model doesn't have to nail details.
10:17
What tool is used for video generation in this workflow?
Seedance 2, built into OpenArt.
06:48
Story First
Emphasizes that planning the story is the foundation for a cohesive film, a key principle often overlooked.
00:44Consistency Through References
Explains how reference images solve the consistency problem in AI video generation.
02:51Seamless Transitions
Using previous videos as references is a clever technique to avoid jarring cuts.
09:38Wide Shots for Complex Elements
A practical tip to avoid AI failures by keeping difficult elements in wide shots.
10:17[00:01] videos still come out looking like a bunch of separate shots instead of one actual film? >> 30 seconds. Go.
[00:31] specific workflow pros follow where they plan out everything before they generate a single shot. And in this video, I'll walk you through that entire workflow from the first idea to the final edit so you can make cinematic films that hold
[00:44] together as one piece instead of a handful of shots stitched together. The first step is the one most people skip and it's locking our story. Without it, most AI videos end up feeling disconnected. The shots all get made
[00:56] with no plan behind them so they never add up into one coherent film. When you plan the whole story up front, every shot has a reason to be there and your consistent because you already know
[01:08] start generating. So, I'll open up Claude since shaping a rough idea into a I'll tell it what I'm after. A short film about 90 seconds with me as the
[01:20] main character and I want something with a bit of action and a simple twist that along with the whole workflow down in the description completely for free so you can follow along in real time. Claude comes up with a ski film shot
[01:33] like a high-stakes mission with me at the top of a mountain and the twist is a warm one. It was never really a mission. My friends are all waiting at the base to surprise me for my birthday. The reason this idea actually works well for
[01:45] AI creation is that all the hard stuff to generate, like the ski tricks and the big crowd of people, stays in wide shots where it's just a small element in the distance. So, the model never has to nail a spinning trick or a dozen faces
[01:57] up close. So, I tell Claude I like it and it tightens everything into a clean step-by-step process with every shot in order from the summit to the surprise. This setup also gives Claude enough context, so it can be your creative
[02:10] need help with the technical stuff, instead of spending hours doing them yourself and risk getting bad results, you just ask Claude to do it for you in is planned out before I've generated a single frame and that is the whole
[02:24] reason it holds together as one cinematic piece later. Locking down your story is the first important step, but for a professional output, you need to control every single element. That's why instead of just heading straight to
[02:36] video generation, we need to create our assets. The tool I'm using for all of all-in-one platform that gives me everything I need to make a whole film along, I've left a link to it in the description below. Now, the whole reason
[02:51] to consistency. With a plain text prompt, you can't actually control what your characters look like or how your locations will be and if I were to just they'd come out looking a little different every single time. So,
[03:05] instead, I'll build my character and my locations once as reference images and which is what keeps everything consistent across all scenes. The first which is based on me. So, I'll get Claude to write me a prompt for all my
[03:20] assets and the first one is a proper character sheet, which is one image split into two panels. A full body shot of me on one side and a close-up of my face on the other, dressed in the Mission ski gear, all on a plain snowy
[03:32] background. The reason I want it split like that instead of just a single photo is that it gives the model more angles of me to work from, so I'll stay looking like the exact same person from every perspective. And the background stays
[03:44] plain on purpose, so when I give this into the video model later, it isn't thrown off by anything else in in frame. Once I've got the prompt, I'll log into OpenArt and click on image from the sidebar to enter the image generation
[03:56] workspace. Then, I'll set the model to GPT Image 2, which is one of the best image models for a photoreal result that actually holds the exact face and details you give it. Then, I'll upload a collage of myself so it knows what I
[04:09] better than a single photo because the model can see me from multiple angles with a lot more detail, which is exactly what will make this seem realistic. Then, I'll set the aspect ratio to 16 by 9 to match the film, push the quality up
[04:23] high, and generate. I always match the quality on these because the cleaner the the model builds off it later. And it comes back looking exactly like me in the full mission ski gear with the face matching my photo spot on. That single
[04:37] which will ensure that I'll look exactly the same in the entire film. With me locked in, I need the actual locations the film takes place in, starting with the mountain I ski down. This one works a little differently though because it's
[04:50] there's no face to match and nothing to upload. It's a pure text prompt. So, I'll take the prompt Claude has made for me, which describes the top of a snowy mountain summit with a few pine trees scattered down the slopes. And in the
[05:03] same image workspace with GPT Image 2, I'll remove the reference photo and with the same settings as before, I'll paste in the prompt and generate. And it comes back looking exactly like I described. The realism is great and GPT Image 2
[05:16] made this seem like a picture of a real place. This is the location for my first two shots. So, by building it once, it will keep the environment consistent and it will ensure that the transition between videos is smooth and natural.
[05:28] thing ends up, the base at the bottom of the slope. Similar to the mountain, it's from Claude and it describes the base of a ski slope next to a wooden lodge. But it specifically points out that there should be no people in the frame, which
[05:43] is very important. This is where the birthday crowd shows up at the end. But image now, I'm stuck with whoever the model puts there, and it will be very easy for everything to break when we start animating. So, I'll leave the base
[05:56] totally empty and add the crowd later in the actual video prompt for that final shot, which means I stay in full control of who's in the frame and when. I'll generate it the exact same way, and it comes back as a realistic lodge with
[06:08] fresh snow and not a single person around. Now, before I build anything on these, I always do one quick check and look all three over because every shot from here is built on them. So, if my character or a location came out wrong,
[06:20] it would show up in every single video after it. This time, all three were spot on. But, if one had drifted, now's the moment to regenerate it, not after you've already built three shots on top of it. So, now the story is locked, the
[06:32] elements, and I've got Claude as a personal assistant writing every prompt when it comes to making the actual videos, most people still struggle, and separates someone that randomly generates from a professional who's
[06:48] directing a film. And that difference comes down to how you actually generate how I do it. The tool for the actual videos is Seedance 2, and it's built right into OpenArt. So, I'll head back to the navigation bar and click on
[07:01] video, this time to open the video workspace. From the model selector, I'll pick Seedance 2, which right now is one of the best video models out there for cinematic quality, and it's built to handle the kind of detailed prompts
[07:13] the images, I'll have Claude write the actual video prompt for me. And the first shot is the mission kicking off up at the summit. One thing I'm doing on purpose here is keeping all the dialogue as audio only, so that walkie-talkie
[07:26] voice is just sound. There's nobody else on screen I have to animate, which is wrong. To actually build the shot, I'll pick the text with reference option, since this one uses my references, and I'll upload two of them, my character
[07:39] sheet as image one, so it's me in the shot, and the mountain summit as image two, so it's the right location. Then, I'll set the aspect ratio to 16x9, the resolution to 1080p, the duration to 15 seconds, and leave the audio on, so it
[07:53] generates the sound along with the video, and generate with the prompt Claude gave me. >> 30 seconds. Go.
[08:15] holding perfectly from the reference, and the walkie-talkie audio coming through right on cue. And straight away, it feels less like a random AI video, and more like the opening of an actual film. The realistic physics also stand
[08:28] out. SeeDrones 2 is currently one of the best models for that, and it shows. With the opening done, the next shot is the descent down the mountain. For this one, I'll keep the same two references to make the transition between shots flow
[08:40] naturally. The prompt this time is a whole sequence packed into one shot. A high aerial drone following me as I go down the slope, then a fast tracking shot as I weave through the pine trees, a close-up on my focused face, and a low
[08:52] angle as I race past the camera. And this is something a lot of people miss. You don't need a separate generation for every camera angle. You can describe a and SeeDrones will move the camera through all of them in one take. It's
[09:05] the same setup as before, so I'll just paste the prompt in and generate.
[09:25] The drone shot and the sheer speed are selling the whole thing, and I'm still clearly the same person from the first shot, which is exactly why we built our assets first. The The shot is the ski park, and this is where I do something
[09:38] different to keep everything flowing together. Normally, if I just generated a fresh ski park video, it would look like a hard cut, a brand new shot that off. So, instead of using the mountain image again, I'll upload the previous
[09:51] video itself as a reference, keep myself as image one, and skip the location image completely because the park gets described right in the prompt. What that does is tell Seederns to pick up from the last frame of the descent. So, the
[10:03] ski park continues straight out of it as one flowing motion with no visible cut specifies me hitting a rail and then launching off a kicker with a spin. And the important part is that all the tricks are framed as wide shots and
[10:17] drone shots on purpose. A fast spinning ski trick is one of the hardest things for AI to get right up close. So, by keeping it wide where the action mostly reads as a silhouette, the model doesn't have to nail every tiny detail and it
[10:29] comes out clean. I load it into Seederns the same way with my character reference and the previous video attached. Keep the settings the same and generate.
[10:51] perfectly from the descent. The rail slide and the spin both come out as one smooth continuous shot. And because I kept the tricks wide, none of it comes out looking unnatural. The last shot is the one the whole film has been leading
[11:03] time, I'm back to using image references, me as image one and that empty base as image two. And now is when I finally bring in the crowd right here in the prompt. Cloud writes it as a wide shot of me skiing down to the base and
[11:17] with balloons and a party banner. And just like the ski tricks, that whole crowd stays in a wide shot in the background because a dozen faces up wrong. So, keeping them back and letting the cheering carry as audio means the
[11:31] model never has to render a load of detailed faces. So, I'll paste in the detailed faces. So, I'll paste in the prompt and generate.
[11:45] Oh my god. I thought this was something urgent. the fake mission paying off as a warm surprise and it honestly feels like the
[11:58] final scene of a real short film. The banner is missing, but honestly, it doesn't affect the final result. You can clearly tell it's a birthday party and four shots done and instead of having multiple disconnected videos, they
[12:11] actually play as one continuous story from the summit all the way to the surprise. But right now, these are still four separate videos and turning them into one finished film takes a couple more steps. So, we need to bring
[12:23] CapCut, but you can use anything you prefer. I'll drop them onto the timeline in order, one through four. And because every shot was built off the same references and the ski park picks up straight from the descent, they already
[12:36] there's barely any cutting to do. From there, I'll just add some background turns warm for the surprise and that music is a big part of what makes the whole thing feel like one film instead of a bunch of separate shots. Then, I'll
[12:50] export and it's done. >> 30 seconds. Go.
[13:43] oh my god. I thought this is something urgent. continuous cinematic piece that goes from a tense mountain mission to a warm
[13:55] birthday surprise, built from nothing but a few images and a planned out story. The control from this workflow will take you from just starting out all every scene. So, if you want to start
[14:07] making your own professional AI videos today, use the link in the description watching and I'll see you in the next one.
โก Saved you 0h 14m reading this? Transcribe any YouTube video for free โ no signup needed.