---
title: 'How to Start Making AI Videos in 2026 (Beginner to Advanced)'
source: 'https://youtube.com/watch?v=0GO3-JzQjzg'
video_id: '0GO3-JzQjzg'
date: 2026-08-13
duration_sec: 771
---

# How to Start Making AI Videos in 2026 (Beginner to Advanced)

> Source: [How to Start Making AI Videos in 2026 (Beginner to Advanced)](https://youtube.com/watch?v=0GO3-JzQjzg)

## Summary

This video provides a comprehensive guide to AI video creation, from basic prompts to advanced cinematic control. It demonstrates a four-stage workflow using Higgsfield and other tools, emphasizing that the process, not the model, is the key to quality.

### Key Points

- **The core premise** [00:01] — The video introduces the idea that the difference in AI video quality comes from the workflow, not the model. It uses Higgsfield as an all-in-one platform.
- **Entry-level prompt** [01:05] — A simple prompt for a motorcycle rider racing across a desert canyon is used as the baseline. The model generates a video that technically matches the prompt but lacks creative control.
- **Chatbot-enhanced prompt** [02:34] — Using Claude to write a detailed, multi-shot prompt improves the output significantly, with deliberate camera movements and a more professional feel.
- **Storyboarding** [04:28] — A storyboard is created using GPT Image 2 to map out shots visually. This ensures consistency and reduces the need for regenerations.
- **Editing and generating with storyboard** [05:06] — Cinema Studio is used to edit the storyboard, fixing details like the road surface, and then to generate the video with the storyboard as a reference.
- **Full direction** [07:18] — The final stage involves creating a reusable character sheet from a collage of photos and a consistent location, then animating shot by shot with emotion settings.
- **Editing and final result** [11:12] — The two shots are combined in CapCut, resulting in a consistent, cinematic video. The video concludes by mentioning the upcoming Seedance 2.5 model.

## Transcript

now, and the biggest thing I've learned is that to make a cinematic output, most think it's about writing the perfect prompt or finding the best model, but that's barely where the difference comes from. So, in this video, I'll walk you
through the complete [music] process for making a simple video all the way to a cinematic shot, so you can finally master AI video making. We're starting at the entry level, which is what most [music] people do, to show you exactly
how much control you hand over to the AI without even realizing it. For this Higgsfield, which is an all-in-one platform that keeps every tool we need in one place. So, if you want to follow along, I've left a link for it in the
description [music] below. Once you log in, you'll notice everything's organized at the top navigation bar. And to make a video, you just click on video, which generation workspace. From the model
selector, I'll pick Seedance 2.0, because right now, it's one of the best cinematic outputs. So, the difference in the tools and not to the model. The whole idea at this stage is that you
don't do anything technical, you just type out what you want in plain words. So, I'll set the duration to 10 seconds, the resolution to 1080p, and the aspect ratio to 16x9, then write one simple prompt for a motorcycle rider racing
across a desert canyon road. That's the idea I'll use for all four stages, so we can really test the difference in outputs based on the workflow we'll use. The settings, the model, and the core idea will all stay the same.
good, and technically, it's doing exactly what I asked for. The bike is kicking up exactly [music] like I described, and it even added its own engine sound and some camera movement. If you want to experiment with an idea
or need a result fast, this will get your job done. But every choice in that video, including the camera, the sound, [music] and how it looks in general, all came down to the model. None of it could capture my exact idea because I didn't
give Seederns enough context to cover everything. At this stage, that's these simple prompts, pick the strongest one, you've got a real video in almost no time, which is exactly what you want when you're just testing an idea. But
the moment you want something specific instead of whatever the model gives you, writing a simple prompt isn't enough. You need something much stronger. But yourself can be very difficult, and it's easy to leave out important details.
chatbot to do it for you in just a few minutes. So, I'll open up Claude and get it to write the prompt for me. I'll ask for a more detailed, polished video prompt for that same motorcycle rider racing across the desert canyon, this
time on a timed rally stage. And what it hands back is on a whole different level from my one line. It's a way more AI-friendly prompt with a full three-shot sequence. Seederns 2 is built for multi-shot generation from a single
prompt, so this works great. It includes a drone shot high above the canyon, the camera following the rider from the side, and even a bit of dialogue. It like this gives Seederns actual directions to follow and real camera
movements, so there's way less for it to guess at on its own. Then, I'll head back to the video generation workspace with the same settings as before and paste in the prompt from Claude. &gt;&gt; [music]
video. The shots cut the way Claude laid them out, and the camera moves actually feel deliberate. So, the whole thing already looks a lot closer to a real nothing to do with the model. It's the exact same Seederns 2.0 on the exact
was giving it a stronger prompt to work dialogue line I asked for never got generated, but it's far closer to what I had in mind compared to the first video. So, from now on, a detailed prompt is a
necessary step. You can do it using any chatbot you prefer on the free plan, and far better result than you would by writing it on your own. But, a prompt cannot actually show the model any visuals. So, without any references,
this is as far as you can get. But, there's actually a way to plan out your entire video before animating a single frame. To do this, we need to make a storyboard, which is basically a single image that maps out every shot of the
video in one go. So, the model can see the whole thing laid out instead of relying on just a prompt. To make it, click on image and pick GPT image 2 as the model. For a storyboard, we basically want multiple images in one.
While they all maintain maximum realism, and GPT image 2 is currently the best model for holding all of that together. So, a prompt for a three-panel storyboard entire scene: a drone shot high above the canyon, a close-up of the
rider leaning over the bars, and a rear shot with the dust kicking up behind him. And the panels come back looking great. Consistency across all three realistic. But, there is one thing I don't like. Currently, the road came out
as smooth paved asphalt with lane markings, which doesn't really match the regenerating the whole storyboard [music] and leaving it up to chance thing. So, I'll switch over to Cinema Studio from the top navigation bar. Once
you're in, make sure you're on image mode, drop the storyboard in as a reference, and tell it to fix only panels one and three, turning the road into an unpaved dirt one, and keeping everything else exactly as it is. This
time, the dirt road matches the scene while maintaining everything else exactly the same. This is the storyboard [music] we're going to use for our Claude for a video prompt, but this time, I'll also upload the storyboard,
so it has the exact context of how the scene is supposed to look like. I'll also point out the duration of the video I want to make so we can plan everything with a detailed prompt that goes through three different shots that match our
storyboard. Then, still inside Cinema Studio, I'll switch over to video mode. here, it makes you run a quick eligibility check. So, go to image generation and check the storyboard, which is just HeyGen Field confirming
it. From there, I'll set the storyboard as my reference, the genre to action, the duration to 15 seconds, the resolution to 1080p, and the aspect ratio to 16 by 9, and turn the audio on. As you can see, Cinema Studio gives us a
lot more to work with compared to plain video generation, and that control is what takes the final result from a cool video to something that feels directed. So, I'll paste in the prompt from Claude and generate.
planned it, with all the elements staying consistent all the whole way through. And the reason this matters so much is to avoid burning through your credits on bad results. Editing an existing image is a lot cheaper than
redoing an entire video after you've animated it. So, not only do you ensure that your output will be exactly like you imagined, but you end up saving money and time. So, by planning everything ahead, the video comes out
right on the very first generation. Everything we've done up till now are the stages that take you from beginner to pro. But to actually master AI video making, you need to learn to control every single element in your video, just
stage is where you stop letting the model pick anything on its own and take control of everything. The plan here is to make a full 30-second cinematic split into two 15-second shots. And to keep it all consistent, I'll build my own
reusable character and location. I'll start back in Cinema Studio 3.5 inside image mode, set the model to auto, and this time, instead of a random rider, I'll make a character sheet of myself using a collage of my own pictures. A
collage like this works way better than a single image because the model will be able to understand what I look like from every single angle and keep my facial identity intact. So, I'll prompt for a character sheet of myself with a full
shot of me next to the bike and a close-up of my face, all on a neutral confuse the model later. We get back this image, and my face looks exactly the same. It changed my outfit to match the scene we're making, and with this
looks like. If you want to make any changes, you can edit it the same way we fixed the storyboard, but I'm happy with this one, and from now on, I can drop that exact version of me into any shot
and stay completely consistent. Then, I'll switch the model over to cinematic built specifically for realistic with a dirt road. This is our environment, and it looks pretty good.
The previous stages generated the canyon from scratch, but this time, we know exactly how it's going to look. And though it might seem like a small step, else, this process starts to feel a lot more like directing a video than just
generating one. Now, I've got both a character and the setting, so I'll head back to the image generation workspace with GPT Image 2 and build a six-panel storyboard this time because I'm planning a full 30 seconds as two
separate shots with three panels each. And because both shots pull from that same saved character, location, and storyboard, they'll stay consistent with one off the other. This is our storyboard. All our elements have
remained consistent. So, with the whole thing mapped out, I'll animate it one shot at a time. For the first shot, I'll ask Claude for a video prompt based on the tense opening, the drone shot, the
jump over the gap. Then, in Cinema Studio, I'll switch to video mode and run the same eligibility check for my assets. And instead of just pasting in the prompt, I'll use the settings Cinema Studio provides. There's no other tool
right now that lets you control a video this much, and every single setting affects your final output. I'll keep the genre on action, 15 seconds, 1080p, 16 by 9 with the audio on. This time, I'll also click on the smile mark and set the
emotion to vigilance, so my face matches the tension of the scene. Then, I'll the tension of the scene. Then, I'll paste in the cloud prompt and generate.
with my face holding the whole way through, and that tense, vigilant emotion has carried through to the final video. For the second shot, I'll go back to cloud and ask for a prompt for the bottom row. And then, back in Cinema
Studio, I'll keep most of the settings the same, but I'll switch the emotion to serenity, since the scene is coming to an end, and generate.
&gt;&gt; And again, each scene comes out exactly as planned. Every single element, the bike, the location, and even myself from behind, look consistent, and the dialogue came through this time. Then, I'll take both finished shots into a
video editor. I'm using CapCut for this, but you can use any editor you prefer. I'll just arrange my videos in order, and they should feel as one continuous built with the same elements. Let's look
built with the same elements. Let's look at our final result.
&gt;&gt; Compare this to the first video we made and the difference should be clear. And the whole reason for all this control isn't just quality, it's efficiency. Because saving your character and location and planning every shot is what
lets you get exactly what you pictured in your own style without burning worth mentioning that Higgsfield is about to release SeeDrones 2.5, which is set to be the best model on the platform yet. So, everything you just saw will
only get better from here. So, the same idea that started as one plain sentence same character in every shot. And that's the whole path from your first AI video on a real production set. Once you master each one of these stages, you'll
be left with the most efficient workflow that produces high-quality videos your credits. So, if you want to get started and make your own cinematic AI videos today, use the link in the description to sign up to Higgsfield.
Thanks for watching and I'll see you in the next one.
