---
title: 'How to Make a Cartoon with AI | AI for Video Creation'
source: 'https://youtube.com/watch?v=IS14tVq9D0M'
video_id: 'IS14tVq9D0M'
date: 2026-08-04
duration_sec: 920
channel: 'Денис Марков'
---

# How to Make a Cartoon with AI | AI for Video Creation

> Source: [How to Make a Cartoon with AI | AI for Video Creation](https://youtube.com/watch?v=IS14tVq9D0M)

## Summary

This video provides a step-by-step guide on creating animated videos (cartoons) using neural networks. The creator demonstrates how to identify trending video formats in a niche, generate character images, animate them into short clips, and assemble them into a final video with subtitles and voiceover, all using AI tools.

### Key Points

- **Overview of AI video creation** [00:02] — The creator explains that they use two to three neural networks with several prompts to generate a stream of content, and this video will show the step-by-step process.
- **Finding trending video formats** [00:33] — Search for your niche on social media (e.g., Instagram) to find videos that have already gathered millions or hundreds of thousands of views. If no such videos exist, the niche may not be interesting.
- **Scraping competitor videos** [01:05] — Copy the nicknames of popular videos and load them into a scraper to get a table of the most popular videos from competitors, showing what content is already performing well.
- **Identifying trends** [01:47] — Examples of trends include TV series formats and characters like animated fruits or organs coming to life. These formats can be used to promote your own account.
- **Using competitor videos as scripts** [02:43] — Watch videos in your niche to get ideas for similar videos. You can reshoot with the same text and script, or create something similar if you see a popular format.
- **Three ways to create neurovideo** [03:41] — 1) Text to video, 2) Image to video (animating an image), 3) Video to video (replacing a character). The creator chooses image to video for more accurate generation and cost efficiency.
- **Creating character and scene images** [04:38] — Create two types of images: character sheets (characters separately) and footage (characters in different rooms/moods). This is necessary for more complex videos.
- **Finding prompts** [05:04] — Prompts can be found in open sources, such as the Telegram channel 'Traffic Chief'. The creator uses the Syntax service for image generation.
- **Generating series images** [06:33] — For a series, indicate characters in prompts and then tell the neural network what each character does in different scenes.
- **Using prompt packs** [07:17] — The Telegram channel provides a set of tools for video editing, including prompts for various characters and scenes, music, and other resources.
- **Creating custom prompts** [08:42] — If you like a generation but don't have a prompt, take a screenshot and use a text neural network (like GPT) to generate a prompt for that character.
- **Animating images to video** [10:50] — Use the video section in Syntax, select a neural network like VO3, set image-to-video mode, add the generated image and a prompt describing the action. Each video lasts up to 8 seconds.
- **Cost-effective alternative** [12:04] — For videos without speech, choose a cheaper model like Image 10 (5 credits) instead of VO3 (more expensive).
- **Editing and adding subtitles** [12:33] — Use a video editor like CapCut to glue episodes together, add subtitles for vertical videos, and add background music.
- **Adding voiceover** [12:49] — Use Eleven Labs or CapCut's text-to-audio to generate voiceovers. Eleven Labs offers more realistic voices but costs more.
- **Multi-accounting** [13:56] — Publish content across multiple accounts on Instagram, TikTok, etc., to increase traffic. Purchase accounts from FB1 store to avoid blocking.
- **General workflow** [14:38] — The skeleton is: create a script, generate images, animate them, and optionally add sound. Neural networks may change, but the process remains similar.

### Conclusion

The video provides a comprehensive, step-by-step guide to creating AI-generated animated videos, from trend research to final editing, emphasizing the importance of using trending formats and efficient tools.

## Transcript

posting videos.  I take two, maximum three neural networks, load several promts into them and get a stream of content.  In this video I will show you step by step how to create long and videos.  Just repeat after me and you will succeed.  Enjoy
is provided for informational and educational purposes only. actions shown in the video. We begin our journey with the question: what kind of videos should we shoot?  On what topic?  What kind of video is this and so on.  We go to any
social network, for me it’s Instagram, and enter the name of our niche in the search.  We entered it and the name of our niche in the search.  We entered it and received videos in the search results that have already collected millions or at least hundreds of thousands of views.  And every niche should have
such videos, because [music] if there are n’t any, then most likely the niche is simply not n’t any, then most likely the niche is simply not particularly interesting to people and it’s simply not worth being in. And in my case, these are neural networks.  I see
videos that have garnered hundreds of thousands and millions of views.  Then we copy their nicknames and load them into the scraper.  And we get a table like this.  These are the most popular videos from our competitors. And, accordingly, we see what content is
already flying into our topic.  Let's look at some.  For example, this one has collected 960,000 views.  We load it and see that this is a
views.  We load it and see that this is a TV series, which is currently a trend in niches, in the neural network niche.   Let's see again.  And this is the same see again.  And this is the same series, by the way.  Just the second episode.
Let's find something else. Here is the video collected.  Yes.  And these characters are also a popular trend.  This is exactly the format I used to promote my
exactly the format I used to promote my Instagram account.  Yes, it will improve your Instagram account.  Yes, it will improve your health.  And here, diseases or human organs at the intersection of health and world networks came to life.  Let's
world networks came to life.  Let's watch something else.  The same goes for the series.   Yes. watch something else.  The same goes for the series.   Yes.
that made it into this table.  And these videos are your scripts. Watch videos in your niche, and you'll have a ton of ideas for similar videos you can make have a ton of ideas for similar videos you can make .  Like, just reshoot it.  If
we see that a regular video has been uploaded, we can create a video with exactly the same text and the same script.  If we see Iiliki, then we can very well
create something similar.  I teach how to create tables like these, professionally analyze competitors, and work with data in my private club, Triffic Club.  This is a closed community of creators where I share connections,
developments, and market observations.  There you can find information on how to analyze competitors in order to catch trends as they emerge.  You can also find ready-made bundles there.  I have more than ten pieces on
bundles there.  I have more than ten pieces on goods, gambling, and so on.  Just take it and apply it.  There I can also answer your questions if you don’t understand anything.  I'll give you a hint.  In general, the link to Trffic Lab is in the description.
Move, study and move forward with traffic together - it's more effective than moving solo.  Let's move on to creating images that we will later bring to life.  Let's get this straight right away. There are three ways to create a neurovideo.
The first is from the text in the video.  That is, we give the text to the neural network, and it immediately creates a video based on it.  The second one is from the picture in the video.  That is, we animate
the image, or animate it, as you prefer.  And the third is from video to video.  That is, we need to replace the character in the video.  I will choose the second method from the images in the video, because this is the
most accurate generation.  And secondly, we can already tell from the pictures whether the neural seeding rules are relevant to the task or not, because in terms of credits, an image is much cheaper than a video.  And when we create an image, we need to
create two types of images.  The first is the character system.  That is, we create characters separately, which will then be used in our video.
And the second one is footage from the future video. That is, these characters are in different rooms, with different moods, and so on.  This is already more cool for such videos.
If we make a simpler video, then this is not necessary.  But let's still take it into account.  And to create an image, we need a prompt.  It can be found in open sources.  This is what it looks like.  By the way, I work for the
Syntax service.  You see, [music] my chats.  That is, I use this service, chats.  That is, I use this service, and you can check it out.  And here is my chat in the design section.  Here is my [music] prompt.  And this is the generation I
got. That is, it is a character system.  I created That is, it is a character system.  I created my character based on the Adalisia template. my character based on the Adalisia template. Well, here we see him.  Alize.
My second pomt is already like this.  That is, I add a picture of the character and in the prompt I indicate what he is doing.  Here we see the second generation.  Here my
character is already in a different pose, in a different situation.  That is, we only created it once situation.  That is, we only created it once , and then we just change the poses and everything else.  And, accordingly, here is another such prompt, that is, the
another such prompt, that is, the same character.  the room is different, the pose is different.  I indicated that the character is in such and such a pose.  And, in fact, we get this result.
Let me show you something more complicated.  This is already a generation of my and [music] series. already a generation of my and [music] series. And here, as it were, I also indicate the characters in the prompt, actually.  Here.
actually.  Here. And then, in other prompts, I tell the neural network who does what.  Here, for example, are strawberries in different poses. for example, are strawberries in different poses. M, these same characters are already in different
situations. And I just add their generations and tell the neural network where they are. Where can I get these prompts?  This can be Telegram channel, Traffic Chief.   Let's
go in.  and we find this set of tools for video editing. Next, go to this folder and select the prompts section.  And we, accordingly, see these generations and the prompts
they are made according to.  For example, you can make a generation like this with a character make a generation like this with a character from the movie Cars.  Well, uh, this one needs no introduction.  TV series with animated fruits, vegetables, berries, and so on.
Next, we copy this prompt and paste it into any neural network that generates an image.  I recommend nanoban. an image.  I recommend nanoban. You may have something else.
Further in this same folder you can find, for example, music for video editing. for example, music for video editing. And here she is.  And the neural networks that I use, other tools and files for editing.  For
tools and files for editing.  For example, backgrounds, emojis and everything else.  Overall, this is a must have if you edit your own videos.  Link to
my open Telegram channel, as well as to this pack, in the description.  Go ahead and explore.  So what if you liked the generation ?  But I didn't find any prom in my pack
.  There is a solution here too.  We are looking for our generation that we like. For example, I liked this.  It's a GPT chat room come to life, but it's not a Pixar character either GPT chat room come to life, but it's not a Pixar character either .  And then I take a screenshot of this
character, and then I go to any text [music] neural network. I'll do it again through syntax.  Ah, I'll choose the GPT chat GPT chat and, ah, well, let's choose
GPT 5.4 Pro, yes, five credits will be written off. And then we add our screenshot, which we just made, with the character for whom we want to
receive prompts, we add it.  Write a prompt for nanoban prompt for nanoban to create
image. Okay, let me change it to Okay, let me change it to chat GPT 5. So, Chat GPT web. chat GPT 5. So, Chat GPT web. Ah, Chat GPT5, yes, it will be free.
in English.  In English the prompt would be more accurate, but [music] now the prompt would be more accurate, but [music] now for clarity I will copy the Russian. Next we move on to design in syntax. Select a new chat.
Oh, and we add this promt.  Yes, I have, but bananas aren't much cheaper.  Nan banana Pro, the ratio
got the same character as in the original.  Yes, in the original it is more fierce, more aggressive, but we can work with this generation and make it like that on others.
Let's move on to creating the video. Accordingly, we move on to the video section.  Here it is.  And here in the settings we select.  There are a lot of neural networks here.  I we select.  There are a lot of neural networks here.  I choose VO3.  Almost all my videos are made
through it.  Here is the VO3 fast version.  The image-to-video mode is image-to-video mode is 9K, and the resolution is 720p. So, we add the generated image, add a
generated image, add a prompt, and see that VO3 brings it to life. In the prompt we indicate what this or that character does, that character does, and, accordingly, we get the result.
And each video lasts a maximum of 8 seconds. Here we see that [music] the characters are in different locations, and action is happening.  The video lasts up to 8
seconds, and therefore we cannot do everything in one generation.  [music] And here I can also choose the player.  When I don't need speech from the characters, I do it this way.  I choose,
do it this way.  I choose, and here is imagм 10, I choose six.  And here it's only five credits.  This is almost three times
cheaper than EO3.  And then we go to any video editor, I use CapCAD, any video editor, I use CapCAD, and, uh, we connect and glue together all of our resulting episodes.  Actually, here they are, I have them .
These are different frames from the same cartoon.  Next, we add subtitles if we're making vertical videos, and we add background music, because the neural network itself is n't yet capable of this.  And if we need a
voice-over, then go to the sound section.  And here we see, ah, two neural networks. We select 11 Labs and, accordingly, speech.  And here
we enter the text and, accordingly, it is converted into audio.  And here, accordingly, we select a voice.  We can upload our own or
choose from the existing ones. And the majority of votes here support the Russian language.  Or the same can be done in Kapkat.  Go to the done in Kapkat.  Go to the text section and select text to audio.
text section and select text to audio. Here we enter the text we need, and then select the voiceover. The voices here are simpler than in 11 Labs, but
The voices here are simpler than in 11 Labs, but we also pay less for it. The library is quite large.  We can try different ones and choose the most suitable one.  and we can publish our content.  By the way, we can do this with
multi-accounting.  That is, we have many accounts on Instagram, TikTok, and so on.  And we publish single-title content to increase traffic volumes. If you have multiple accounts, I recommend purchasing
Instagram and TikTok accounts from the FB1 store. Go to the website and select the appropriate section.  Ours is either Instagram or TikTok.  We select and see the products available to us.  We choose the most suitable one, purchase it and don’t have any
problems with [music] blocking and everything else.  Link to FB Store in the description.  Go ahead and explore.  I showed only one connection, how you can showed only one connection, how you can generate video.  However, the skeleton is
more or less the same everywhere.  First we create a script, then we generate images, then we animate them and optionally add sound separately. Neural networks for creating and editing images can also change.  And
this is worth keeping an eye on, because video is the only trend for 2026 so far.  By the way, regarding trends and getting into them.  I recently discussed five sources of conditionally free traffic, their pros and cons, how to get
recommended, and so on.  To do this, watch this video, and you will have no questions about which format to use and which one is most suitable for to use and which one is most suitable for your project.
