---
title: 'Generate AI Images Without Paying A Dime | Stable Diffusion Tutorial'
source: 'https://youtube.com/watch?v=CWoZbpgL0aw'
video_id: 'CWoZbpgL0aw'
date: 2026-07-28
duration_sec: 989
---

# Generate AI Images Without Paying A Dime | Stable Diffusion Tutorial

> Source: [Generate AI Images Without Paying A Dime | Stable Diffusion Tutorial](https://youtube.com/watch?v=CWoZbpgL0aw)

## Summary

Stable Diffusion is a free, open-source AI image generation platform that offers more control and customization than paid alternatives like Midjourney or DALL-E. This tutorial covers installation, basic usage, and advanced techniques like inpainting and image prompting.

### Key Points

- **Stable Diffusion is a hidden gem** [00:00] — Unlike paid platforms, Stable Diffusion is completely open-source and free, allowing anyone to create images with text prompts.
- **Local GPU usage** [00:49] — Stable Diffusion runs on your own GPU instead of the cloud, giving you full control over every setting.
- **What is Stable Diffusion?** [01:30] — It's an image generation model by Stability AI that turns text prompts into images. Being open-source means it's free and customizable.
- **Using Focus as a front-end** [01:53] — Focus is a user-friendly interface for Stable Diffusion, like a car built around the engine, providing buttons and controls.
- **Installation steps** [02:24] — Download from GitHub, extract the file, and run the .bat file. The download is about 50 GB. Default models include standard, anime, and realistic.
- **Generating images** [03:15] — The interface has an image generation window, text prompt input, and options like input image and pain enhance for post-processing.
- **Crafting effective prompts** [05:14] — Specific details like 'cyberpunk city at night with glowing neon lights and cinematic lighting' yield better results.
- **Image upscaling** [06:34] — Use image upscale to increase resolution by 1.5x or 2x, with a fast option for quicker processing.
- **Inpainting for refinement** [07:34] — Inpaint allows you to brush over areas to refine details. Options include inpaint (subtle variation), improve detail (increase resolution), and modify content (dramatic changes).
- **Advanced image prompting** [09:29] — Use image prompt with Pyrocheni (maps character positions) or CPDS (uses contrast/color) to generate images similar to a reference.
- **Negative prompts for safety** [10:06] — Add negative prompts like 'not safe for work' to avoid unwanted content. Default negatives include unrealistic, saturated, big nose, etc.
- **Iterative refinement with inpainting** [13:11] — Drag images into inpaint to fix specific issues like broken limbs or incorrect outfits, using references to anchor details.
- **Image expansion** [15:20] — Use content expansion to extend an image to landscape format, generating left and right sides with a prompt for background details.

### Conclusion

Stable Diffusion offers powerful, free image generation with extensive customization. With practice, you can create high-quality images and refine them using tools like inpainting and image prompting.

## Transcript

Generative AI has opened the door to
allowing anyone to create incredible
images just using a computer and a few
text prompts. Now, the platform that
gets the least amount of attention is
Stable Diffusion, but it might just be
the hidden gem you're looking for.
Unlike other major generative AI
platforms like Adobe's Firefly or
OpenAI's Dali or even Midjourney, Stable
Diffusion is completely open source,
which in short is free. So, I'd want
Padme to wear her iconic [music] white
turtle attire. See that image is coming
in. It's looking great. There's a few
issues going on in the lower parts of
the frame. In this particular section,
I'm changing how the body structure
looks. Again, let's increase that to
four images. Okay, we're going to pause
it right there. Remember, there are no
limitations in the software, which means
the responsibility is on you.
Stable diffusion lets you run on your
own GPU instead of on the cloud, which
allows you to tweak every minute setting
that you can think of. Now, if that's
not enough to convince you to learn this
software, it's also the only platform
with zero limitations, but it also comes
with its own drawbacks. It's tough to
get started with, but that's where we
come in. Today, we're going to help you
with everything from getting started,
installation to even more complicated
things like perfecting your prompts and
getting the best possible image. Now,
today's episode is a little bit
different. It's just you, me, and a
computer. And we're going to get through
this together. It's also the first time
that we haven't had someone behind the
camera, so expect things to go wrong.
First off, what is stable diffusion? In
short, it's an image generation model
created by stable AI. It turns text
prompts into images. What makes stable
diffusion really special is that it's
open source, which in short means that
it's free, but it also means that you
can download it, bring it into your own
computer, and customize it to your own
art style.
Now, in a majority of this video, we
will work with a software application
called Fucus. This is a front-end
interface that allows us to interact
with the stable diffusion model. Now,
don't let any of those words scare you
off. Think of stable diffusion as the
engine, the code that's doing all the AI
image processing. Focus is like the car
that's built around the engine. It's
what has all the buttons and the
controls and the interface that allows
you to interact with the engine without
really messing around with the tech.
Okay, enough talking. Let's jump into
the application. The first thing you
want to do is go to the GitHub
repository for Focus. We'll leave a link
in the description. After you download
the file, you want to extract the file
and you'll see and run.bat. You want to
click on the run.bat file and start the
installation process. Remember, the
download file will be around 50 GB. So,
make sure you have that capacity before
you get started. By default, you'll get
the standard model, but you'll also get
the option to run the animate and
realistic models as well. For right now,
we're not going to change any of the
default behavior of the software, which
means as it launches, it'll look for new
models, keep the software updated, which
is generally what you want. But if you
want to change that, you can use these
two command lines. Let's do realistic
right now. And now we just let it do its
thing. Another quick tip that I like to
use is keeping task manager open. As
long as I see activity on the GPU, I
know that the software is working with
the GPU to get it collected.
Awesome. And we are finally ready to
start generating images. You'll notice a
few things. The image generation window
at the top, the text prompt window at
the bottom where we can type things like
cats on a window lid. And at the bottom,
you can see input image. We'll go into
that in detail. It's something that I
use in great detail. Pain enhance, which
we won't talk about too much.
essentially post-processing steps that
you can take on your final image to
increase the resolution and detail. As
we hit generate, you can see that there
is an immediate spike in your GPU. So,
you'll notice the image comes in a
little fuzzy, then it cleans up over
time. Okay, with those images in place,
we can look at them in detail by
clicking on each individual image. The
second image is a really good example of
the kind of problems that this kind of
model has. You can see that the eye on
the left is perfect, but the eye on the
right just doesn't have the detail we
need. So, there's two things that we can
think about at this point. We can click
on advanced and we can first look at the
presets. We are using the realistic
model which is what generates such fine
detail but there's lots of other
variations that you can try out here.
There's also performance depending on if
you want to take a qualitative or
quantitative approach. Right now we're
prioritizing speed of delivery, but we
can also change that to quality to
generate higher quality images. Now a
feature that I use all the time is
changing the number of images, but for
right now let's do four images.
>> [music]
>> That's pretty incredible. The first
three images have come in. I love the
detail here. There's lots of detail in
the hair and in the eyes. Both eyes look
perfect. The second image is not quite
as good. It has really great
perspective. Lots of detail in the
bricks and the window, but not quite as
much detail on the cat itself. You can
see that the eye on the right isn't
perfect. Okay, this is great perspective
cuz there's a beautiful window. It's a
low angle shot. Again, not quite as much
detail on the cat itself. And the final
image has also come through. Oh, this is
amazing. And you can see the cat looking
through the window. Lots of detailing on
the glass and the light and the texture
on the skin. We're going to build a
crazy Star Wars poster no one has ever
seen of Padme fighting off Anakin
Skywalker. But first, let's build a
cyberpunk city with crazy [music]
detail.
So, you can notice the kind of detail
that I'm putting in here. I'm I'm
writing out very specifically what I
want to see. So, a cyberpunk city at
night. So, I'm specifying what the
lighting looks like. Glowing neon
lights. someone have that look where
there's lots of light from the buildings
themselves. And I'm specifically calling
for cinematic lighting. With AI models,
the more specific details you can give,
the more likely you are to get a
successful final image. Okay, so those
two images have come through. They're
both fairly similar, which is something
I don't really like. Ideally, I want
variation on shot. Also, something I
like to do is watch the model as the
image generates just to see if it's
something that's in alignment with what
I want. If it's completely off, I can
choose to skip through that image. Okay.
And that's pretty great. Everything's
generally looking okay, but nothing is
specifically looking good. That's also
another problem with the AI generation
model for wide and detailed shots. It's
harder to get fine details correct when
there's lots of details spread across
the image. Let's say we love this image
in general, but we want to enhance
certain details. With that in mind,
let's refine our model some more. Okay,
now we've got two new images. Let's look
at both of them. I love the detailing on
this. Generally, everything seems okay.
There's lots of lost details in the
buildings in the background, but I think
I can generally fix those as we go. The
second image is a lot softer and there
is a lot of billboards and written
detail which [music] I probably won't be
able to fix. What if we have a great
image, but we want to refine that image.
That's where image input comes in.
First, if you want to take this image
and we just want to scale it up, bring
in new resolution, we can use image
upscale. So, you just drag that image in
here. First, we look at the bottom row.
Upscaling by 1.5, upscaling by 2x. that
essentially just expands the resolution,
maintaining as much of the image as
possible. And then you have upscaling
fast 2x, which is the same thing, but it
processes it a little faster with a
little less accuracy. So, first let's do
a 1.5 upscale. Now, as that image comes
in, we can see we already have a lot
more detail. Okay, the first image is
in. Let's take a look at what that looks
like. So, you're already seeing so much
more detail in the building in the
foreground. All the problems of the
background have now been resolved. All
the lines are straight. All the windows
are visible. Generally, everything's
looking good. There is still a little
bit of specific problems I'm seeing,
especially on this billboard on the
right, parking lot down below, but the
cars don't look perfect. We can work on
those specifically. Okay, let's look at
this image with some detail. That's not
going to work cuz it's so close to the
foreground element, which is this
building. So, to refine this, we're
going to use impaint. This allows you to
refine and work on specific details of
your image with great detail. There's a
brush, and anything you brush over will
then be changed. Anything that's outside
the brush radius will not be affected.
So, we've got three options here.
Impaint, improve, and modify. Impaint is
a great technique if you want to change
something of your image with subtle
variation and keep the general look of
the image the same. The first thing
you'll notice is that the GPU is using
all of its processing power on just that
one zone of the image, which means
you'll get a lot of resolution in just
that one area. Has a little motel
looking building with a swimming pool in
front. It's lots of details in terms of
cars and people, which may not be what
we want because it's going to attract a
lot of attention. And then the second is
a black building with a few windows. You
can see that I can actually drag the
image from my image generation window
into my impaint window. This is a pretty
common technique of refining your image
as you move it back and forth within the
software itself. And this is a good time
to talk about the other two impend
features as well. Improve detail is used
very often. This is when you want to
increase the resolution of something in
the background, right? We can see far
more detail in that image. You can see
the specific floors of the building. You
can see through some of the windows. You
can see some trees and shbery brought in
front of the building. Really quickly,
let's look at modify content as well.
This is a powerful tool when you want to
make a dramatic shift of your image and
then later go back and refine it using
the refineer tool. So, let's look at one
of these. And that's what that car path
looks like. Again, here you can see a
lot of imperfections in the car, in the
floor, in the building next to the car
park. All of which will need to be
refined if you want to use it in your
final image. But for right now, we're
just going to revert back to the image
that we had and look through our final
settings. At this point, you probably
get the idea of image generation, but
let's dive into something a little bit
more obscure that uses more advanced
features of the software.
So, [music] let's think of something
that we can't find on the internet.
though. Anakin Skywalker and Padme
lightsaber dual Star Wars franchise high
detail cinematic lighting at night.
[music] All right, let's see what this
looks like. Again, let's increase that
to four images. Okay, we're going to
pause it right there. So, this is where
you need to be careful as you generate
images. You need to make sure that you
don't generate any images that's not
safe for work. And this is a good time
as any to talk about the power of these
creative tools. Remember, there are no
limitations in these software, which
means the responsibility is on you. Make
sure you don't create any content that's
harmful, misleading, or violates
privacy. Another tool that we can use to
safeguard our content is negative
prompts. Now, these are things that you
don't want in your image. So, by
default, I'm getting things like
unrealistic, saturated, big nose,
painting, drawing, sketch. These are all
prompts that have come in default by the
software. But now, I'm going to also
include not safe for work as a prompt.
And you can expand on that list. Okay.
And those images have come through. They
both have their own challenges. So over
here, both Anakin and Padme are sharing
a lightsaber, which doesn't really make
the most sense. In the second image, I
kind of have their hands crossed over.
Both which we can fix. And we can do
that by bringing this image, dragging it
in using impaint and correcting for. But
we're not going to do that right now
because there is a holistic problem in
the image. And that's the fact that I
don't like the perspective that we're
getting. Ideally, I'd want to see
something similar to the Star Wars
poster. It's a low angle shot with lava
in the background and each character
fighting aggressively against each
other. Now I could give this in the form
of a prompt, but I can also use image
prompt. Image prompt essentially allows
you to input an image and generate a
similar image. Now, but for right now,
let's turn off all of our advanced
features. You can go over to image
prompt and you can drag your image in
and you have a few features that will
show up. Now, if you don't have this bar
at the bottom, you just want to scroll
down to the bottom and click advanced.
That'll give you four options. Now,
image prompt will generally scan the
image very similar to a textbased
prompt. It'll attempt to understand
what's happening in the image and it'll
generally use that as a suggestion for
generation. You can see that it's
generally taking the idea of the image
and generating something similar. But in
this case, what we really need is
something that's visually similar. So,
we have two real options. Pyrochni and
CPDS. Pyrochani essentially maps the
position of characters and takes it into
your new image which would be ideal for
this particular image and CPDS
essentially uses contrast, color and
saturation to generate a similar image.
So in this particular case, Pyrocheni
would work perfect. Okay, we've got four
images in. Let's let the remaining
images come in. Let's look at the first
one. So this first image, the first
thing you'll notice is that it's fairly
low resolution. There's not a lot of
detail in the face structure. The
lightsaber's broken. I don't like how
the sand looks relative to the mountain.
The second image is much better. And
then it's converted Obi-Wan into some
version of Anakin. So, that's really
great in terms of perspective, in terms
of each lightsaber. Lightsabers are red,
but I'm assuming I can make those
changes. Third image. I don't like it as
much. The resolution is not quite there.
The perspective isn't great. Anakin's
looking a lot bigger than Padme. It's
probably one we're going to avoid. Image
four. That's looking much better. I love
the perspective of this. This next image
is completely broken. perspectives are
wrong. Seems to be a dual lightsaber
which Padme is holding from the wrong
side. So, we're not going to use that
image. Final image looks like two Padme
fighting each other. Again, it's less of
a problem because I know I can change
one of those characters, [music] but I'm
not a big fan of the perspective. I'm
not a big fan of the background. So,
with all of that in mind, I think I'm
going to go with this image as the image
I'm going to refine to final.
Okay. Okay. So, the first thing I'll do
is drag that image into impaint and
start erase my previous painting and
start working on specific sections of
the image. So, I'd want Padme to wear
her iconic white battle attire. So,
these images are already more in
direction of where it needs to be. This
time, I'm only going to change up the
outfit itself. Just highlighting very
specifically. So, our first section of
images are coming in and they're all
looking pretty great. So, I love the
position here. Doesn't look perfect in
terms of outfit. Second outfit looks
much better, but there's lots of
problems with the hand structure. The
hand seems a lot smaller than it needs
to be. The foot structure isn't great.
This third image is better in terms of
hand structure, but there is a broken
limb right here, and there's a few
issues going on in the lower parts of
the frame. So, here's a pro tip. In this
particular section, I'm changing how the
body structure looks, but I don't really
want to change the general position. I
find that it's helpful to highlight only
parts of the image. So you'll notice
here I've left the foot open and the
back of the ankle open as well along
with the waist of the character. That
way I can change just the middle section
of the image without affecting the
overall look of the Now I'm liking the
overall structure here. I just want to
fix this broken limb. I'd like to change
the shoes the character is wearing.
Maybe reduce some of this armor. And
that's what I'm going to do now. Now you
can see that after a certain point I'm
going to start running parallel
processes. In this section I've
highlighted Padme's hair very
specifically. So I'm trying to get her
hair tied up and it's proving to be a
little bit difficult. But the important
point here is that you need to keep a
reference ready. I'm trying to match
those references. Having a reference
will always help get the details right
and anchor your character in reality.
So, one thing we'll definitely need to
do is work on character features,
especially in the face. You want to have
as much detail as possible in the face.
So, even if things aren't perfect, most
people won't notice. And you can see I'm
already painting in the second image
while that first image is generated. I'm
going to need to work on that next. So,
Anakin, Star Wars, detailed face, it's a
lot better. We're not going to perfect
anything. We're going to do the best we
can and really just go through all of
the different tools and functions. If
you'd want to see us explore that in
another video where we deep dive into
creating hyper realistic images, let us
know in the comments below. Now, in this
next section, I want to show you how to
expand an image. Now, you've probably
seen something similar if you've used
Adobe's Photoshop with content of
airfield, but it's quite powerful in
focus because you can really control and
refine how the expansion work prompt you
put in here is very specifically the
information that you want in the
background and in the expanded areas.
But ideally, I'd like to have this in
landscape format. So, what I'm going to
do is I'm going to take that image, put
it back into content expansion, and just
generate the left and right side of the
image. Okay, that's looking really
great. I love how that looks. Now, I
don't really like how the characters are
standing. I'm seeing a lot of
imperfections in the hand structure,
mannequin's leg structure, but generally
this is a great starting base
considering we generated this image in
under an hour. That's a really great
starting point. Now, there's a lot more
you can do here using the refiner tools.
You can bring in things like smoke and
fog. You can work on the character
outfits, bring in detail, so things like
the shoes and the hair. I'll also look
back at my references and really
understand where I missed the ball. And
that brings us to the end of this
episode. If you like this video, hit the
like and subscribe buttons. Now, there
is a lot more we can do to perfect this
photo.
