---
title: 'Building AI Agents in Pure Python - Beginner Course'
source: 'https://youtube.com/watch?v=c9AnqCeyxbI'
video_id: 'c9AnqCeyxbI'
date: 2026-09-20
duration_sec: 2244
channel: 'Tech With Tim'
---

# Building AI Agents in Pure Python - Beginner Course

> Source: [Building AI Agents in Pure Python - Beginner Course](https://youtube.com/watch?v=c9AnqCeyxbI)

## Summary

This video demonstrates how to build AI agents from scratch using pure Python, without any frameworks or third-party tools. The presenter explains the core components of an agent—model, tools, and loop—and then walks through coding a functional agent step by step, culminating in a mini coding agent that can create and refactor a Snake game.

### Key Points

- **What is an AI agent?** [00:38] — An agent is essentially three things: a model (LLM), tools, and a loop. Context memory can be added but is often wrapped into tools.
- **The model is just text generation** [01:16] — The LLM is the brain that does reasoning, but it only generates text. It doesn't call tools or access memory; it simply takes tokens in and produces tokens out.
- **Tools can be anything** [01:32] — Tools can be MCP servers, third-party services, web search, or custom Python functions. In this video, tools are Python functions, but they can be extended later.
- **The loop enables problem-solving** [02:00] — The loop allows the agent to make multiple requests, call tools, get responses, and iterate until the task is complete. Without a loop, it's just a chatbot.
- **Big players use the same pattern** [02:30] — Tools like Claude Code, Cursor, and Codex all implement the same three components—model, tools, loop—under the hood, just packaged cleverly.
- **Step 1: Calling an LLM** [03:15] — The first step is to call an LLM via an API. The video uses OpenAI's API with the 'openai' Python package, set up using the UV package manager.
- **LLMs have no memory** [07:31] — LLMs don't remember previous conversations. You must pass the entire conversation history as messages each time to maintain context.
- **Message roles: user, assistant, system** [08:23] — User messages are questions from the user, assistant messages are the LLM's previous responses, and system messages set the initial context or instructions.
- **Step 2: Managing conversation history** [13:25] — To give the agent memory, you maintain a messages list that grows with each user input and assistant response, and pass it back to the model.
- **Step 3: Tool calling** [17:18] — Tool calling involves sending a list of tool schemas to the model. The model responds with a request to call a tool, and the developer's code executes it and returns the result.
- **The developer writes the glue** [19:40] — The LLM only generates text; it doesn't execute tools. The developer must write code to actually run the tool and feed the result back to the model.
- **Building a mini coding agent** [28:13] — The final example combines everything into a mini coding agent with four tools: list files, read file, write file, and run command. It can create and refactor a Snake game.

### Conclusion

Building an AI agent in pure Python is achievable by understanding the three core components: model, tools, and loop. This foundational knowledge demystifies how commercial agents work and empowers developers to create custom agents without heavy frameworks.

## Transcript

In this video, I'll show you how to build AI agents in pure Python. That means no framework, no vibe coding, no third-party tools. Everything will be written directly in Python code.
And by the end of this video, you'll not only have a fully functioning agent that you can use and extend, but you'll understand how AI agents work a lot deeper. So with that said, let's get into it. And just a quick thank you to HubSpot for sponsoring today's video.
They've got an awesome free resource on AI agents that I'm going to share with you later. so stay tuned. So before we build an AI agent, we should understand what it actually is. I'm going to go through a little bit of theory, and then we're actually going to go into the code editor,
and I'm going to show you multiple steps. We'll code everything out, and then you'll have a functioning agent at the end. So what is an agent? Well, an agent is really just three things. We have a model, we have tools, and then we have a loop. That's pretty much it. And then you can
add, you know, context memory in there as well, but that can kind of be wrapped into the tools. Now, the model is really just typically some kind of API. So, I mean, it could be a local model, something running on our own computer through a Lama
or through Docker Model Runner or LM Studio, what we call an API to use it. Or it could be a closed source model, something from Anthropic or OpenAI or some other provider, okay? Now, the model or LLM is really the brain.
This is the thing that's doing the reasoning, and really all it's doing is actually just generating text. So we give it some text in, it gives us some text out, that's the model, okay? And these models are smart because the text they give us out is meaningful and we can process that and use it in some way.
Next we have tools. Now tools can really be anything. It could be an MCP server, it could be a third-party tool, it could be web search, or it could be a Python function that we write ourselves. Now in this video we're going to make our own tools, and those tools will be Python functions,
but you could do anything inside of them and then later connect to other tools as well. We also got to have a loop. Now, the loop is the important part because if we just send one request to a model and it gives us a response back, that's really just like a chatbot, right?
An agent can actually solve problems. It can do tasks, and the way it can do that is by looping an indefinite amount of time. It could have 20 requests, 30 requests, 40 requests. It could call a tool, get a response, call another tool.
Hopefully, you get the idea, but we need this loop that allows the agent to kind of keep running in the background and then just tell us when it's finished. If you think about any of the big players out there when it comes to AI agents, like ClogCode, Perse, Codex, whatever, under the hood, they're effectively all the exact same
thing, and they just implement these three things that I talked about really cleverly, right, in a really well-defined piece of software. So ultimately, they have an MLM, they have tools, they have a loop, and then they have a bunch of other kind of side features that really decores the three things right here.
So what we're going to do now is we're going to start coding this out. Now, of course, I'm going to explain it more in depth, and we're going to understand this deeper as we go through, but there's three steps that we're going to go through. The first step is actually just calling an LLM or using an API.
This is kind of what it looks like, but of course, we'll put it in a code editor. The next step is going to be understanding kind of memory or conversation history. And then the third step is actually going to be calling tools. Then what we're going to do is we're going to wrap this into a full kind of piece of software to model like a mini-clad code.
So we can actually have tool calls go back and forth, loop around. you'll see what I mean, but let's go onto the computer and do step one. So I'm on the computer here, and like I mentioned, the first step is actually going to be calling an LLM
or triggering some kind of response through an API. Now, in order to do that in Python, there's many ways that we can set this up, but what we're going to do is we're going to use OpenAI. That's because this is super popular, it has frontier-level models,
and this is something you'll commonly do. But the first step when you build an agent in Python is kind of picking an LLM. These LLMs are again typically available through an API, Application Programming Interface. The API could be local, meaning it's running on your own computer through something like a LAMA for a local model,
or it could be public, right, or something available on the internet like the OpenAI API. So what I'm going to do is I'm just going to start by setting up my project using UV. So you'll notice that I'm inside of Cursor here. You can use any code editor that you want.
I'm going to go to my terminal. I've just opened a new folder called Mini Cloud Code because that's what we're going to build. and I'm going to type uv init and then dot. uv is a package manager for Python. If you're not using it already, I would highly suggest installing it and setting it up.
It's very basic. What this will do is create a new uv project for us. Now we can delete this main.py file as well as the readme.py file because we don't need those, but the rest of them we'll leave open.
From there, we're going to type uv add and then I believe this is just going to be openai. That's because we're going to install the openai module. This is just going to allow us to send a request to OpenAI super easily. It's not really a framework
It's just gonna prevent us having to write a bunch of tedious code that we don't need to. Alright? This will be the only thing we install, but we just need it for this video. So, uzadd OpenAI I believe that's the correct name of the package and then it should be installed.
So what I'm going to do here is just copy and paste some sections of code. That's because I don't think it's that valuable for you guys to watch me write it manually and in most cases cases, AI will generate the code anyways. And if you guys want all of the code that I write in this
video, I'll leave a link to it in the description. I'm just going to put it inside of my free school community. So if you join the community, you can access this as well as a bunch of other stuff. I'll leave a link to that in the description. Again, join, and there'll be a direct link to
where the code is available. Okay, so what I'm going to do is I'm just going to paste this in here. And this is kind of the step one, which is triggering an API. So you notice we say from OpenAI, import OpenAI. We create an OpenAI client, which we're going to look at more in the
second, and then what we do is we generate a response. Now, what I want you to understand is that when we're using an API or we're using an LLM, the only thing that the LLM does is generate text. That's all. It doesn't call the tool. It doesn't go save a memory.
It doesn't go to the web and research something. It simply generates text. You give it tokens in, it gives you tokens out. So how do you turn that into an agent? Well, that's what we're going to do in the next steps, but effectively what's going to end up happening is we're going to
be sending multiple requests to the API, and we're going to ask the model to tell us what we should do. So we're going to tell the model, hey, we have five tools available, would you like to call one?
It's going to give us some text back that will effectively say, yes, I want to call tool one. Then we're going to call tool one, so we're actually going to run the tool call on our own computer in software. We're going to get the response from that tool, and we're going to give
it back to the model. We're going to say, hey, you asked me to call this tool. Here's the result of that tool call. What do you want to do now? And this is that looping feature that I'm talking about that we're going to get into. But in order to do any of that, we need to create a response
or a completion. So this code chunk that you see right here is exactly how you do this. It will vary depending on what model or API you're using, but it will look very similar to this. So you'll So the code is that we create the client. We're going to put an API key here in a second, by the way, so we can actually get peeled for our usage.
And then we say response is equal to client.chat.completions.create. If you don't know how to find this code, you can just ask AI to tell you how to make the completion with an LLM. But in the case of opening AI, this is how you do it.
Now, when we do that, we specify the model we want to use. It's pretty straightforward. We're just using a cheat model for this video. And then, importantly, we pass in the messages. Now, messages is how the model actually knows what's been said and the conversation that's occurred so far.
And that's because LLMs don't have any memory. They don't actually know what was just said. All they do is they take all the information we gave them, and they give a text back. And if you think about it, they're actually a lot stupider than people make them out to be,
because they don't have all these features built in. So what we're actually doing here is we're passing in an array or a list of messages, and these messages represent the conversation history so far. Now at this point, there's been no conversation
and we're just manually putting something inside where we're saying, hey, we have this message. The role, which we'll talk about in a second, is the user and the content of the message is explained what an AI agent is in one sentence.
So this is effectively the prompt or what we're passing to the LLM. Now let's go back to these roles for a second. So you see role user. Now a user role just means that this is a question that the user, right, or the person who's using this LLM has asked.
But there's two other rules that are important to understand. And those other rules, I'm just going to simulate a message, is assistant. Assistant is actually the rule that represents what the LLM has said. So, for example, the assistant might have said something like,
an AI agent is simply an LLM connected to tools in a loop. Whatever, okay? So the idea is that assistant is tracking what the LLM has said previously. So user asked a question, assistant responded with this.
We actually need to keep track of that and pass it back to the LLM. And then we have one more role, which is a system role. And this system role specifies something that is non-assistant or user message,
but that should be loaded in and understood by the LLM at the beginning of the conversation. So typically the system role is going to go at the top of the messages. And what you would have in system is something like the following. you are a helpful AI assistant answer concisely in short sentences okay and that what you put there now one thing just to call out here if you guys notice I use this kind of AI dictation tool If you haven seen it before it called Whisper Flow
I'm just going to open it for you because it is super useful, and you can see I use it literally all the time. So if you notice me in this video transcribing the text, I'm using this. I'll leave a link to it in the description. It's free to try out, and I need pretty much everyone that watches the channel that uses
it, because if I did, it did. Anyways. Point is, right, we have these three types of messages, system, user, assistant. We actually need to keep track of what these messages are in order for the LLM to know what's going on.
So for now, we'll just go back to the user message and we'll test this out in a second. But I want you to understand what those are because you're going to see them quite frequently. Okay, and then the last thing here is we just have this print statement where we're taking the response from the model and we're printing it out.
Pretty straightforward. Now, the model will give us a bunch of information, but what we care about typically is just the raw response. The raw response, or the text response that it's giving us, comes in this particular field.
I'll show you what the total response looks like in a minute, but for now we're going to go with that. Okay, so if I try to run this right now, you're going to notice, let's just go in UV run, and then step1.py,
that it's not going to work because I'm not connected to the OpenAI API. The reason for that is I haven't provided an OpenAI API key. So it says, hey, you need to set this OpenAI API key. So what we can do is actually just pass a variable directly inside of here that says API key,
and then pass in our OpeningAI API key. Now, you also can store the API key in an environment variable and then pass it in there. For now, I'm just going to put the raw text just to save us and make it a little bit simpler. But just keep in mind the API key, obviously, you don't want to expose,
but it's better if it's in something like an environment variable. But what I'm going to do is generate an API key. To do that, I'm going to go to platform.openai.com slash API keys. I'm going to make a new secret key. I'm going to call this delete later because I will delete it after the video.
And then I'm just going to copy it and I'm going to paste it here inside of my cursor. Okay, so now we have the API key. And now if I go here and I run this, it should actually give us a response. Hopefully it'll be pretty fast because we're using GPT-5 Mini.
And let's see if we get it. And you can see if I scroll up here, it says an AI agent is an autonomous system software robot that perceives this environment, reasons, whatever. You guys get the idea, and you get the response.
Okay, so that was step one. Whenever you're building an agent, you need an LLM. You need to be able to call it. This is kind of how you call it and the formatted way that it works. Now let's move on to step two. Okay, so we're going to dive into step two here.
but I want you to understand that once you build the AI agent, the difficult part is actually knowing which tasks you should hand off to this agent and how it's going to be useful in the first place. Now, that's actually a perfect segue to tell you about a great free resource.
Now, HubSpot put together a free playbook called AI Agents Unleashed, and it answers exactly that question. Now, it's built from interviews with HubSpot's CMO and their SCP of marketing on how they actually deploy agents inside of HubSpot.
So it's not theory, it has real implementations. And my favorite part of this guide is the, is this an agent job decision tree? Now you take any task and run it through four quick checks. Is it repetitive? Does it use structured data?
Is 90% accuracy good enough? And can you measure success? Now if it passes that, you can automate it. If not, you keep it human on that task. Now that one framework is going to save you from building a bunch of agents that nobody ever needed.
And you also get real use cases that are working right now, like agents that research every new sign-up and predict how they'll use your product, plus a full-phase roadmap for rolling agents out without breaking everything. And they include a bonus checklist PDF that you can literally check off as you deploy agents.
So if you're watching this video because you actually want to build a useful agent, this is a great layer that goes with it. It's completely free. I'll have the links to it in the description, also in the pinned comment, and a massive shout out to HubSpot for sponsoring this video.
So let's get back into it. Okay, so like I mentioned, step two. We know how to call the API. Now what we need to do is actually keep track of the conversation that's occurred so the agent doesn't just forget what we asked him.
Now what you need to understand is that memory is not just a list, right? And we kind of talked about this a little bit already. You'll see that if I simulate this by just sending a few messages here, you can see we have user messages, like I discussed, right?
And then assistant messages. And then eventually, I'll show you later, we'll get to the point where we actually have tool calls that are coming in. Now, the reason we need to keep a list of all these messages and kind of repass it to the model is because the model doesn't remember anything by default, like we discussed.
So all of this stuff that you see that we're going to pass into the model is really what we'll refer to as the context window. So when you hear people talking about, oh, this agent has a 128 token context window or whatever,
really that means how much space or how many messages effectively that it can store. And these messages are broken down into tokens, which you can think of like individual words. So when we hear terms like context engineering, right,
really that's us deciding which things we give to the model and which things we don't. Because the model doesn't store any of them itself, it's up to us, the developer, to actually pass them in. So I've just copied in the code here for step two.
It's pretty simple, so we'll just go through it line by line. Same thing, we have declined as before, so we can call open amp. Now notice that we keep track of a messages list. This is a list that we're going to use, right, and actually populate and keep adding to
to keep track of the conversation history. We then create a while loop. This is just going to allow us to keep chatting with the agent and getting the response back, so we'll just kind of go back and forth. We ask the user for some input, so they can type into the terminal.
If they type agent or quit, of course, we'll get out of that. And then notice the first thing that we're going to do is we're going to take whatever the user was typed thing, right, whatever they asked, and we're just going to append that into our current messages. So for the first message, this will be the only thing that's inside of here, right, just
whatever they asked, but then we're going to keep appending more. Next, we generate the response, like you already saw, right, we get the text content, and then we append whatever the assistant said back into the conversation history.
So this message list is just going to keep expanding and expanding and expanding and expanding and expanding as we chat with the bot so he knows what we talked about before. Now in step one, we didn't do any of that, right? We just asked it one kind of blank question,
and even though we put this in a loop and we kept asking it one question, it would never know what we said before. And I'm going to show you what I mean by that by actually breaking this example in a second, but for now I want to run this so you can see what it looks like.
So I'm going to type uv run, and then what is this, step2.py, and let's test it out. So I'm going to say hello, my name is Tim. Okay, and it's going to give us a response back here in a second, and then I just want to confirm that it does have memories.
I'm going to say, what's my name? And it should know that it's Tim and says, my name is Tim. Okay, cool. Boom. And this is our conversation, right? So it remembers what we're talking about. Now, if I quit this or get out of it,
and I simply do this where I remove this messages.append, right? So we remove that and we change this to just be the following so that messages becomes just this one list that's the same every single time
like we had in the previous example. You're going to see what happens here when we run the demo. And I know I'm going a little bit fast, but I just want to show you. So I'm going to say, my name is Tim. Like this. Okay, it's going to say hi or something. Nice to meet you.
I'm going to say, what's my name? And let's give it a second. And you will see, come on, I don't know. You haven't told me. But I did just tell it.
The reason it doesn't know is because, again, we didn't keep track of that conversation history. So that's why this is so important, and I put this as a separate step, so you understand that the model doesn't remember things you need to manually tell it.
So this is kind of the setup of, okay, how do we actually keep track of the conversation history, know what was done previously? Now let's move to the next step, which is definitely the most exciting, which is tool calling. Now, tool calling is really what allows agents to feel useful and like they can produce meaningful work.
And the way that this works is that you send a list of tool schemas, kind of like a menu of options, to the agent in the prompt. And typically this will come in as part of the system prompt.
It also can come in in different types of messages. But you'll tell it, hey, these are tools that you have available to you. These are what the tools look like. This is the parameters. This is how you call them. This is how they function. And then the model can actually give you a response that indicates the tools that it wants to call.
So you'll notice on the left side of my screen that I have what's kind of my menu of tools. Now in this case, we just have one tool. The tool's name is refile. It has a description of what it does, and then it has some parameters. This is how you actually call it, right?
Then what we would do is we would send this to the model, send it to the LLM, and you'll see that the LLM here will actually give us a response back, sometimes, not always, depending on if it wants to call a tool, that says, hey, this is a tool that I want to call.
It wants to call a function. The name is this. These are the arguments that it wants to pass to the function. So what will happen is, like, the model doesn't actually call the tool. It can't actually do anything. It's just producing text, right? But it will tell us that it wants to call the tool.
So it's like asking us, hey, can you call this tool for me? So then what will happen is our software, our agent, right, or our harness, will call the tool by actually executing the code that does some operation right runs the tool It will then get a response So hey this was successful this worked or here what the tool gave you right and it will pass it back into the model
And think about all of these models when they're running bash commands, for example. What are they really doing? They're not running the bash command itself. They're not using grep. They're not creating a file or making a directory. They're simply saying, I want to make a directory.
I want to open this. I want to know the content here. and then our system executes that command and gives it the response. And that's why whenever the model tries to do something that maybe it shouldn't be doing, you'll get a gateway or kind of a guardrail pop-up that says,
hey, do you want to approve this? Do you want to run this tool call? Should this happen? That's the harness or the agent software in the background saying, hey, I don't know if we should run this tool, so I'm going to ask the user if we should.
Hopefully you get the idea that that's kind of the guardrail component. So I just want you to really understand that agent doesn't do anything, sorry, LLM doesn't do anything other than generate text. Calling the tool requires the developer to write the glue
that actually calls the tool and gives the response back. So what we're going to do now is go back to the computer, and we're going to look at example three. I've copied in some more code, but again, a lot of it is similar to what we've seen already. Okay, so same thing, step three code.
We have our kind of import. I also imported JSON. This is built into Python, so you don't need to install it. We have our client. We then have a tool. Now, this tool is a read file tool. What it can do is it can take a path, and it can read the contents of the file.
And notice that all it does is it just returns whatever the content of the file is at .read. Okay, so just a super simple tool. Now, whenever we have these tools, we need to provide what's called a schema. A schema is essentially a breakdown of how the tool works, what the parameters are, what it's going to return, etc.
So, we have to write this manually because we're not using any frameworks. However, if you use some frameworks like LandChain, for example, they'll be able to automatically generate tool schemas for you based on things like Python functions.
And of course, you can have other types of tools like MCP servers, but for now, we're just going to stick with the basics, but it makes a little bit more sense. So we'll need to create this schema, again, if we're doing it purely manually, which is the point of this video, that defines what the tool looks like.
I don't pass this function anywhere. The LLM can't see the function. It just sees this schema. Now you don't need to memorize what the schema looks like, because most times you won't need to write it yourself. If you are going to do it fully manually, then what you can do is just have AI generate them for you.
But you'll see that we have a type, so this is the type of the tool, which is a function. We then say, okay, here's the function itself, the name is readFile. The description, so this is important because this will tell the LLM when it should actually use this tool.
So it's really important that the tools are named accurately and that the descriptions make a lot of sense. And then same thing with the parameters, we define that it's type object. We have a property, which is a path, right? And then this is type string.
The description is path of the file to read. So the model knows how to actually call this tool. Okay, then we specify the required parameter is path. So we know we need to pass that in. Okay, if we keep going, you'll see we have messages.
And this is just a manual message I've put in. So we don't have to actually type something in. It's just going to try to look for notes.txt. And this is how we actually call the tool. So what we do here is we say, all right,
We're going to get a response here from OpenAI. Now, this time, we've passed an additional field. Notice this. It's called tools. This is built directly into the ChatGPT API or OpenAI API that allows us to pass in different tools.
Different APIs have different ways of providing tools, so you'll have to change how you provide the tools based on what model you're using. But typically, it will be in this kind of format because this is what's called OpenAI compatible.
OpenAPI compatible. It's a little bit weird how they have the standards, but it's a way that all these APIs typically behave. They have this kind of set format, so you can usually call it like this. Okay, so then what we do is we get the message response. Now, notice this time that when
I gave us the message, I didn't add the additional field here, which is the actual content, right? I'm just getting the message. So, if we look at the other stuff, just so you can see, when I want to print something out, I'm saying, hey, I want to print out the content. The content is human
readable. This is something that I should actually see as the user who's using the LLM. However, the LLM can give us another response, and that response can be tool calls, right? Just like we looked at, where it's telling us, hey, I want to call these particular tools. So just like before,
in my messages, I'm going to append whatever the message was that we got from the LLM. This time, I'm not just appending the human readable content, I'm appending everything, including the tool calls. Okay, so then we go down here, and we say no more tool calls. Alright, so if there are no
more tool calls to be done, that means that the model is finished, because it doesn't need to call any tools. However, if it has a tool call, what we're going to do is we're going to loop through those tool calls, because they've actually given us multiple in one response,
and we're going to execute. Now to execute them, we're going to take whatever the tool call.function. arguments was, because this is the format that is going to give us the response back in, We're going to say, hey, it wants to call this function with these arguments,
and then we're going to call the function, and we're going to append whatever the responses are back into the messages. And the next time we call this, it's actually going to see whatever the content of these tool calls are.
So you'll see what this looks like in one minute. I also want to just make another print statement here so you can see what the tool calls look like. So let's just print out the message itself so you can see this. And let me run this so we can get kind of a quick example,
and I can go into this a bit more in depth. So then we're going to do an even better example in a second where we have like a full agent running with multiple tools. But for now, let's run step3.py. Okay, so give this a second here.
It also should fail, yes, because there's no file in those.txt. So that's fine, but let's look at what we've gotten so far. Okay, so you can see that when we printed out the, let's see here, message, it gave us a chat completion message.
There was no content because it's not giving us anything human readable. And if we scroll through here, you can see the role as assistant. And if we keep going, it says, hey, there's actually a tool call I want to do, right? It's under this little bit difficult to read tool call.
The tool call is a chat completion message function tool call. The ID is this. The function is the function, right, that we told it. And then we want to call with arguments path notes.txt. The name is read file.
The type is function. And then it says the model wants to run read file with path notes.txt. So our code tried to execute that function, but we got an error because that file didn't exist. So, of course, we would need to improve the tool so that this works.
For example, we could say, you know, try, right? And then just accept file not found error and tell it, hey, file not found. And it would probably just keep retrying. In fact, let's run that and see what's going to happen now
when we fix the tool so that it can handle an error. So same thing, it says model wants to run this, read file, okay? So we've now run this and got back to it. says, hey, I couldn't find notes.txt, you know, tell me what it is, and it didn't give us any
more tool calls, so we stopped executing, and it told us, hey, I couldn't find this, upload the file. So now let's just quickly add the file, notes.txt, hello world, okay, and now let's go
and run this one more time. I just want to see what the response is, but I'm trying to show you the tool calling process here so you get an idea, and if we go up here, you can see it says, I want to call this tool. I want to run it with notes.txt.
Okay, the file contains, you know, whatever, hello world, and then we get the response. And then the last thing to quickly look at, just so you guys can understand the flow here, I'm just going to print out at the beginning of this,
should be beginning, let's do it at the end of this, what the messages list actually looks like. So I'm going to print messages like that. Now, when I do that, you're going to see the list will just keep getting larger with all of the tool calls.
So let's give this a second to run, and you can see, where is it? We have the initial list, right? And sorry, because of where I put it, we're only going to see it one time, but you can see it says role, user, write content, what's inside notes.txt, okay?
Then we get this chat completion, which is what the agent wanted to do. It wanted to call this particular tool. Then we have the role of a tool call, right? So we're actually telling it, hey, when you called the tool,
the content of that was hello world, and then it gives the response, notes.txt contains this content. So all of that information is being stored in this message log, which allows the LLM to reason on it and perform the tools.
So that's kind of the three steps of the agent, right? And you've seen the loop as well. This loop is kind of implemented so that it can just keep calling tools. If you wanted to see this, I could say, what's inside of notes.txt, notes1.txt, and notes2.txt.
And now if we run this, it should actually attempt to do three tool calls. So let's see if it does it in one or if it does it in three steps. But there we go. So you can see now it gives us like three tool calls that it wants to do.
We get all three tool calls, and then it gives us a response, right? And if I told it to do it step by step or check notes one, then check notes two, then check notes three, et cetera, you'd see they would loop through, and it would run multiple times until no more tool calls are required.
And that's what this is doing here. It's saying, hey, if it doesn't want to call any more tools, nothing more for us to do. Let just print out the content and keep going So hopefully this is clear Now what I want to do is build this into a much better example that combines all of this and makes it a lot more realistic where we kind of make a mini code Okay so for our mini code we need to make this something that actually
able to write code, and in order to do that, I'm going to give this four tools. Now the four tools that we're going to provide here, just for this simple example, is listing files, reading a file, writing files, and then running a command. Now a command can be a shell script, where it actually
executes code, you know, does a cd, makes a directory, whatever. However, for something like a shell command, which can be destructive, we want to make sure that it actually asks permission before it runs that. And to be honest, probably same thing with write file,
because that means it can override or delete a file, but for now we'll just leave it without the guardrail. Okay, so those are the four tools, that's kind of how we're going to build this, and then we need to build the guardrail in, we need to wrap the context around, we need to keep track of the messages, handle the tool calls, you get the idea. So let's
go back in here and we're going to start constructing this kind of line by line. So let's start by going to step one and I'm just going to copy in my kind of client because of course we're going to need that. Now what else I'm going to do is I'm just going to copy in a model as well as a system prompt
and then a few other modules that we're going to need. Now again rather than me just writing it line by line I'm going to kind of copy and paste in sections and then explain what we're doing. Okay so of course we've specified the model we just put it here so we can change it really easily
and same thing like in Cloud Code when you do slash model, then you just switch the model so you switch where the requests are going. Now for the system prompt, what we're doing here is just giving some kind of like general info
to say, hey, you're a coding agent, you have access to files, you can do these things. So that's a bit of context that's always in the prompt before we move on to the next step. Okay, now after that, we need our tools.
So I'm just going to copy in all four of the tools here. The tools look like this, right? So we have list files, let me make another space here, we have read file, we have write file, and then we have run command.
And you'll notice inside of run command, we have this kind of little user input thing that says, hey, do you want to run this command? And if you type no, it won't run it. If you type yes, it will run it, we'll get the result, and then we'll output it.
Now we're doing that using subprocess in Python. Don't worry too much about every individual line of code. I just want to show you the process and again, so you understand the theory and you know how to build this out yourself. Okay, so we have four kind of, again, commands, tools, whatever that we can use.
Now next, we're just going to make a list of these tools, just going to make this easier for us to call them later. So we have this text list files associated with the function, read file, write file, you get the idea.
Okay, now next, this is the most annoying part. Again, that's why I'm just copying it. This is the tool schema. Now, let me zoom out a little bit. This is all four of the tools just kind of translated into a format the model can understand.
So, you see we have list files, we have read file, we have write file, we have run command. Exact same thing we saw before, just for all four tools. So, you'll notice like a lot of the code is going to be related to the tools. That's just because the more tools we have, the better the agent's going to perform, right,
in terms of what it's capable of doing. So, we have those four tools. We have the list. We have the schemas. and now we need to move on to actually being able to execute them. So in terms of calling these tools, first we're going to have a function that will actually run the tool.
So this is just a simple function. What it's going to do is just print out the tool call so we can see a log of that. It's also just going to grab the actual name of the tool and then just kind of parse this in so we can go from string to a real Python function. Then if there is an error which can occur,
it's going to tell us what it is rather than failing the tool call. Okay, next we have a function, and this function is going to represent one run of our AI agent. So one run means effectively one user prompt.
So I give the prompt, hey, write me a snake game in Python and then this will execute. And then we'll need to have a loop again that will allow us to continue to chat with the model, which you'll see in a minute. So from here, we'll be out of the while loop
and this while loop is responsible again for kind of one sequence, or one prompt that we give to the model. So what we're doing is we're saying, hey, I want to create a response. I want to get the message from that. I'm going to append the message back into my conversation history, right?
And that is going to be as the parameter that we provide here. And then what we're going to do is check, okay, do we have any tool calls? If we don't, we're good. We can just return wherever the content is. If not, we need to actually make the tool calls.
So we run the tool, and then we get the response, and we put it back into this messages list. And that's it, right? We already looked at this. It's exactly the same thing we have in step three, where we're just doing this loop where we're effectively calling the tool, getting the response, giving it back to the agent. And then lastly, we need to just have kind of the
main driver code here. So we're going to paste this in down here. And what this does is just have a main function where we have a system message with our system prompt. We just say, hey, it's a mini agent ready to use. We have a while loop, which is kind of the loop for the
users that can keep prompting the model. We get some user input. They type quit. We quit. We take their user input, add it into our messages, and then we run this agent, right? So whatever they said, we run the agent on it, agent sees all of the tools, if it wants to call a tool,
it can, we run the tool call, and then we continue. And then we show the agent's reply, and we keep going until the user wants to quit. And we keep track of all of this context so at any point in time we know what was said previously. So this is 160 lines of code,
I know I just copied and pasted it, but I think this is a pretty efficient way to demonstrate it, and what we want to do now is run it and then play around with the code a bit so we can see what actually happened. So let's clear this and just go uv run and then this is agent.pline. Okay, mini agent
is ready. Type exit to quit. So I'm going to say, hey, my name is Tim. Okay, so let's just start with that just to make sure that this is going to work. Hey Tim, nice to meet you. I'm going to say, can you make a new
directory in the current folder here that is for a snake game that I'm going to create? Okay, let's ask it and let's if it wants to trigger or call a tool. Okay, so it says, may I run the command to create it?
I need, I'm going to say yes. Okay, that's kind of a weird way to ask me because it should just call the tool. And there we go. Okay, so now it wants to call the tool. It says tool, run command, command, mkdir. Okay, yes. So let's go ahead and run that. And you can see that now this folder
was created. It's now asking us to run another command. Let's go yes and just don't do that. and now it's created the directory. Okay, cool. So I'm going to say, now in that directory, make a simple snake game in Python using Pygame that I can play.
Okay, and let's see if it runs the tools to do that. Okay, so it's giving me a bunch of text here, which is a little bit difficult to see, but it did call a tool. The tool was writefile, and then for writefile, it has this main.py file, which is just created,
and this is a full snake game in Python. Okay, so now let's tell it something a little bit better. Let's say, oh, stir, yes, do you want to run that? Okay, I'm going to let it run a few of the commands here. I'm going to say, can you please refactor this so it's multiple files
and it's a little bit easier for me to read? And I just want to see if it can, like, read the file and then reproduce multiple of them. When we do the code, it's, like, a lot of stuff in the terminal to look at, but you get the idea. So let's see what we get.
Okay, so you can see now it did write file. It wrote, like, a smaller file here. I'm assuming we're going to get some more tool calls. Yes, so we have another one coming in here. Okay, another one, and you can see more of them are showing up. Boom, great.
It also is able to remember what it did previously, so it doesn't need to keep rereading all of the files, although it can if we need to. And the second year, we should get the output. Okay, so we now got the snake game.
I'm just playing it right now. You can see that it did a bunch of tool calls. It just told me I need to install Pygame, so I just did that. And if we go here, not that. Let's go to Pygame. You can see we have snake. I can move around, and we can play the game.
And it's fully functioning, you know, like it should. Can I kill myself? Yeah, there we go. Game over, and then R, and we're good to go. Okay, so that is Snake. And there we go, we have a fully functioning AI agent.
I'm just going to exit out of this for right now, and I just want to go back and kind of, you know, recap what we've done here, because I think it's really interesting how we can build this, and how it kind of explains what a lot of these things actually do.
A lot of times we look at these agents and we think they're super complicated. Yes, there's a lot of things that go behind the scenes, and of course it's going to be much more complex than what we show here, but ultimately it's just software that ends up using an LLM to orchestrate and handle different tasks.
You already just saw here that we can have the LLM call pretty much any tool. With that alone, we're already capable of building really interesting agents, and you can do them without fancy frameworks or super complex layers on top.
Now of course, using frameworks will simplify this quite a bit and give you code that can be much shorter and easier to read and parse, but I hope that this video was helpful in terms of demonstrating to you how this actually works behind the scenes.
If you guys want more content like this, then definitely let me know in the comments down below. Reminder that I do have a free school community you can join that's all about AI agents, so I linked that below as well. And with that said guys, I will see you in the next video.
Thank you.
