---
title: 'How AI Agents Actually Work (Every Piece Explained & Built)'
source: 'https://youtube.com/watch?v=HzGOWq5UyjY'
video_id: 'HzGOWq5UyjY'
date: 2026-09-20
duration_sec: 1168
channel: 'Tech With Tim'
---

# How AI Agents Actually Work (Every Piece Explained & Built)

> Source: [How AI Agents Actually Work (Every Piece Explained & Built)](https://youtube.com/watch?v=HzGOWq5UyjY)

## Summary

This video demystifies AI agents by explaining the model is only 20% of the equation. It breaks down the five essential components—harness, MCP, skills, sandbox, and production layer—in plain English, then demonstrates building a functional research agent using the open-source TrueForge harness. The goal is to give viewers a practical understanding of how real agents work, surpassing most people currently discussing them.

### Key Points

- **The Model is Not an Agent** [00:00] — A model is just a function: text in, text out. It can't search, read files, or remember. An agent is a model inside a loop that can take actions and feed results back.
- **The Five Pieces of an Agent** [02:18] — The five components are harness, MCP, skills, sandbox, and production layer. The model is only 20% of the agent; the rest is the unseen software around it.
- **The Harness: The Runtime** [02:30] — The harness is the actual software loop that sends prompts, executes actions, manages context, and keeps state. Tools like Claude Code and Cursor are harnesses.
- **MCP: The USB Standard for Tools** [03:39] — Model Context Protocol is a standard way to give agents access to tools, like web search or databases. It eliminates custom integrations, working like a USB plug-and-play.
- **Skills: The How-To Documents** [04:44] — Skills are markdown files with instructions for specific tasks, such as report formatting or tool usage. They are loaded only when relevant to save context.
- **The Sandbox: Safe Code Execution** [05:40] — A sandbox is a disposable, isolated environment for running agent code. It prevents arbitrary code from affecting your machine, crucial for production.
- **Production Layer: Sub-agents, Approvals, Observability** [06:46] — Sub-agents parallelize work, approvals enforce human checkpoints, and observability provides a trace of every step—separating demos from production-ready agents.
- **Building a Research Agent with TrueForge** [08:52] — Using npx truefoundry/trueforge, the video demonstrates setting up the harness, adding Exa for web search, enabling the WebArtifactsBuilder skill, and creating a reusable researcher agent.

## Transcript

Everyone is talking about AI agents right now, and almost nobody can tell you what one actually is. Now ask someone to explain the difference between an agent and chat GPT, and you're going to get a lot of random hand-waving.
But here's the thing. The model, so the actual AI, is maybe 20% of an agent. The other 80% is stuff that nobody explains. So what a harness is, what MCP actually stands for and why it matters, why agents need something called a sandbox.
So in this video, I'm going to explain every single piece of a real AI agent in plain English, and then I'm going to prove that I'm not just making it up by building one from scratch, live, with all of those pieces built in.
Now, everything we use is going to be open source, and by the end, you're going to understand agents better than most people shipping and talking about them. And a quick disclosure before we start here, this video is sponsored by True Foundry,
who just open sourced the tool that we're going to be building this later today, but everything you see is free, you can run it yourself, and I'm going to link the GitHub in the description so you can follow along. Anyways, let's get started with the single most important idea in this video, because if you understand this, everything else will make sense.
Now that idea is that a model is not an agent. A model is really just a function. Text goes in, text comes out. That's it. It can't search the web, it can't read your files, it can't run code, it can't even remember what you said in the last message.
Now, when ChatGPT feels smarter than that, it's because there's a bunch of software around the model that's actually doing most of the heavy lifting. Now, an agent is what you get when you take a model and you put it inside of some type of loop.
Now, that model looks at a goal, picks an action, and then something executes that action for real, and the result gets fed back in. Then it goes again, over and over, until the job is finished.
So it will search the web, read the results, decide what to do next, maybe write some code, run the code, look at the output, fix the result, you get the idea. Now, the model is the brain, but a brain that's just sitting inside of a jar can't do anything.
So the interesting part that nobody really explains properly is everything that's around that brain. And there's basically five pieces here, and I'm going to walk you through all of them, and all of this will start to make sense. But understand that a model is not an agent.
An agent requires a bunch of things around the model, And again, the model is really just 20%. The other 80% is what I want to get into now. Okay, so the first piece to discuss here is the harness.
Now, this is effectively just the loop that I described. So it's an actual piece of software, so it's written in code. It's not this fancy, you know, AI model, LLM, whatever. And the harness effectively holds the runtime that ties everything together.
So it sends prompts to the model. It executes the actions that the model asks for. It manages the context windows so the model doesn't drown in some history, and it keeps state so a long task can survive a reset.
If you use something like Cloud Code, ChashGPC, or Cursor, you've used a harness. You maybe didn't know that it was called that, but this is how you really use AI models in today's era. Now, the harness that I'm going to show you today is called TrueForge.
It's MIT open source, you can run it with a single command, and importantly, it's model agnostic. So that means that you can plug in any model that you want, even a local model or something with open weights, like something like Kine or DeepSeek, and then you can swap it whenever you want so you're not locked into a provider.
Okay, so that's the harness. Again, it's the runtime. It's really a piece of software that orchestrates the model, sends prompts to it, gets the results, executes tool calls, and it's probably the most important part, but the harness has some other components kind of inside of it or that it connects.
So the second piece here is MCP. Now you've probably heard of this before, and it stands for Model Context Protocol. It's really just a standard way to give AI agents access to tools.
Now here's the problem that this solves. Let's say I want my agent to search the web. So maybe I want to check my calendar, query a database, whatever, right? Before MCP, every one of these was a custom integration that you had to write and maintain yourself.
Now MCP turns each of these into a server that can speak one common language. So the agent can connect to it, ask what tools it offers, and then it can immediately use them. You can think of it literally just like a USB.
You don't have to write a custom driver every time you plug it into the keyboard. You just take whatever USB you want, you plug it in, and it works. So any MCP server will work with any AI agent as long as the harness will support that and kind of handle the communication and the bridge between the model and the tools.
In our build, we're going to connect an MCP server that handles web search. We'll talk about that in a second. and this is going to allow our model to actually see some other tools, start executing those, and pull real live data.
Okay, now piece three is skills. Now if MCP server is what an agent can do skills are how it knows how to do things well So skill is literally just a markdown file which is kind of a way to format text It typically called something like skill
standing for markdown, and it's full of instructions for a specific kind of task. So think of it like an onboarding document or an SAP that you might hand a new hire. Skills might include things like how you build a report,
here's a format that you need to follow, here are the steps that I want you to go through, here's the specific tool to use, here are gotchas or things to watch out for, you get the idea. Now, the agent is going to load this only when it's relevant. And that's important because a lot of times you'll have a bunch of
skills that can take up a lot of room and context. So the agent will know these skills are available, but it will only load the entire skill and read through it when it thinks it needs to use that. Now, piece four is the sandbox. And this is something that most people actually never even
think about or talk about. Now, real tasks need code execution. So analyzing data, processing files, generating a report, whatever. And this raises a pretty fair question, do you want an AI agent to be able to run arbitrary
code on your actual computer? Now, in most cases, you do not, right? Because one bad command, and it could be touching all your files, your credentials, everything, right, especially if it has sudo or root access.
So instead of doing that, the agent can get a disposable computer. This is called a sandbox, which is an isolated environment that can spin up whenever your agent needs to execute something. So we can run code with scoped access, and then as soon as it's
done, we can throw that sandbox away. So if it messes something up, it's not a big deal because it's not on your own computer. So the agent will get all the capabilities it needs with full power, but your machine is taking zero risk. And we'll talk about that more later, but when you're
building production agents, having these sandboxes is super important to make sure they're not running on your own hardware, and there's no risk factor of actually executing code. Okay, now the fifth piece here is the production layer. Now, this is what separates a demo from something that you could
actually run in production. And there's three main things that kind of go into here, so I've just put it into, you know, section five, or really it's three separate things. The first is sub-agents. Now, this is so a big task can be split across multiple parallel workers instead of just one
context window doing everything. This means that you're able to parallelize work so you can do it faster, and you can run the model, you know, multiple times without having to rely on just a single instance of it running. Next is approvals. So these are things that actually force the agent
to stop and ask for human approval before doing something sensitive. And then lastly, we have observability. So this is a full record of every step, every tool call, every result. So if something does go wrong, you can actually see what the agent did instead of guessing. And that's a pretty big
deal because when you go into production, you want to be able to trace through the results, see any optimizations you can make, if tool calls are failing, you get the idea. So that's the basic anatomy, right? We have the model, we have the harness, we have MCP, skills, sandbox, and then
that production layer with things like our sub-agents, our approvals, our guardrails, and then, of course, the observability. Now, I want to stop just talking about this. I want to actually show you how to build an AI agent that uses all of these different things that you can spin up on your own computer
and that you could actually deploy and access from different types of applications. So now, as promised, I'm going to show you how to build an AI agent so you can understand all of the pieces in practice. Now for this video, the harness that I'm going to use is something called TrueForge by TreeFoundry.
It's new, it's open source, it's free to use, and importantly, anything that you build here in terms of agents can actually be deployed. You can spin it up with Kubernetes, you can deploy it on your own server, and that way it can actually scale and you're not locked into a model provider's harness.
For example, if you were to try to do the same thing using something like Cloud Managed Agents, you're now locked into Anthropix pricing and you have to use Anthropic models. Even if you want to use entropic models, running the same model through an open source harness like this can be significantly cheaper based on how the harness is optimized, and you can see here you can actually run up to 50% lower cost using the exact same model.
You'll see what I mean as we set it up, but let's get into it. Okay, so in order to get this set up, it's literally just one command, npx truefoundry slash trueforge. Of course, you don't need to pay for anything, and if you're on Linux or Mac, you can just directly run it in your terminal.
If you're on Windows like me, then you're going to need to open up WSL, which you can see I have open right here, and then let's just copy the command, bring it back, and simply paste it on our machine. Assuming that we have Node.js installed, it should start running, and you'll see immediately that it spins up.
It may bring us through a quick install process, but of course I've already installed this, and then all of a sudden, the harness is running. Now once the harness is running, it will show you this local host port right here. You can hold control and then click on this, or just copy it and open it on your browser,
and when I do this you'll see that it opens up this graphical user interface where we brought into this harness I was just playing with it before so you can see I have some chaps on the left hand side Now once the harness is running we be able to actually create our own age instead of our own tools skills resources whatever But in order to do that we first need to configure some models
So because this is an open source harness, you can add any models you want. To do that, you just simply go into settings. From here, you have the ability to add some provided ones like OpenAI, Anthropic, Gemini, whatever. So if you want to configure Anthropic,
just enter your API key, you know, choose the endpoint, you can start using their models. Or alternatively, you can add a custom provider and add something like a local model. So in my case, I actually have my own DGX Spark running. I have quite a bit of compute here.
It's a $7,000 computer. So I'm running Gwenn 3.6, 35 billion parameters. I've just connected it to this base URL. And now I can do all of this completely locally and private from this harness, which I can't do in something like QuadCode.
So I have this set up here. You'll also notice we have connectors and skills, which we're going to have a look at in a minute. But for now, we just want to test that the model is working. So we can just go to the chat interface and just type something like hello world.
Now from here, you can see the model is selected down here, and immediately we get the response. It's also extremely fast, and we can see the full process and observability of this agent. So we're going to start building out a proper research agent here in just one second.
But before I add all of the tools and the connectors, I want to show you what happens if I try to treat this current chat like a fully functioning research agent. You'll see what I mean in a second, but let's give this a prompt. hey, I want you to act like a research agent, specifically an expert in local AI models.
I want you to do research and figure out what the best current local AI models are, at least open weight models. You need their speed, their cost, their availability, how they're performing, their intelligence levels, and all of this stuff in a research paper and an interactive dashboard.
Okay, so we're just going to give this a prompt, press enter, and see what we get. And by the way, if you're wondering what I'm using to dictate here, I use a really cool tool called WhisperFlow. it's free to try out. I'll leave a link to it in the description. You can see I've written like two full books using it,
and that's what I use for pretty much everything when it comes to actually chatting with the agent. So anyways, for now, you can see that it is reasoning through this process. It's going to start doing this. It's spun up a sub-agent, and we will get like a decent result here from this particular,
what do you call it, prompt. However, we haven't added any tools, right? So I just want to show you kind of a before and after here. So let's see the result that we get, and then we can compare it after we add all of our tools, skills, etc. Okay, so we've gotten the results. It's quite long, but what you're going to notice immediately is that all of the results it's giving us is just coming from its training data, because we haven't actually connected this to a web search tool, so it still thinks that we're talking about, like, 2025 here, and if I scroll down, it's showing you, sorry, like, mid-2025 updates, anticipated late 2025, because that's when this model was trained, or at least when the training data was available.
So, like, we're currently on Gwenn 3.6, and it doesn't have access to that because we haven't given it the correct tools. So anyways, point is, what I want to do now, if I want to start connecting up the tools to make this agent useful, let's do that by going to the settings.
So from settings, there's three things we're going to look at. Connectors, skills, sandbox providers. For the connectors, I want to add an MCP server. Now, you can see there's a bunch of available MCP servers. We also can just add our own custom servers here.
But what I want to do is I want to add this Exa server. Now this is something that allows you to search the web to actually get the relevant information and up-to-date data. So we just connected it by pressing this button here. You can see we don't actually need any authentication,
and now it has these two tools where we'll be able to go inside of our sandbox and do the search. Next, we want skills. Now for skills, of course, we can create our own skill or import something from GitHub, but for now, I'm just going to have the WebArtifactsBuilder skill enabled,
so to build self-contained web artifacts. So this way we can take our research report and we can actually have it as an interactive dashboard. So let's enable that like so and now we have this skill provided. And then lastly we have sandbox providers and we are able just to use our own local
sandbox. What I mean by that is that by default if we have a look here, let me open up the terminal, you'll see that TrueForge actually can create its own sandbox. So you can see it says local sandbox fallback is available. So what it will do is kind of just create its own isolated shell within the Linux operating
system or WSL, so it's not running in our main computer. However, if we wanted to deploy this and run it more at scale and have like tons of agents running at the same time, we could connect to Daytona, which is a remote sandbox provider, which is going to allow you to have this spot up
in the cloud and just work a lot more reliably. Anyways, for now, I'm not going to configure Daytona because I'm just using the sandbox that's on my own machine, which is fine. And what we'll do now is we'll go into a new chat, and I'm going to start actually crafting an AI agent here and
then just pressing the Save Agent button so it will be kind of available reliably afterwards. So for the model, we'll stick with GREN 3.6, however we could change that. And for the tools what I want to do now is enable all of these tools and skills So I going to enable the Web Artifacts Builder the ExaSearch and then for capabilities we can have all of these as well Generative UI Dynamic Sub Ask Clarifying Questions all of that is fine
And then we'll give it the exact same prompt that we gave it before. Do research on all of the top local AI models that are available right now. Use Web Search, find all of the stats, the metrics, the parameters, how fast they are, and which ones are the highest intelligence.
I want you to talk about open weights and open source models specifically. Then it gave me an interactive web dashboard so that I'm able to view all of this. So now, if we press Enter, we should get a significantly better response because we've added these tools.
Let's see what we get. You can see now that it's able to actually list the tools. You can see we have the MCP server EXA available. It's going to start spinning up different subagents, getting tool information, reasoning through what we need to do. And this is the observability that I'm talking about,
where we're able to view all of the logs of exactly what's going on. And then same thing if we go back to the terminal, of course, you're going to see a trace of everything that's running here. So let's wait for this to finish, and I'll be right back. Okay, perfect. So you can see it went through all of the steps here.
We had multiple sub-agents that were spun up. We went through all of the different tools, and then it created a research summary for us with the updated models 2026 by actually doing the web search. Now we have DeepSeek R1, Maymax M3, GLM, you know, MISROW, right,
all of the current models, at least my understanding, with the top parameters. Now what I wanted to do was create a web dashboard, so I did make one. However, the button wasn't working, so I just said, hey, can you just make sure you're using the skill to render the dashboard?
So let's see what that looks like. Okay, perfect. So now it's building the full dashboard for us automatically, so it's actually viewable inside of this chat interface. And you can see it's kind of building it live time, and we can click through the different components here, benchmarks, all of that.
And that's exactly what I was looking for. And we don't need to wait for all of this to finish. Point is, what I can do now is I can save all of this as an agent, and then I'm actually going to be able to use this reusably in the future with all of the skills, context, and everything set up.
So what I'm going to do here is just go save agent for the agent. Let's call this, I don't know, researcher or something like that. We can give it a description if we want. We can choose the model, connectors, skills, tools, and then save this.
And if we go to our agents library now, you can see we have the researcher agent. And now that this is here, we can use the agent directly with all the configuration done, or we can actually trigger this from something like an API or another tool,
which I'm going to quickly show you. Okay, so I've just got a super simple JavaScript file here. I'm just using the TrueForge SDK, which will allow me to actually connect to the harness array that's running on my machine.
So you'll see it's just very simple code. All I'm doing is using my local host port, saying, okay, I want to access the researcher agent, telling it to do some research on tech with Tim. And now if we go here, let's just clear and go npm start.
And what it should do now is actually trigger this to run straight, create a new session, and then start researching tech with Tim by calling the different tools and giving us the response. Point being, you can use this from different ways. You don't have to use it in the web UI.
We can use it now in our code. We can use it like a llama, right? You get the idea. And then if I want, I can actually deploy this even for like my team, for example, where I have Docker Compose that I'm using or Kubernetes with something like Okta SSO support.
Anyways, I'm going to link the docs below in case you want to do that, but it's very, very cool. All right, so you can see it's spinning up a sub-agent here. I'm not going to wait for all of this to finish. What I quickly want to do is just cover a little bit of the cost angle here, because they've actually released some reports on how their harness fares compared to something like Clogged Managed Agents.
And the quick thing here is that you can see that it actually matches the exact same accuracy as using someone like the Opus model. However, it can be up to 75% cheaper based on the optimizations that this specific harness has.
So you can see TrueForge ran 30% cheaper than Clogged Managed Agents on the same model, and 75% cheaper by routing to an open model. So if you scroll down here, for the exact same task in Quad Manus Asians with Opus 4.8, it was 10 million tokens per run.
When we used Opus 4.8 in Trueforge, again, same model, it was only 3.8 million. And then when using it with GLM 5.2, we had 3.7 million tokens, but also significantly cheaper price due to the price per token, right, in terms of input-output tokens.
Point being, the harness matters a lot, and sometimes actually even more than the specific model that you're using, so make sure you use the right one. So that's pretty much all that I want to cover. Trueforge, as I showed you, is completely open source. It just launched. So if you found this useful, the single best way that you can support these types of projects is to go and actually star the repo on GitHub.
I'm going to leave the link below in the description as well as in the pinned comment, and it really helps out all of these companies that are doing this completely for free. You know, not charging you anything, you get the idea. Anyways, that's going to wrap up this video. Next time someone says the word agent, you'll actually know what that means and the difference between a harness and a model.
If you enjoyed this type of content, leave a like, subscribe, and I will see you in the next one.
