TubeSum

This NEW FREE Voice API is CRAZY! (Build Voice Agents in Seconds)

0h 12m video Published Jul 28, 2026 Transcribed Aug 10, 2026 R Robert Benjamin

AI Summary

This video showcases Fish Audio's S2.1 Pro, a free text-to-speech API that outperforms competitors like 11 Labs, OpenAI, and Gemini in blind tests. The creator demonstrates how to build a custom AI podcast generator using Fish Audio and Claude Code, highlighting features like instant voice cloning, emotion control, and multi-language support.

[00:01]
Fish Audio S2.1 Pro is free and top-ranked

Fish Audio's S2.1 Pro model is free with no usage caps and beats 11 Labs, OpenAI, and Gemini in blind tests. The same model is used on free and paid tiers.

[00:42]
Supports 83 languages with no extra cost

Fish Audio covers 83 languages without additional charges, unlike other tools. It has a fair usage policy instead of hard caps.

[01:08]
Multiple products: TTS, voice cloning, sound effects, and more

Fish Audio offers text-to-speech, instant voice cloning (10 seconds of audio), professional voice cloning (1-2 hours), voice design, sound effects, audio separation, voice changer, and speech-to

[00:01] went free with no usage caps and it's beating 11 labs open AI and Gemini in blind tests. It's called Fish Audio and their brand new S2.1 Pro model is the same model on the free tier and on the paid tier which makes this a no-brainer

[00:16] going to know exactly what it is, how you could use it and I'm going to show build your own voice agent. In fact, my team uses this agent all the time in order to create viral content for my brand. Now before I actually walk you

[00:30] through the exact voice agent that I've built, how I actually created it and how you could build something just like it, I wanted to walk you through Fish Audio before, this tool has been ranked number

[00:42] one when it comes to blind studies against tools like 11 labs, open AI and Gemini's voice tools. It covers 83 different languages and you don't have to pay extra for certain languages like you do with other tools and other

[00:55] models. And they don't have a hard usage cap, they just have a fair usage policy. What's even crazier, I said there's a 90 millisecond latency which is incredibly we can see that they have several different products. From home, we can

[01:08] allows us to convert any speech that we want into natural sounding speech in seconds. For example, we can come over here, we can see that we can choose Fish Audio S2.1 Pro for free. We can come over here and actually choose what voice

[01:22] actually click into this to add another voice, we can see that there are tons of voices, there are samples here and we can go through and actually explore related models here which is pretty cool or if we wanted to add another voice, we

[01:35] could simply just click on this, we could go through and we can explore all there are literally tons of them. Bunch of different accents, bunch of different Larry from advertising, there's professional and energetic middle-aged

[01:49] male, smooth and confident, there are tons and tons. In addition to that, we can also actually create our own voice here which we can see right here. And we can create voice collections of voices, which is pretty cool. If we actually

[02:01] where we could go through and you can instantly clone your voice with only 10 seconds of audio. You can actually describe a voice that you want to hear in plain text and be able to look that out, or you can do a professional voice

[02:14] clone, which requires 1 to 2 hours of audio. And what's pretty incredible is if you're just looking to clone your voice quickly and you don't want to put the instant voice clone. But, if you really want to take things to the next

[02:27] voice clone. If you're going to do something like create a podcast you in just a little bit, or something along the lines of that, where you might be using your voice inside of your content. Or even this voice design is

[02:40] this tool is that when you actually design your own voice, you are not going to be held accountable when it comes to using an AI voice on YouTube. And what I mean by that is that you likely won't get banned for that. A lot of people

[02:54] because what's happening is they're all using the same voice. But, if you custom decrease the odds that that actually happens for you. Now, if we come back handful of other things that we could do right here. Like, we could create sound

[03:09] could actually choose them, or we could actually create our own sound effect if we wanted to by simply describing exactly what it is that we want to over here, we could see that we could also do audio separation. This allows us

[03:24] drums, bass, and other stems from any audio if that's something that you might need to do from inside of your content. And they even have this voice changer right here, which allows you to actually upload an audio source. We could select

[03:38] actually have a conversation and be able to change so you could transform your voice into any character with AI voice conversion. In addition to that, they also have speech-to-text. So, if you want to be able to upload a long-form

[03:50] audio, you can actually go through and see what the transcript is. What's even you to preview the credit cost before you actually do this, which most other tools don't allow you to do. Now, all of this is what's possible from directly

[04:05] can see that I've actually gone through and created multiple different projects could do is if we come over here into developer mode, this is where you're actually able to get your API key and you're able to go through and you are

[04:18] able to take this into other tools like Claude or ChatGPT or something that you build custom and be able to access this through that by putting it into your own which is exactly what we're about to go over. Now, before I walk you through

[04:32] order to build a voice agent inside of Claude, I wanted to remind you that you could go to the pinned comment below and grab your free API key today so that you could start also building with Fish Audio. On top of that, Fish Audio is

[04:44] currently celebrating their first anniversary. As a result, from July 24th anniversary. As a result, from July 24th to August 24th, creators 50% off all creator plans. And you're also going to get free access to S2.1 Pro from July

[04:57] 24th to August 7th. Okay, so in order to actually do this, here's what I did. I personally to do this, but you could do this with any tool that you want and API access right here. So we could see that I said to this, "I want to build an

[05:12] AI podcast generator. It should allow me to type in an article or a topic and choose the voice that I want from a list of voices and choose the emotion that I should get done on that topic and then an AI podcast should be generated from

[05:25] it. To generate the podcast, I want to use Fish Audio's S2.1 Pro API." Now, through and this coded this up for me. And it actually generated this directly inside of here that I can now go through and I can actually look at. So we could

[05:40] this works. So we could go from a single idea to a finished podcast episode. We simply just come over here, we drop in a topic, or we paste an article, and then actually researches the topic on the live web, writes a host-ready script,

[05:55] and records it in the voice and the mood that we actually choose. We can see uses Claude Opus 4.8 in order to actually do the research, and this is going to use Fish Audio's S2.1 Pro voice right here. So, we can see that we can

[06:08] simply just drop in the source material, we can choose a voice right here, we can like, how long it's going to be, and this will actually go through and produce the episode. And this also goes through and actually saves all of the

[06:20] before in the past, because we have created a few. And part of what really makes this powerful right here is that if we actually come back over here, we can see that we can actually come in here and we can create a voice. So, I'm

[06:33] actually going to clone my voice by clicking on instant voice clone right here. And now we can go through and we can actually record what is going to be used here, and we can see that this can actually go up to 210 seconds right

[06:46] minutes. So, I'm going to go through, I'm going to do this quick. I'm going to We're then going to add that into the app, and I'm going to show you what the went through, I recorded 30 seconds right here. We can see that it's now

[06:59] going through, it's analyzing this, and this is actually going to allow me to create my voiceover from my voice in just 30 seconds. Now, while it's doing thing that again makes this incredibly powerful, because if we come over to

[07:12] text-to-speech right here, we can see that we can actually add in different emotions here, which is absolutely incredible. You can also control all of say that I came over here and I wrote out some random stuff, like the TikTok

[07:26] out some random stuff, like the TikTok algorithm just released a massive update. This changes everything, and you shouldn't skip any part of this video. Now, what we can actually do is come

[07:39] automatically going to tag this up with changes the voice output that you get, which is absolutely incredible. So, we changes with excited, changes with emphasis, and changes with emphasis. So,

[07:55] different voice than if the tags weren't used, which again is what makes this tool so powerful and why it's so good for building your custom voice agents. could see that this has actually been created. What I'm actually going to do

[08:09] is I'm going to name this Aaron Knows AI. It's kind of like just his name that it has a description, has all this stuff. I'm going to keep this as public confirm that I'm actually able to use this right here. We're going to click on

[08:22] create. And now what we should actually be able to do is from inside of our voice model, we should be able to actually see and use this voice right of what this voice sounds like, check this out. Our planet's diverse

[08:35] ecosystems harbor countless wonders from the smallest microorganisms to Now, literally in 30 seconds taking me a voice that sounds exactly like me, which out. What we're actually going to do now is we're going to come back over here

[08:51] source material from something that I want to create. So, let's say that I about the TikTok algorithm update for 2026. What I could do is I could go through and I could grab an article about this. For example, if we come over

[09:05] here, Hootsuite has a blog here and I'm actually going to come through. I'm going to grab the key takeaways right here and I'm going to grab a lot of this information and we're actually going to paste this inside of here. So, we're

[09:18] here and what this is. So, I'm actually going to copy this. We're going to come this as the source material. We're going to come down here. I typed in Aaron actually created. Again, we could test it by looking right here. Our planet's

[09:33] diverse ecosystems harbor countless one We we see that is my voice right here. terms of what the delivery is going to look like, I want it to be energetic. short right here, but we could make this long, we could make this standard, we

[09:49] going to click on produce episode right here and then this is going to go through and actually create this. And now this is actually going off and doing We could see that this is going through, this is actually researching our topic

[10:01] live on the web. It's going to take 1 to 3 minutes. It's then going to write this this entire podcast episode for us and this is exactly how you could go through and you could create your own podcast generator. You could use this for

[10:14] basically anything that you want because now you're able to turn your real voice, somebody else's voice, or a voice that you designed into whatever kind of content that you want to create literally in minutes just by using

[10:27] literally in minutes just by using Claude Code and Fish Audio S2.1 Pro, which again you could get started with today literally for free. And so here it just a few minutes. We could actually come over here and we could read exactly

[10:40] tags that are in here, which is incredible. We could come over here and the playback speed and I want to show you exactly what this looks and sounds Rob TikTok podcast right here. We're going to click on save. Now let's check

[10:55] this out. Ever wondered why TikTok knows you better than your best friend? Here's the wild part. TikTok doesn't run on a social graph, it runs on an interest graph. So think about that. I literally just typed in a topic, something I got

[11:09] here, which literally took me 30 seconds to make and now I have this automatic podcast generator. What's even crazier is I can actually go through and I could get this to create podcast about every topic that trends, everything that comes

[11:22] without knowing how to code, without knowing how to do anything because I'm knowing how to do anything because I'm using Fish Audio S2.1 Pro. And again, model is. I've used 11 labs and literally uploaded two to three hours of

[11:36] content and it's not as good as their 30-second voice model. So, what are you create a voice agent that will help you create better content without getting stuck paying for cap trials or watching your credits disappear before you've

[11:50] that's over now because you could go to the pin comment below and get started with Fish Audio's S2.1 Pro for free fact that Fish Audio is currently celebrating its first anniversary and

[12:04] celebrating its first anniversary and from July 24th to August 24th, creators can get 50% off all creator plans. In addition to that, creators can also enjoy free access to S2.1 Pro from July 24th to August 7th. So, go to that pin

[12:18] comment below and get started with it today for free.

More from Robert Benjamin

View all

⚡ Saved you 0h 12m reading this? Transcribe any YouTube video for free — no signup needed.