TubeSum ← Transcribe a video

This New Chinese AI Agent Is Insane! (Free + Open Source)

0h 08m video Published Jul 28, 2026 Transcribed Jul 28, 2026 Julian Goldie SEO Julian Goldie SEO
Intermediate 4 min read For: Business owners, AI enthusiasts, and developers interested in efficient AI agents for automation.
AI Trust Score 65/100
⚠️ Average / Some Fluff

"Title promises an insane new agent, and the model is genuinely impressive, but the 'insane' hype slightly oversells a practical efficiency gain."

AI Summary

Ling 3.0 Flash, a new AI agent model from Inclusion AI (Ant Group), boasts 124 billion total parameters but activates only 5.1 billion per response, achieving efficiency and performance that beats its predecessor on 11 of 12 benchmarks. Designed for multi-step agent tasks, it features a 256K token context window, a mixture of experts architecture, and demonstrated capabilities like building 3D cities, coordinating agent teams, and managing schedules. The model is free to test via OpenRouter and Versel AI Gateway until August 3rd, offering a practical glimpse into efficient AI agents for business automation.

[00:02]
Model Introduction

Ling 3.0 Flash is a free AI agent model with 124 billion parameters, but only 5.1 billion are activated per response, making it efficient.

[00:17]
Performance vs Bigger Model

Despite using only 1/12 of the parameters of its sibling model Ling 2.6, it beats it on 11 out of 12 benchmarks.

[00:45]
Mixture of Experts Analogy

The model activates only needed parameters for each task, like a business calling in only the relevant staff.

[01:22]
Company Background

Inclusion AI is the research group inside Ant Group (Alipay). Ling is their agent-focused model line.

[01:36]
Key Specs and Improvements

Total parameters increased from 104B to 124B, but active parameters dropped from 7.4B to 5.1B, enabling more knowledge with less compute.

[02:05]
Context Window

Native 256,000 token context window, scalable to 1 million tokens, allowing handling of large documents and conversation histories.

[02:18]
Attention System

Uses a mix of two attention systems in a 5:1 ratio for long-term memory without slowing down, designed for long multi-step tasks.

[02:59]
Demo Examples

Demonstrated building 3D cities (BlenderMCP), coordinating five-agent research teams, generating design systems, building browser synthesizers, and managing schedules.

[04:45]
Industry Shift

AI industry is moving from chasing size to efficiency; Google Gemini 3.6 Flash and Ling 3.0 Flash exemplify this trend.

[05:52]
Free Access

Available for free on OpenRouter and Versel AI Gateway until August 3rd. No setup required.

[06:35]
Advice for Businesses

Test with real tasks now; agents are capable of repetitive work like scheduling, content repurposing, and proposals.

Ling 3.0 Flash demonstrates that efficient AI agents are now accessible for practical business automation, with a free trial window available to test real-world tasks. The model's mixture of experts architecture signals an industry-wide shift toward smarter, more cost-effective AI.

Mentioned in this Video

Study Flashcards (9)

What is the total parameter count of Ling 3.0 Flash?

easy Click to reveal answer

124 billion

00:02

How many parameters does Ling 3.0 Flash activate per response?

easy Click to reveal answer

5.1 billion

00:02

What is the native context window size of Ling 3.0 Flash?

easy Click to reveal answer

256,000 tokens

02:05

Which company developed Ling 3.0 Flash?

easy Click to reveal answer

Inclusion AI, the AI research group inside Ant Group (Alipay).

01:22

What architectural approach does Ling 3.0 Flash use to achieve efficiency?

medium Click to reveal answer

Mixture of experts, activating only a fraction of parameters per task.

00:45

How does Lin 3.0 Flash's active parameter count compare to its predecessor Ling 2.6 Flash?

medium Click to reveal answer

It dropped from 7.4 billion to 5.1 billion active parameters.

01:36

What is the ratio of the two attention systems used in Ling 3.0 Flash?

hard Click to reveal answer

5:1

02:18

Until what date is Ling 3.0 Flash free to use?

easy Click to reveal answer

August 3rd.

05:52

How many benchmarks did Ling 3.0 Flash beat its bigger sibling model on?

medium Click to reveal answer

11 out of 12 benchmarks.

00:17

💡 Key Takeaways

💡

Efficiency Breakthrough

Shows how mixture of experts enables massive models to run with minimal compute, a key trend in AI.

00:02
📊

Outperforming Larger Models

Demonstrates that efficiency doesn't sacrifice performance; beats Ling 2.6 on almost all benchmarks.

00:17
🔧

Real Agent Capabilities

Showcases practical automation tasks like scheduling and document editing, not just toy demos.

02:45
⚖️

Industry Shift

Highlights the move from model size to efficiency, making AI cheaper and more accessible.

04:45

[00:02] yes, it's free right now. It's called Ling {dash} 3.0 {dash} flash and it might be one of the most useful AI agent models to come out this year. Here's the number that matters. This model has 124 billion parameters, but it only

[00:17] activates about 5.1 billion of them every time it generates a response. That's roughly 1/12 of what its own bigger sibling model uses and despite that it beats that bigger model called Ling 2.6 on 11 out of 12 benchmarks. Let

[00:32] me explain what that actually means because the number alone doesn't tell the full story. Parameters are like little decision points inside an AI more knowledge packed in, but using every single one every single time costs

[00:45] every single one every single time costs speed. Ling {dash} 3.0 {dash} flash only wakes up the parameters it actually needs for the task in front of it. Think of it like a business with a 100 staff on the books, but only calling in the

[00:57] five people who actually handle that specific job. Everyone else stays home. That's why it runs fast even though the full model is massive. Hey, if we haven't met already, I'm the digital avatar of Julian Goldie, CEO of SEO

[01:09] agency Goldie Agency. Whilst he's helping clients get more leads and customers, I'm here to help you get the latest AI updates. This model comes from a company called Inclusion AI, which is the AI research group inside Ant Group.

[01:22] Alipay, one of the biggest payment platforms in China. Ling is their agent focused model line and this flash version is built to be their fast, efficient tier sitting below their much bigger flagship models. Now, here's what

[01:36] makes this one different from the last version called Ling {dash} 2.6 {dash} flash. The total parameter count actually went up from 104 billion to 124 billion, but the active parameter count went down from 7.4 billion to 5.1

[01:51] billion. More room to store knowledge, less power needed to use it. That's a hard balance to hit and it's exactly the direction the whole industry is moving in right now. It also holds a 256,000 token context window natively. And

[02:05] Inclusion AI says it's built to scale up toward 1 million tokens. In plain terms, that means it can hold a huge amount of information in mind at once. A long document, a full conversation history, a big set of instructions without losing

[02:18] track of the details. Under the hood, it uses a mix of two different attention systems stacked in a 5:1 ratio to help it remember things over long stretches of text without slowing down. You don't need to memorize that term. Just know

[02:31] this, it's specifically engineered for long multi-step tasks, not just quick one-off questions. That matters because most real work isn't one question and one answer. It's a chain of steps where each one depends on the last. And this

[02:45] questions, it's built to act as an agent. That means it can take a goal, break it into steps, use tools along the way, and actually finish the job. Not just describe how you do it. Inclusion AI showed this off directly. They had

[02:59] the model build and render a full 3D city using a tool called BlenderMCP. They had a coordinate a five-agent research team working together on a and check the numbers in Excel spreadsheets. They had it generate a

[03:13] full design system without using any outside images. They had it build an interactive synthesizer that runs right in a browser. And they even had it read messages, check calendar availability, and update a schedule all on its own.

[03:25] That last one is the part worth sitting with, reading messages, checking a touching it yourself. That's not a demo trick. That's the exact kind of task that eats up a business owner's morning. Or take the Word and Excel example.

[03:39] Editing a proposal document and checking the numbers behind it at the same time different apps and doing the math by hand. That's the kind of grinding repetitive work that that takes hours out of a week. For example, imagine

[03:53] using this to take a single announcement about the AI Profit Boardroom and turn X, one for Instagram, one for LinkedIn, each written in the right tone for that platform, all from one input. Right now, this is exactly the kind of tool we're

[04:06] building playbooks around inside the AI Profit Boardroom. If you want to learn how to save time and automate your business with AI agents like Ling dash 3.0 dash flash, this week's tutorials

[04:18] handle repetitive admin work, the kind of stuff that normally eats your whole morning, like turning one announcement into posts for every platform you use. You get four live coaching calls a week where you can bring your actual setup

[04:31] built specifically for agent workflows like this one, and a member map so you can connect with people near you who are already testing these tools for their description, or head to aiprofitboardroom.com.

[04:45] Now, let's talk about why this release matters beyond just this one model. For a couple of years, the AI industry mostly chased size. Bigger model, more results. That was the whole game, but

[04:58] that approach gets expensive to run at scale, and agents make that problem worse because an agent isn't answering one question. It's thinking, checking its work, using tools, retrying when something fails, and doing that across

[05:10] dozens or hundreds of steps in a single task. Every one of those steps costs compute. If the model behind it is heavy, that adds up fast. So, the how do we make the model bigger, companies started asking how do we get

[05:23] the same intelligence for less effort. Google did this with Gemini 3.6 flash, which uses fewer tokens per task than its last version. Ling dash 3.0 dash flash is doing the same thing from a completely different angle using the

[05:38] mixture of experts approach where only a fraction of the model activates at once. not technical at all. You can access Lingdash 3.0-flash right now for free through August 3rd on

[05:52] platforms like OpenRouter and Versel's AI Gateway. No setup, no download. You testing it. So, here's what to actually do with this. If you're not technical, go try it for free before the trial window closes on August 3rd. Give it one

[06:07] real task, not a toy example, an actual thing you'd normally do yourself, like calendar. If you've got someone technical on your team, have them test it against whatever agent tools you're already using and see if it holds up on

[06:21] tasks. If you're running a business, don't wait for the perfect agent model to show up. The ones available right now are already capable of real repetitive work. The businesses moving fastest are testing

[06:35] this stuff the week it drops, not months later. And if you're managing a team, start thinking about which weekly tasks are actually just repeated steps in disguise. Updating a calendar, turning one post into five, checking numbers

[06:48] against a document, those are exactly the tasks agents like this one are built to take off your plate. Look, I get that this space moves fast. New model, new company, new number to remember every single week. But, that's exactly why

[07:00] staying close to it matters more now, not less. This model just showed that a business can automate real admin work, content, proposals, scheduling using an AI agent that's free to test right now, and that's exactly what we're building

[07:13] around inside the AI Profit Boardroom this week. Our tutorials walk you through setting up an agent like Lingdash 3.0-flash covered, turning one piece of content into posts for every platform, checking

[07:26] your calendar, drafting the first version of a proposal, four live bring your setup and get real feedback, a prompt library built specifically for agent workflows like this one, and a member map so you can connect with

[07:39] people near you already testing these tools for their own business. Links in the comments and description or go to aiprofitboardroom.com. And if you want the full process behind everything I just walked through, plus

[07:51] over 100 other AI use cases just like it, join the AI Success Lab. It's free. You'll get all the notes from this video and access to a community of 87,000 people who are already deep into this stuff, sharing what's working every

[08:04] single day. Links in the comments and description.

More from Julian Goldie SEO

View all

⚡ Saved you 0h 08m reading this? Transcribe any YouTube video for free — no signup needed.