-
Cryptocurrencies
-
Exchanges
-
Media
All languages
Cryptocurrencies
Exchanges
Media
Share
Everyone who has actually used AI in the past year will have the same feeling - AI is getting more and more expensive. In the past, a monthly subscription of US$20 might not be enough to spend, but since agents and vibe coding became popular, tokens were burning like water. A coding agent hung around all afternoon, and the bills piled up. So everyone gradually learned to be careful about whether this task is worth running, and whether this code should be rewritten by AI. As soon as many ideas came up, they were put back by the sentence "How many tokens will this cost?"
AI is supposed to allow everyone to create and use it to their heart's content, but instead it has become a matter of billing by the meter and saving as much as possible.
Now there is a company that allows you to use all three AI models of text, pictures, and videos without spending a penny. And it’s not a seven-day trial, nor is it giving you a credit that can only be used up, but it is enough to control it. Is this the “Cyber Bodhisattva” of the AI era?
On June 1st, a start-up team called Agnes AI made API tokens for its three models of text, pictures, and videos free of charge. As soon as the news came out, more than a dozen groups were crowded within a few days. In just the first week after it became free, the call volume of Agnes-2.0-Flash soared to more than 1 trillion (1T) Token; Agnes-Image-2.1-flash generated more than 2 million pictures; Agnes-Video-2.0 even produced more than 2 million seconds of video. The first ones to come in, almost all the first ones to come in were geeks who came overnight to "experience" it.
But soon, the style of the group changed.
Some people use it to generate videos that are several minutes long, some people use it to match workflows to create a complete set of materials, and some people cut clips of their two daughters growing up into short videos with AI narration. At the past price, he probably wouldn't be willing to try this kind of video of a few minutes. This is actually the most interesting thing about "free". What it really unlocks is not the money you save, but the ideas that you didn't even try because it was too expensive and you didn't dare to try it.
What is even rarer is that when most companies only focus on a single model form, Agnes chooses to make text, pictures, and videos together, and all of them are free.
The excitement continues, this week Agnes AI will update the 1M ultra-long context and 4K ultra-clear pictures model.
Of course, the question also arises: Does free mean the model is not good enough? How do you control the cost so that it can be used by so many people? If no money is collected, how can the team survive?
And most importantly, What is the purpose of the Agnes AI team that did this?
With these questions in mind, Geek Park had a chat with Bruce Yang , the founder of Agnes AI. The following is the summary of the conversation:
Models with higher prices are better, while models with lower prices have poorer performance. This is a big misunderstanding. DeepSeek is also very cheap, and it is actually as good as many more expensive models in many indicators.
What is really unlocked by free Token is not the little money saved, but the ideas that you didn’t try at all because it was too expensive and you didn’t dare to try it. User potential should not be limited by cost.
Because of the Harness constraint, the gap between models is actually getting smaller. The role of Harness is, first, to weaken the gap between models, and second, to make the upgrade and optimization of models more directional.
We want to be the first to fly the free flag now, be the first to get to the table, and be the first to become a significant player.
Ten years ago, if you didn’t know Chinese or English, you might be illiterate; 10 years later, if you don’t understand AI, you might be illiterate. In fact, it’s not that they are afraid of AI, but that people who don’t understand AI are afraid of people who do, and feel that they will be replaced at any time.

Agnes AI |Source: official website
Geek Park: Let me introduce myself and the Agnes AI team.
Bruce Yang: Let me first talk about my own experience. I went abroad at the age of 15 and went to high school at Raffles Institution in Singapore. After that, I was admitted to the University of California, Berkeley, majoring in computer science and mathematics. Fortunately, I studied under two Turing Award winners Richard Karp and David Patterson. Ion Stoica, who taught us the operating system at the time, is now the founder of Databricks. Later, I worked in Silicon Valley, worked for Microsoft and LinkedIn, started a business in Silicon Valley, returned to China, and am now in Singapore. There was an opportunity in the middle. During the country's lockdown during the epidemic, I returned to Singapore to study for a Ph.D. in AI at the School of Computing at the National University of Singapore. This experience gave me a lot of inspiration, and it was also a very important route for the founding of Agnes.
Agnes will really start around the end of 2024 or the beginning of 2025. It is a very young company. We have been working on models from the beginning, but last year we were working more on applications. Because the models were not that good yet, we first made our own harness, which is now the agent, and then slowly optimized the capabilities on top of the agent. In the beginning, we relied more on the external API of "Yusanjia" to realize capabilities, but the cost has always been high, especially after the scale reached more than 10 million users, it was no longer affordable, so we accelerated the promotion of self-developed models and did so-called domestic substitution. By the end of last year, the story was more about products and domestic substitution.
By the beginning of this year, we found that the model was doing pretty well. Compared with some closed-source models, it still had advantages in some places. So we made a bold effort and started to open up the model API this year, scaling it up from a small scale to a full mode. Now we just make a big move and make all modes free. It has only been three days since it was announced on June 1st and it has been officially launched now. There are already more than a dozen groups, with hundreds of people in each group, all of which are geek users. Our token consumption exceeded 100 billion yesterday. It is not bad to reach this number in three days, and it may increase three or four times by the weekend. So far it seems to be within expectations.
Geek Park: Not only in China, but also all over the world will be very excited when they hear it for free. But everyone will also question, is it because the stuff is not that powerful that it is free? What is the current level of your three models?
Bruce Yang:I think this is a misunderstanding, and it is not only for free models, but also for low-priced and cost-effective models. I always feel that the more expensive models are better, and the cheaper ones are lowered so low because of poor performance. But you see, DeepSeek is also very cheap. It is actually as good as many more expensive models in many indicators.
While our model is currently free, it does not represent any compromise in performance. Judging from the current results, our text model is among the top ten AI labs in the world in some agentic scenarios, such as PinchBench and ClawEval; the image and video models are also among the top ten AI labs in the world on Artificial Analysis, the world's most authoritative blind evaluation list.
The model is still being optimized and will be updated this month, and possibly every month thereafter. Our requirement for ourselves is that the leading SOTA model may not necessarily reach the same strength immediately, but it must keep up quickly and stay within the same generation. For example, when a new version of it comes out, we can achieve the capabilities of the previous version. It is not easy to do this, and since it is free, I believe it will be favored by many users.
Geek Park: You are very confident in your model, why not show and explain to everyone the demo made by the Agnes model.
Bruce Yang: There have been many reviews of our model online. We took a look and found that 95% of them were not provided by us. On the first day after it was announced that it would be free, there were a lot of spontaneous user promotions. The reviews were quite pertinent and pointed out some of our problems, but overall everyone recognized our capabilities.
For example, this kind of particle effect was an important indicator for testing text model capabilities when Gemini first came out. Another friend made an operating system using text models, and there was also a mini-game about flying in it.
In addition to text, pictures and videos are also OK, especially pictures. We have done a good job in optimizing some content with high information density. Of course, there is still some distance between the image model compared to Nano Banana and GPT, and some high-density text details have not been fully optimized, but overall it should be considered relatively high among domestic models.
In terms of video, we support the simultaneous production of audio and video. Characters can speak in the video. Both Chinese and English are supported. Some small details still need to be optimized. We will probably launch the next version of the video model in the second half of this month. The goal is to be close to the stage of HappyHorse, but there is still a gap with Seedance. But generally speaking, as a commercial model, being free does not mean that it has no commercial value. We have achieved the capabilities of many closed-source models and can also release a lot of commercial potential.
Geek Park: Are the tasks just shown completed end-to-end by a single model, or do they involve the collaboration of multiple agents?
Bruce Yang:The API we provide has only three models, text, image, and video. At present, these three APIs have not been unified together. We want to release them together next week because many people will be confused when configuring them. Many harnesses do not support direct uploading or downloading of images and videos and need to be loaded as skills. So now it's three different models. The content you see is basically completed on the basis of harness. The harness can be our own Agnes harness, or it can be Codex, OpenClaw, or Claude Code. After connecting our single model, we can achieve capabilities.
Currently we do not use multiple text models, or multiple image or video models to support harness work; however, during the execution process of the harness, due to its own understanding, needs and dependencies, multiple agents may be dispatched at a certain time to implement it. We can support this.
Geek Park: How much difference is there in the computing power cost to complete these tasks compared with the current popular models and tools?
Bruce Yang: Let’s talk about our quotation first. Although it is free now, there was a quotation before it became free, and there is still a token plan. In terms of text models, generally only the output token is linked to cost, and the input token is basically zero cost to the model company.
The token we input is 0.15 US dollars per million, which is about 1/100 of GPT and Anthropic, and about half cheaper than the flash version of DeepSeek. We still make some profits. Pictures are $3 per 1,000 pictures, which is $0.003 a picture, which is an exaggeration. The actual cost of video is about US$0.3 per minute, which is about one cent of RMB per second. According to this cost, it is about 1/100 of the price quoted by the leading model in the market.
This is the original quote. Now it’s free, so everyone can use it for free. We only slightly limit the QPS (queries/requests per second) and RPM (requests per minute), but we still give you a lot, and you can request 20 times per minute. Normal individual developers have not yet encountered situations where the amount of resources is insufficient.
Geek Park: Free makes people worry about whether the team can sustain it? There are very few teams that do all three models, especially since Agnes is not a giant company. Why do all three types of models need to be done together?
Bruce Yang: There is indeed pressure. Our scientific research team already has more than 100 people. Currently, there are not many Lab companies that are ranked in the top ten of the global model list in terms of text, pictures, and videos. Overseas, they are Google and OpenAI, and domestically, they may be Alibaba and Byte. There are not many companies that are doing the other three.
We didn’t think much of it at first. Because our own harness products originally support text, pictures, and videos, and are quite balanced in terms of usage, the first step is to replace them with domestic products. In this process, we discovered that there are actually synergies between the three models.
When Nano Banana came out, they mentioned a point. The reason why Nano Banana’s instruction following ability is very strong is because it uses the then flagship model Gemini 2.5 Pro for visual content analysis, and its ability to reverse prompt words is very strong. The same goes for the video model. Anyone who has actually trained it will know that the first premise is that the text and picture model must be strong. The video model also requires a large amount of data, many of which come from film and television slices. After slicing, this part of the video must be well described with text. These descriptions can be used for reverse training. This process also relies heavily on text models. Therefore, the three models actually have certain dependencies during training. Including some new routes now, such as image models that have begun to use AR, which requires the ability to combine understanding and generation.
So generally speaking there are two reasons. First, start from the real usage needs. For many one-man companies and small workshops, it is quite challenging for him to configure three interfaces from different companies; if they can be combined together to make an Omni-mode API, the cost and threshold of use can be better reduced. Second, there is synergy between training. The better the multi-modal text understanding model is, the more it can support the generation of pictures and videos. The two complement each other; a lot of new data will be generated in multi-modal scenes, which is very helpful for us to synthesize data and further train. In particular, the picture and video models need text harness to help them with prompt word enhancement.
Only by integrating the three models together and creating a scenario for users to continuously explore can we understand the direction of the next model upgrade.
From another perspective, the difference between training three models at the same time and training only one depends on the vision and cognition of each company. The vision of Anthropic and OpenAI is to use the strongest text model to achieve qualitative changes in capabilities and realize AGI as soon as possible. But our understanding of AGI is a little different. We hope that our AGI will be used by the widest range of users and in the largest scenarios, and it will be a broader AGI. Under this route, we may not have the strongest model in every model, but we must stay at the forefront, maybe the top ten, and not fall behind a generation. At the same time, we hope that the model capabilities will complement each other and make progress together. We also hope that more and more users will use our products and build an ecosystem. Let the ecosystem promote our progress, understand market demand, and understand how to lower the threshold for use.
Because the vision, starting point, and technical route are different, we will choose a route that others may not choose, but this does not mean that we have any downgrade or compromise on performance. We will still remain at the forefront of the world.
Geek Park: What’s so charming about you that you can bring together talents who do text, image generation, and video production?
Bruce Yang:Actually we have four teams. There are one team for text, one for pictures, and one for video. There are more than a dozen people in each team. There is also a team dedicated to performance optimization. How to further reduce costs. In fact, cost is not the best word. Efficiency may be better. How can we achieve some staggering numbers in both the training phase and the inference phase, such as 1% of the inference cost.
One of our core logics is that we have been doing an optimization problem with strong constraints from the first day, but our constraints are different from others. Many people's restriction is to give you enough resources to increase your capabilities; but for us on the first day, the resources themselves were not that big. That’s why we need a team that spans three vertical categories and specializes in performance optimization, whether from the GPU, Codex level, or algorithm level, using the smallest possible parameters to achieve the most perfect match between user satisfaction and performance. When we first started doing this, we actually had no idea.
As for the charisma you mentioned, we are actually latecomers, whether it is the Singapore team or the domestic team, because most model companies are not in these two regions. But when we made some signs of success, we attracted a large number of outstanding local students. Many students from NUS and NTU in Singapore, as well as domestic teams from Nanjing University, Dongda University, University of Science and Technology of China, Zhejiang University and even Tsinghua University, chose to come to our company.
The entire scientific research team now has nearly 100 people. They are all smart and excellent, working hard for a great vision. On June 1, we launched a big move to announce the capabilities and some scientific research findings we have accumulated in the past, and we will open source some new discoveries next week. The team is very motivated and wants to be not only a recipient but also a builder in the AI era. This is our corporate culture.
Geek Park: In your opinion, in what scenarios can the three types of models really enter production and commercialization, and then be able to run and make money? Are there any clear scenarios?
Bruce Yang:Still my point just now, paid and expensive models are not necessarily better. Students in the group who have tried our model said that it is as good as any paid SOTA model, even comparing it to Gemini and Claude. Of course, we know in our hearts that there is still a gap. Because of this misunderstanding, it is no longer meaningful to simply reduce prices. If you lower the price, many people think it's because your performance is not good enough, so they won't use it even if you lower it, because they would rather use Yusanjia.
The way to break this deadlock and change this stereotype is to let everyone try it boldly and find some surprises in the process. After three days of opening, there are more than a dozen groups and thousands of friends. In fact, there are far more, but only about 10% of users will scan the code to join the group. The QR code is under the API key on the official website.
From the feedback, we can realize most of the functions they use the paid advanced model. Even if there are some deficiencies, such as particularly complex instruction following and particularly long-term agentic tasks, these can be made up for and can be optimized in the next version, possibly next week, such as some of the capabilities of tool calling. So the big logic is thatwe can implement 90% of the scenarios that everyone is using now.
If we have to focus more on where, we have spent more time optimizing agentic capabilities. This is why I pay attention to PinchBench and ClawEval. The next version will further optimize coding. For example, we are currently playing SWE and upgrading coding capabilities. We hope that SWE can also become the top ten in the world. At present, there are still opportunities. In terms of text, we focus more on Agent and Coding, which are the most heavily used by users. I think the picture is quite capable. Although there is a gap with the GPT image model, it is still acceptable among domestic models. The video gap is a little bigger, and there is a gap with Seedance and HappyHorse, but whether it is free or at the original price, the price/performance ratio is absolutely OK. I can look forward to the next version this month. I hope it can be close.
To sum up the three models, even if there is still a gap with some SOTA closed-source models, we know how to shorten the distance and will always promote scientific research with the mission of being infinitely close to the closed-source model.
Geek Park: Without the hot wave of agents, the token issue might not attract so much attention. But now as soon as the token comes out, you spend it all at once.
Bruce Yang: Yes, coding has agent harnesses, so-called OpenClaw, Hermes, Codex, and Claude Code. Their architectures are actually very similar. Because of the harness constraint, the gap between models is actually getting smaller.
I went horse riding in Xinjiang some time ago just to feel the harness, so I rode a few different horses. The first horse was very obedient, but not very fast; the second horse was very fast, but not very obedient, but when the reins were in my hand, I found that the difference was not big. Those who can't run fast can go faster by kicking the stirrups; those who don't obey instructions can become obedient by pulling the reins. Therefore, the role of harness is, firstly, to weaken the gap between models, and secondly, to make the upgrade and optimization of models more directional.
What we need to do more is not to train a wild horse without a harness, but to train a horse with a harness. After putting on the harness, many directions and dimensions have actually been compressed, and the direction of progress is very clear. There is also a fast and obedient horse that I am not riding. The guide is riding it. It is a thousand-mile horse and has no share in me. It is equivalent to a SOTA model. What I have to do now is to train a less talented horse on the basis of a harness, so that it can be infinitely close to the SOTA model.
Geek Park: In order to experience harnessing, I went to experience horseback riding, which was also awesome. Claude Code is so strong not only because Anthropic’s model is great, but also because its entire harness is very well done, and there is a lot worth learning in it.
Bruce Yang: Compared with OpenClaw, I think Claude Code has two greater advantages. The first is Memory processing and compression, which is much better than OpenClaw. It has done a lot of optimization of long-term memory capabilities; the second is the optimization of KV Cache, which can reduce token usage and improve token hit cache.
The hit cache is basically zero cost to the model company. Although the user is charged, it is zero cost to the model company, and the input token is also zero cost. So many times you will see why some companies can reduce the prices of cached tokens and input tokens so low? Because everyone’s cost items are mainly in the output token and the output layer.
Geek Park: More than a dozen groups have been created since it became free on June 1st. How is the current situation? How do users use free tokens?
Bruce Yang: They helped us find a lot of problems that we couldn’t find when making our own products, including some stress testing methods, usage scenarios, settings for adapting to different harnesses, error logs, etc. We used to have a testing team of seven or eight people, and now many active users in the group have helped us find problems that they couldn't catch and gave us very good suggestions. Many people are development engineers and operation and maintenance engineers, and they also pointed out some stuck points of our gateway.
Second, what moved me even more was the discovery of many scenes. It turns out that what we use the video model to do is tens of seconds, 5 seconds, and 10 seconds, because the model only supports 10 seconds. But users use their own harness and skills specially written for us, and some people create ComfyUI workflow to produce videos of a few minutes, 3 to 5 minutes, and send them to the group.
I saw a user posted a short video of his two daughters growing up, and used TTS to add a very touching quote to put the video together. My first reaction was surprise. Is this done by our model? I thought it was pretty good. Many people make 5-minute videos. If they don’t use our free model, they may not be willing to try it due to cost. We are opening up a new scenario, a new right. Our company has a saying that users’ potential should not be limited by cost. We give users the right to unleash their potential.
There is another point that is quite touching. We originally tried to write an email to OpenClaw, saying that the models you access by default are all well-known models, and our rankings are also good. Can our models be included?
Geek Park: What does OpenClaw say?
Bruce Yang: I responded to an email saying that we do not allow and will not access unknown models. As a result, I searched OpenClaw and Agnes on GitHub today. From June 1st to 3rd, there were dozens of comments every day asking why Agnes AI is not supported and why I need to configure it myself. So we gave some sharing and got very touching feedback.
Geek Park: I chatted with Yang Pan of Silicon-based Liquidity before, and he gave me a suggestion - subscribe to a $200 version, and you will find that your ambitions will become bigger when you have unlimited tokens.
Bruce Yang:Yes, that’s what we think too. In fact, before promoting free services, the company had not fully figured out what to do next and what the business model would be after free. There was only a rough idea. But we have a big understanding. When you take something to the extreme, such as reducing the price to free, it will definitely open up a new opening mode for the entire ecosystem. It is a paradigm shift and many scenarios will burst out. We don’t need to think about these scenarios now. Many users will help us think better, because the power of the masses is unlimited. This is what we have already seen, some seeds are already blooming.
Geek Park: Are you worried that someone will not only engage in prostitution for free, but also set up something similar to a transfer station to transfer your free tokens to more people and start charging themselves? Are you worried about the emergence of such second-rate dealers?
Bruce Yang:We limit the RPM, which is the number of requests per minute, to about 20 times per minute. It will definitely be no problem for individual users, but it will be more difficult for enterprise users. If you give a 20 RPM product to 10 users, they will feel stretched. Therefore, for enterprise users, there may still be a paid model in the future. Of course, the price is also very cheap. You can use the free one to do POC and pilot first.
Geek Park: In a CLI environment, which tasks should be paid and which ones should use the Agnes free model to maximize the economy for individuals?
Bruce Yang:Most people, unless you are a geek. I think there are two types of users who can be a little more cautious. The first category is absolute geeks, such as those who need multiple codex instances and run continuously for 3 to 4 hours. We currently have no support for this. Of course, we are optimizing it and are optimizing it with our coding harness for this long-range, multi-instance scenario. The second category is very professional in making short plays. It’s not that we can’t be used, but that in certain scenes, such as particularly complex movements and scenes that require consistency, we can be used together with some higher-end models.
In addition, our model should currently be able to solve more than 95% of the scenarios on the market, which has also been verified in more than a dozen of our WeChat groups. About 80% of users would say that you are similar to other models we have seen. There are also some users who will ask questions. These questions are divided into two categories, one can be solved quickly, and the other is temporarily unsolvable. Those who can solve it quickly account for 80% of the 10% to 20% of users who ask questions. Calculated in this way, there are only about 1% of the scenarios and problems that have not been solved and do not know how to solve them. In addition, we have lowered the usage threshold to free, I think it is a very popular direction and worth trying.
Geek Park: How did Agnes reduce the cost of the three modal models to be free? Hundreds of billions of tokens were sold out in just three days.
Bruce Yang:Yes. Moreover, the hundreds of billions of tokens are only 1/5 of our reserve cards, which can be multiplied by 5 times based on daily consumption. I have also prepared a second batch of cards. You can boldly collect them until we can no longer collect them.
The logic is this. First, what we are doing is an optimization problem, but the constraints are different from others. Most mainstream companies believe in scaling law, which means that parameters and data can be improved equally if computing power permits. But it does not answer the question of how big the marginal benefit is: in many cases, when the parameters are increased by 10 times, the benchmark only increases by a few percentage points; and most of them are now reverse distillation. For example, Gemini uses Pro to distill Fast, and the parameters are reduced by 10 times. The difference on most lists is not big.
So we made an important assumption on the first day. We will not do models above 200B. We will only optimize models within 200B and find a suitable range within them. Relying on environmental stability, synthetic data and online data of our own products to continuously expand, and then expanding on similar issues in list data, this area is now very mature, and we will soon open source some synthetic data methods.
On top of this, we only focus on two key points: agent and coding, hoping to be as good as the SOTA model. Strategic abandonment of other areas is not unimportant, but it is not the first step to solve. Because tokens are now consumed on a large scale, it must be a coding harness or a white-collar office harness.
In addition, there is a slightly advanced attempt. We published an article on the official website about how to approximate the effect of a larger model by cyclically calling Transformer layers without increasing parameters and depth. This is called recurrent depth transformer. In small-scale verification, the PPL in one cycle dropped by 10%, which is equivalent to a 10% increase in parameter utilization; calling the same MoE model multiple times can better utilize the capabilities of each unit parameter. This is the next step to focus on. The long-term vision is to continuously optimize performance within 200B and approach SOTA. Resources are limited, but it seems to be quite effective so far.
Pictures and videos are different. They have not broken through the scaling law. Basically, the more data, the better the effect. Many products fail to achieve results, not because of capability issues, but because of data issues, and synthetic data is very complex. For example, if you want 100 million videos, it may take several months to crawl and cut them by yourself. By the time you finish, the opportunity has passed.
So how do you get the data you want in the shortest time? What kind of pipeline is used to train this data? How to let the picture model empower the video model? Should I choose DiT or auto regression as the technical route in the process? There are actually many small know-hows here, which are more important than a one-time big concept upgrade. A bit like Yann Dubois, a person in charge of post-training at OpenAI, said that training models is actually more like a manual job, not a conclusion that can be systematically derived.
In the past year or so, hundreds of our scientific research colleagues have made a lot of innovations and given full play to the power of academia and open source, so we are also feeding back to the open source ecosystem. For example, the previous paper on recurrent depth transformer has been open sourced; next week we will open source a VAE module that allows us to optimize text in images; later on in the video model, the most important thing is how to synthesize data, and we will gradually open source it.
This ecosystem is still very helpful to us. Many raw materials and many dishes are actually available, but do you have a big enough and strong enough team, and do you have enough confidence to invest and cook this dish? I think we've had a pretty good burn so far.
Geek Park: What is the business thinking behind “Token Free”?
Bruce Yang: I have a rough idea, but I haven’t thought it all through. I can share some. Let’s talk about numbers first. We made hundreds of billions of tokens in just a few days. I took a look and found that the current number one on OpenRouter is DeepSeek V4 Flash, with about 3 trillion tokens a week. I did some math and found that if we reach this kind of weekly usage, our actual server cost will be around a few million yuan, which is not a big number at all. A very important reason is that we have compressed the cost to the extreme. I don’t see anyone in the market currently who can achieve our cost, which is a bit exaggerated.
How free do you want it to be this time? The goal is to reach twice the size of OpenRouter’s first place. If there are new users after the doubling, we may continue to support it, depending on our financing situation; but within the doubling, we can fully support it. At present, our team is as big as OpenRouter ranking first. It is mainly provided to individual consumers. We have not made any large-scale publicity to enterprise consumers for the time being. You can do POC, but the RPM given is not that big. If the volume reaches twice the size of the largest model on OpenRouter, it can be supported for free. Because rather than saving this cost, we hope that more users will experience our model, like our model, and become our loyal users, which is very worthwhile.
We have several ideas on how to commercialize it next.
The first is enterprise users. It’s tiring to make sales, but it will be much faster if you open a free platform for him to try and let him take the initiative to come to us. This is a very important commercialization path for us.
Second, we see that OpenAI and Anthropic’s fastest growing business-side products are their harnesses, namely Claude Code and Codex, so we will soon launch our own harness products. This is a bit of a hiccup first, but this is also a very important commercialization path.
Third, for geeks with particularly large usage, this is not the focus. We will upgrade to a better model. When it reaches a very SOTA and top three in the market, we can consider charging a small range, or give priority to paying users. After paying for a period of time, we can still make it free. But these are not the highest priority, the first two have higher priority.
极客公园:今天是「Token 免费」,下一步会出现「给用户钱让他们用 Token」吗?
Bruce:有这种可能性。但总体来说,在 AI 时代,想保持一两年的门槛和壁垒是很困难的。我们现在趁着有这个能力,全模态模型都能达到可用状态、能达到全球模型榜单前十的 Lab,率先打出免费的旗帜,希望先把愿景推出来,因为这个行为背后跟我们的愿景是符合的。能完全匹配全模态、同样能力又免费的,目前市场上公司不多,大部分公司选择在某一个领域发力,其他领域虽然也在慢慢发力,但需要时间。
所以我们想借这个机会尽快先上牌桌、先成为一个重要玩家。我们后面也有后手,别人匹配我们时,我们还有别的招没出,harness 产品就是我们现在正紧锣密鼓准备的,具体什么时间点、推什么样的产品暂时还不能说,但后面还有新的增长曲线。
极客公园:大厂会跟进吗?例如把过去的模型也免费出去?
Bruce:看他们多快能匹配,我觉得有难度,毕竟已经有那么多用户在付费了。我们作为新参与者,没有那么多包袱,没有那么多企业用户和规模性付费用户,所以可以快速掉头;但对很多公司来说船大难掉头,整个规划、预算、年度计划都要调整,大公司的决策路径没有那么快。
极客公园:刚才说到很多普通用户用免费 token 生成和女儿回忆的影像。这是不是你和团队的一种情结,希望把 AI 作为工具免费给大家,让大家释放创造力、让生活更美好?
Bruce Yang:我先介绍一下我的背景。我从小在国内一个四线城市长大,初中靠竞赛和中考成绩拿到奖学金,去了新加坡莱佛士书院,相当于新加坡最好的高中。在那里我认识了很多来自东南亚、家庭不富裕但成绩很好的同学,有了很多新认知。我参加新加坡全国的数学、物理、化学竞赛都是金牌、全国前几名,也进了学生会。 靠这份经历,我拿着leadership奖学金去了 UC Berkeley 读书。
整个硅谷有两所学校,有人说富人的孩子去 Stanford,穷人的孩子去 Berkeley。 Berkeley 的同学很像一个社会,不是那么标准的精英,但每个人都很聪明、有很多想法,很纯粹、很干净。
之后我在硅谷创业,这次回新加坡读博也拿了总统奖学金。我运气非常好,来自四线城市、父母也不富裕,但一路都有奖学金和支持。今天的很多成绩都是当时的积累,加上一颗不服输的心,虽是后来者,也愿意挑战现在的市场玩家。 但 AI 现在变得没那么平权了,因为成本,很多有创意的人都在意 token 消耗,不敢大规模用,反而没那么有创造力、没那么有效率。
回想我的经历,无论是莱佛士那些拿奖学金的同学,还是学费不贵、让加州很多普通家庭聪明孩子都能去的 Berkeley,这颗种子是我自己得到的,我也到了一个时间点要回报社会,把火种传下去,就是平权:能力的平权、价值的平权。
在这个时代,AI 平权是最核心的。 10 年前你不会中文、英文可能是文盲;10 年之后,不懂 AI 可能就是文盲。
我硅谷的朋友很多很反 AI、害怕 AI。其实不是害怕 AI,而是不懂 AI 的人怕懂 AI 的人,觉得自己随时会被替代。解决的办法不是压制 AI,而是让它变成一种更平权的能力,让每个人都知道如何借 AI 创造更多。这也是我们公司很重要的愿景,让世界级的 AI 属于每一个人。我们能做的可能微不足道,但这个愿景非常长久、持久。
极客公园:很多大厂已经不开源了,但是你们还在做开源。除了 AI 平权,背后还有哪些思考?
Bruce Yang:现在很多公司在尝试做开源,但只开源了参数、没开源方法。既开源参数又开源方法的,就是 DeepSeek,所以我对 DeepSeek 非常 respect。梁文峰确实是在做 AGI,如果你现在问我,全世界这么多做 AI 的人最崇拜谁,我肯定还是梁文峰,一年前是,现在还是,他有大局观、大格局。我们也是一样的想法。如果开源了模型,但模型太大没法自己部署,又不开源方法,那更多只是证明自己有这个能力、证明自己的模型跟别人不一样、可以被别人蒸馏调用,并没有为社区反馈太多信息。
所以我们很想做的是,如果真能做到一些别人做不到的成绩,还是想把方法论开源出去。无论是上周开源的 recurrent depth transformer,还是下周要开源的、让图片文字更清晰的 VAE,还是后面告诉大家训练视频模型最大的卡点其实是如何快速合成数据,这些能力我们都会想着分享出去。
一方面是想证明我们有能力创新,不希望大家认为我们只是个跟随者;另一方面,得益于人、也反馈于人,希望能在开源社区、开源生态里成长,也希望能反馈给社区。我们各个群里很多小伙伴都在帮我们写 skills,很多我们自己都没写,但你现在搜 GitHub「Agnes 模型」,很多 skill 都写出来了。我知道的群里小伙伴大概就写了四五个,还不断在 OpenClaw 的 issue 里催更,问为什么不支持 Agnes。
极客公园:催更 Peter(OpenClaw 创始人)是吧?
Bruce:对,催更 Peter,而且好几个还是中文的催更 Peter。 这样的生态是大家比较希望看到、比较期待的,这也是为什么我觉得国内的 AI 现在在领跑全球。
极客公园:如果让你给大家传递一个信息,token 都免费了、门槛已经降到很低,普通人在这样的时代应该怎么做、应该有什么样的态度?
Bruce Yang:越早拥抱 AI,越能理解 AI 的世界,而 AI 世界和非 AI 世界是不一样的。我在 NUS 读博时上过一门机器人课,博士课程,我拿了全班第一名。教授 David 第一天就跟我们说,你们可以用 AI,但要说明自己是怎么用的,最好把提示词写出来。结果那门课,我读博时已经比同龄人大 10 岁,花的时间其实不多,但无论做项目、做研究、写论文还是做演示,我都大规模用 AI,居然在大部分同学都比我小 10 岁、可能更有精力的情况下拿了全班第一。这说明如何充分发挥 AI 很重要,AI 能发挥的维度可能远超我们的理解,尤其这波 harness、Codex,包括理解屏幕、做很多新的 skills、对接 MCP 插件,已经在完全改变这个世界了。
我身边有些朋友在做 AI 应用,我们自己也做过一段时间,现在不是公司重点。我有个很重要的观点,当一个产品越做越复杂,它就不是一个 AI native 的产品。因为 AI native 的产品大部分是越做越简单,越来越依赖模型;短期内可能会部分依赖 harness,但这种依赖会不断迭代、可能越来越少。
所以更先进的 AI 认知、更早地接触 AI 产品,再加上免费的资源让大家大胆尝试,我们就把门槛降得很低。很多人不敢尝试,就是怕费太多 token、太多时间;如果 token 都免费了,每天都可以尝试、不断和 AI 互动。 AI 本身是双向的,不一定需要一份操作手册。这样你可以越来越全面地理解 AI 的每一个角落、它的边界在哪里、它的脾性在哪里,这才是新时代的 AI 平权。
有时候我们用 AI 去改造传统业务,有点像把马车装得更豪华、让马跑得更快;但真正 AI native 逻辑,其实是换一辆汽车,是彻底改变对行业的认知。这种认知有些地方比较根深蒂固,我们希望通过免费的、足够多的 token,让大家在这个转变中更快地适应新时代。
我们后面还会出大量的场景和案例,让大家快速上手,包括给没试过 vibe coding 的同学,把我们的一些提示词和生成效果都分享出来;以及如何连接大家想用的 harness。最简单的我们自己也提供了 harness,叫 Agnes super agent,现在还没做得那么好,但已经可以尝试。
如果你自己有 harness,比如 Codex、Claude Code、OpenClaw、Hermes,都可以快速对接。这些资源我们都会快速分享出来。我们的逻辑就是让大家无门槛上手,而且是真免费、没有任何套路。案例和提示词都会慢慢分享出去,让大家无论已经是极客,还是想快速开始 vibe coding,都能快速体验起来。