-
Cryptocurrencies
-
Exchanges
-
Media
All languages
Cryptocurrencies
Exchanges
Media
Share
Author: Boyang, Tencent Technology
On March 6, 2026, nearly a thousand people lined up downstairs of the Tencent Building in Shenzhen, not to grab mobile phones, but to ask someone to help them install a software. Its scalper price once reached 1,000 yuan. Longgang District and Wuxi High-tech Zone directly included this software in government subsidy documents. Sam Altman admitted that when faced with similar self-driving products, his initial decision not to let AI fully control the computer "only lasted two hours."
This software is called OpenClaw, which is an open source AI Agent.
In the first quarter of 2026, it emerged at the same time as four other completely different Agent product forms during the same window period. OpenClaw is a personal assistant, Cowork is an office collaboration, Codex App is a long-distance engineering task, Perplexity Computer is a unified workstation, and Tencent Cloud ADP is an enterprise platform.
It is no coincidence that each of the five companies has taken its own route. Coincidence is when a company happens to make a good product. Five companies taking action at the same time can only mean one thing - a certain underlying condition has just matured, and everyone smells it at the same time.

In this quarter, we're not focusing on single point events, but on structural changes, the things that really change the game.
There are four filtering criteria.
First, the whole industry resonates. It’s not a single company acting alone, it’s multiple people running in the same direction at the same time. Five companies launched Agent products at the same time, dozens of teams were building constraint frameworks at the same time, and at least three independent routes completed recursive R&D at the same time.
When everyone takes action at the same time, it is not a question of who has the vision, but the foundation has changed.
Second, causal bite. It's not four things that happened to collide in the same quarter, but the former one directly gave birth to the latter one. If any link is removed, the latter one will not hold.
Third, qualitative changes are perceived. This shows that these trends are not small improvements within the industry, but have crossed a certain critical point and have greatly improved to the point where the public can perceive them. People queuing up to install OpenClaw in Shenzhen made social news, and the "Lobster War" became a popular topic. The government included Agent in subsidy documents, and 22% of employees secretly used it without telling the IT department. When a technology trend overflows the technology circle and enters public discussion, it is no longer an "industry trend" but a signal of a turning point in the times.
Fourth, cognition is irreversible. Specific products will be replaced and specific frameworks will be iterated, but the ideas behind these trends will not disappear. For example, the consensus that "Agents need discipline" will not go back, the direction that "experience should be reusable by Agents" will not go back, and the expectation that "Agents should be able to improve themselves" will not go back. The form will change, but the cognition will not.
Q1 has exactly four that meet these four criteria.
1. Automated AI Agent enters productization. Agent can finally do things independently, moving from minute-level demonstration to day-level execution.
2. Constraint engineering. Agent learned to abide by the rules, and within 6 weeks the industry forced a set of disciplinary frameworks.
3. Recursive research and development. Agents begin to grow themselves, not just by performing tasks, but by improving the way they perform tasks.
4. Skill ecology. Agent began to inherit previous experience through the Skill model. For the first time, human industry know-how has a format that can be directly reused by Agent.
And these four forces are not parallel, but a flywheel.
After Agent was able to work independently, the problem of unruliness was exposed, which forced constraint engineering; constraint engineering gave discipline, so recursive R&D could run; recursive R&D created a strong need for experience reuse, giving birth to a skills ecosystem; the skills ecosystem in turn allowed Agents to handle more complex tasks, and the flywheel turned to the next circle.
Q1 is the first quarter in which this flywheel turns completely.
On April 10, 2026, Tencent News released the "AI Trend Research White Paper 2026Q1" (hereinafter referred to as the "White Paper"). This 59-page report focuses on the operating logic of the entire flywheel.
This article is to streamline the content of the white paper and give 25 specific judgments based on these four forces.
Agent used to be like a kid in a talent show. It's amazing to ask it to perform a bit, but you really don't dare to leave it to it. In the past, model demonstrations were astounding in just three steps, but in the fifth step, they completely lost their overall vision and started acting randomly.
Q1 This thing has changed. The point of change is not simply that the model has a higher IQ, but that the Agent can finally do "you go to sleep, and it will work there by itself."
Cursor Agent single task has been running for 36 hours. Claude Code submitted a maximum of 4% of the world's public GitHub code in a single day, with annual revenue of approximately US$2.5 billion. Dario Amodei confirmed that more than 90% of Claude's new code was written by the AI itself. There is even an engineering lead within Anthropic who said, "I don't write any code anymore, I just let Opus do it and I edit it." Anthropic posted 74 updates in 52 days. Codex has more than 1.6 million weekly active users and more than 1 million desktop application downloads.
The most dazzling OpenClaw GitHub star number soared from 9K to 247K and 2 million monthly active users in 60 days.
In addition, Karpathy called Moltbook, which has driven 1.5 million Agent registrations, "the closest reality to a sci-fi takeoff in recent times." OpenAI announced the acquisition of the founder of OpenClaw on Valentine’s Day.
China reacted even more violently. At least nine companies launched desktop agent products in the same quarter. Tencent tied it to WeChat and Qiwei, Byte anchored Feishu and cloud SaaS, Alibaba switched from coding tools to general office, and Baidu relied on search skills to lower the threshold. The industry calls it "Lobster War" - named after OpenClaw's logo, a lobster, which means that the Agent has finally grown pliers that can grasp things.
Agent can indeed do things independently. But why now?
None of the six dimensions of OpenClaw - continuous online, heartbeat mechanism, externalized memory, Skill (skill package), browser takeover, remote node call - is original. AutoGPT and various browser proxies have already drawn this pie. But OpenClaw welded them together and produced a qualitative change.

What really breaks it out of the circle are two simpler things, IM (instant messaging) access and 7×24 initiative.
Cowork almost fully matches or even surpasses OpenClaw in terms of capabilities. Anthropic's three-tier product system - Claude Code command line, Cowork desktop application, Computer Use (computer operation) + Dispatch (dispatch) cross-device remote control - is far more sophisticated than OpenClaw in terms of technical depth. Computer Use matched human performance on the OSWorld benchmark (72.5% vs. 72.4% human). But it was missing only two things.
IM lets Agent wait for you in the interface you are most familiar with. 7×24 Let it wake up and patrol on its own without waiting for you to say anything. Put the two together, the Agent no longer waits for you to speak, it comes to you on its own initiative. OpenClaw does not explain to the user what a contextual window or search enhancement is, and simply throws out a plain saying - "I will always be online, I will remember what you said, and I will get things done myself." Let’s look at the efficacy first and then talk about the principle. This market approach directly breaks through the technical barriers. 22% of employees secretly use OpenClaw without the IT department knowing about it.
Accessibility trumps ability. The technical depth is not as deep as Cowork's OpenClaw, which takes away the user's mind because it appears in front of the user on the right interface, at the right time, and in the right attitude.
In the automation products of the same period as OpenClaw, we can see five routes.

Behind them are two conditions in place at the same time.
First, the model finally crossed the passing line of "sustainable execution." The current model still makes mistakes, but at least it can hold on through a cycle of dozens of steps without suddenly forgetting what it is doing halfway through. This difference is decisive - local mistakes can be corrected by system scaffolding, but global collapse cannot be cured.
Second, the Harness (scaffolding) engineering methodology is stable enough. The memory is changed from a black box vector database to a plain text file that users can directly browse and edit, supporting Git version control. The execution environment has a gateway, heartbeat mechanism, browser takeover and remote node call.
The ability and scaffolding are in place at the same time, and long-range agents have become a common choice for the entire industry. Codex App's Worktree architecture allows multiple Agents to work in parallel in the same code repository. Five parallel Worktrees reduce the 42-minute task to 14 minutes, with zero merge conflicts. The execution span of Agent has officially moved from minutes to days.
OpenClaw coupled with the Skill Market turns Agent from a developer tool into a general work assistant, capable of everything from research, monitoring, content generation, and customer service. 13,700 Skills cover scenarios far beyond coding.
The pattern here is very clear. As long as there is Skill to standardize know-how (industry experience) in the field, any long-term and high-cognition work will be the choice of Agent. But short-range low-cognition operations are not like this - ordering a cup of milk tea, you can do it on your phone in 30 seconds, but the Agent is even slower.
in the country, nine major manufacturers have tied up their own IM to compete for entrance. This battle is about "which App should the Agent live in?" and the battle is for the right to enter the ecosystem. But by the end of Q1, this wall began to loosen. QClaw supports Feishu and DingTalk, and OpenClaw-CN has five built-in IM adaptations.
Because when the Agent needs to help you handle customer messages in WeChat, team collaboration in Feishu, and approval processes in DingTalk, it cannot live in just one IM cage. The more demanding the Agent is, the greater the cross-platform pressure will be, and the harder it will be to sustain the walled garden. The competition now is "who is closest to the user", and in the future the competition will be "who can make the Agent work seamlessly in all places".
The styles of play in Silicon Valley and China are very different. Silicon Valley started fighting over "model suppliers vs. middle layers" - Google suddenly banned users who called Gemini through OpenClaw on a large scale in mid-February without any prior warning, and hundreds of paid accounts were shut down overnight. The superficial reason is that "malicious use causes the computing load to far exceed expectations." In fact, OpenClaw's heartbeat mechanism checks the complete context of tens of thousands of Tokens every 30 minutes. The actual consumption of a single Ultra subscriber is converted into an API price of US$1,000-3,600, which is far more than the US$250 monthly fee. This is a direct impact on the subscription business model.
Anthropic directly characterized this behavior as "Token arbitrage", requiring users to use API key access (the price is 5-10 times that of the subscription system), and finally directly blocked Openclaw's subscription entrance in early April. OpenAI chose the opposite path, acquiring the founder of OpenClaw and then whitelisting it.
To put it bluntly, when an open source middle layer allows users to bypass official pricing to obtain model capabilities, the platform must choose between blocking and inclusion.
The same OpenClaw spawned a pricing debate in Silicon Valley and an entry war in China.
Julien Bek of Sequoia calculated that for every $1 companies spend on software, they spend $6 on services. Accounting, legal, IT hosting, recruitment, insurance brokerage, all services. The pricing unit of Agent is changing from seat (per person) and feature (per function) to workflow (by process) and outcome (by result), but the real opportunity is not to replace people, but to replace outsourcing contracts.
Think about it, a job has been outsourced, which means that the company inherently accepts external execution, has a ready-made budget line, and buys results. Replacing outsourcing is equivalent to changing a supplier, and replacing internal employees is equivalent to organizational adjustment. The resistance of the former is an order of magnitude smaller.
This explains why vertical autopilot (autonomous driving agents) such as Harvey (legal), Anterior (medical approval), and WithCoverage (insurance) are launched much faster than general agents - they are not targeting the political minefield of "AI replacing people", but the natural business area of "AI replacing outsourcing". A comparison with OpenClaw also proves that the first batch of tasks that individual users ask Agent to perform are exactly the tasks that they used to pay virtual assistants to do, such as monitoring, research, and social media management. The service market that is 6 times larger than the software is the real foundation of autopilot.
Block (the parent company of Square + Cash App) shows an extreme form. The company is directly restructured into a four-layer intelligent agent architecture, middle-level management is eliminated, and the product roadmap is automatically generated by the failure signal of the intelligent layer. There is no longer a product manager to decide what to do. But the premise of Block is high-frequency structured data for bilateral trading platforms, which most companies do not have.
What really blocks the Agent from copilot (co-pilot) to autopilot (autonomous driving) is not the model capability, but three organizational interfaces with independent rhythms, including the evaluation system (how to judge whether the Agent's performance is up to standard, accounting has GAAP standards, but the legal advice and recruitment screening criteria are far from established), authorization boundaries (who allows it to execute, whether the execution boundaries are written in the system prompt or written in the business rules system), and responsibility attribution (who will take the responsibility when something goes wrong, copilot In autopilot mode, there is no professional qualification or insurance behind the system).
Model capability is a function of computing power and capital, and you can increase it by spending money. These three things are functions of time and system, and they will not automatically keep up as the model becomes stronger. Between copilot and autopilot, there is a high probability that there will be a severely underestimated transitional form - AI executes autonomously within clear boundaries, the boundaries are defined by business rules rather than prompts, and automatically escalates to humans when encountering anomalies. This is closer to the money than pure autopilot and better value than pure copilot.
Being able to do things independently is the starting point of the flywheel. But Agent’s fatal shortcomings were exposed as soon as it hit the road. OpenClaw’s 512 security vulnerabilities, 341 malicious Skills, and bills of hundreds or thousands of dollars are all obvious pitfalls.
The ability to be independent brings new problems, and the new problems force out the second force of the flywheel.
The number one problem exposed after Agent can do things independently is that it does not follow the rules. Memory is like a goldfish. After taking three steps, you will declare that you are done. You will give yourself a high score but it will not work from end to end.
Q1 It took 15 weeks to force a set of solutions. From Anthropic publishing the first Harness blog on December 5, 2025 to LangChain’s generalized definition on March 10, Harness Engineering (constraint engineering) has achieved industry consensus.
The speed itself shows how anxious everyone is. Because the Agent is already on the road, the rules haven’t been established yet.
Context Engineering manages the information layer, that is, "what the model should see." Harness Engineering is in charge of the structural layer, that is, "how the model continues to do things right in dozens of rounds."
The latter is Q1’s true paradigm jump.
Claude Code generates 135,000 to 326,000 public commits on GitHub every day, accounting for 4% of global public commits and is expected to reach 20% by the end of the year. Agent has run so deep in the code warehouse that it does not deserve a special set of disciplines. Sooner or later something will happen.
Tencent Senior Executive Vice President Tang Daosheng set the tone for the Chinese version at the Shanghai Summit on March 27: "AI implementation is not just an algorithm problem, but also an engineering problem. As the gap in the capabilities of mainstream large models narrows, the core of enterprise competition is no longer the strength of the model itself, but the ability to leverage the value of the model through engineering means."
Other major manufacturers have also entered the game. Byte's DeerFlow 2.0 directly writes "Super Agent Harness" in the GitHub description. It is probably the first time that a Chinese open source project has used this term in product positioning. It soared from 22K to 52K stars on Github within a month, and everyone is eager for Harness.
To understand Harness, we can imagine Agent as a car. The model is the engine and the prompt is the steering wheel. But just having an engine and a steering wheel doesn't make it a car. You also have to install a gearbox, instrument panel, and brakes. Harness is these components. In current engineering practice, there are three layers in total.

The first level, process control, specializes in treating disobedient people. Memory is like a goldfish. After taking three steps, it is declared complete. There is a bug in the environment and you don't even know it. The governance method is to externalize the state (AGENTS.md/progress file), split tasks, and force steps to be followed.
The second layer, concurrent scheduling, specializes in group fishing. With one hundred Agents running at the same time, it is easy to avoid risks and focus on the simplest small modifications. No one touches the real problems. The governance method is a multi-Agent hierarchical structure, role separation (Planner planning → Generator generation → Evaluator evaluation), and an anti-fishing mechanism.
The third level is to verify and correct errors, and to gain confidence in the treatment of confusion. Giving yourself high marks is what Anthropic calls "self-deception". The governance method is independent Evaluator, sandbox isolation, and Git transaction boundaries (Branch is a sandbox, PR is for approval, and Merge is considered a submission).
CLI, Skill, externalized memory format, etc., which are often mentioned, have also become trends in the development field in the first quarter. In fact, in engineering practice, they cannot be completely considered Harness, but Agentic Infra (Agent infrastructure). Harness only cares about "how to drive the car stably", while Infra cares about "road conditions and gas stations" that can increase speed and speed.
Harness does not appear out of thin air, but comes from a series of specific problems encountered in engineering practice. When people want to let the model perform longer-term tasks, bugs force it out again and again.
The origin of the first level. The user throws a large demand to the Agent at once, and it tries to complete it all in one go and crashes in step 30. The solution Anthropic came up with is very simple, like a relay race - "Initialize Agent" sets up the environment and writes a handover list (claude-progress.txt), and then exits. "Coding Agent" first reads the handover list every time it comes on stage to figure out what it has done in the previous session, only performs one function, and then exits after completing the updated list. The main point is that the conversation history is not shared between Agents, and information is only transferred through files. Because by the tenth round, the conversation history has been completely drowned out by the noise of the first nine rounds.
The origin of the second level. Cursor found that Agent was extremely risk-averse under a flat structure, preferring to make meaningless small modifications rather than deal with difficult problems, and the entire system was idling. Anthropic has introduced a "Party A and Party B" structure - the Planner writes the specifications, and the Generator implements functions one by one according to the specifications, and writes a Sprint contract before starting work on each function. Evaluator uses a real browser (not looking at the code) to test functions, and scores based on four dimensions: product depth, functionality, visual design, and code quality. If the standards are not met, the sprint will fail and must be reworked. An interesting finding is that it is much easier to make the "scorers" strict alone than to teach the "coders" to learn to self-criticize.
The origin of the third level. The Agent ran the test he wrote and said "there are no bugs", but it didn't run end-to-end at all. Anthropic calls this "self-deception" - it's the same as asking students to grade their own essays, and the grade will never be low. There must be independent Evaluator and sandbox isolation.
The AGENTS.md in the open source project Ghostty by Mitchell Hashimoto is not a design document at all, it is an accident file - if the Agent touches files that should not be touched, just add a "Do not modify the vendor/ directory". If the Agent uses an outdated interface, add a line "Use v2 API instead of v1". If the agent writes randomly in the commit message, just add a format specification. He found that only 10-20% of the time during normal working days could effectively run the background Agent. At the beginning, "it took longer than doing it manually." But once the inflection point is passed and the rules accumulate to a certain density, the Agent's error rate drops significantly.
OpenAI discovered a more chronic disease. Without a dedicated maintenance mechanism, the warehouse used by the Agent will deteriorate significantly after about 2-3 months. It's like ten interns come in to work every day and leave behind a bunch of "temporary solutions." Three months later, no one can tell which code was written seriously and which code was makeshift. They set up three mechanisms - architectural constraint declaration (clearly write out in AGENTS.md what framework, pattern and naming convention the project uses), Doc gardening (regularly clean up expired comments and redundant documents left by the Agent), and Anti-slop routine (anti-deterioration inspection, clean up the style inconsistencies and duplicate code accumulated by the Agent). It’s not something you can do right once, you have to keep running.
LangChain conducted an experiment. Using the same model but using Harness, the Terminal Bench 2.0 pass rate increased from 52.8% to 66.5%. The weight did not change even one byte, and the ranking soared from thirty to the top five.
This is the effect of Harness.
But this effect is very expensive.
Anthropic's cost data shows that Solo Agent cost only $9 and took 20 minutes to run the same 2D game naked. However, the main functions of the product were damaged and it was impossible to play. Putting on the complete Harness, spending 200 US dollars and 6 hours of use, the finished product is fully functional, visually excellent, and can be played normally.
What 20 times the cost brings is not "better", but the difference between "being usable and not being usable".
From this point of view, Harness is currently the most cost-effective power amplifier. But it’s really not cheap. It’s not at the same level as maintaining a continuously running Agent that understands the rules and a chat assistant that occasionally asks questions.
Bills of hundreds or even thousands of dollars have caused many users to be directly persuaded to quit after a few weeks of experience. Cost remains the biggest stumbling block on the road to popularization.
While the entire industry is working hard to lay the bricks of Harness, Anthropic is already smashing the shell it has built.
After Opus 4.6 was released, they removed Context Reset because the model's context management capabilities were already strong enough that there was no need to reset the context. The Sprint Contract was removed because the new model can control its own pace and does not require an acceptance contract to be signed before each round of construction. The Evaluator has also been changed from each round of confrontation to the last round of QA.
According to Anthropic's own words, "Each component of Harness encodes an assumption about what the model cannot do. When the assumption is no longer true, the component should go."
Being able to disassemble it means that it was built effectively. The decisiveness of the demolition shows that they always knew what they were compensating for.
The difficult thing is not to dismantle it, but to judge when to dismantle it. If you dismantle it too early, the model will not be able to hold on, and the system will collapse; if you dismantle it too late, the shell will obscure the true capabilities of the model.
When the model evolves, you think the shell is helping, but actually the shell is getting in the way.
The road to simplicity must pass through complexity. But currently, Anthropic is the only company that has completed the complete cycle from “addition” to “disassembly”.
The five consensus points hotly discussed in Q1 are Markdown as state carrier, Git as transaction boundary, CLI Renaissance (Agent needs a structured text interface, and a git diff --stat can get the entire change overview), real-time testing (Agent changes a file and immediately runs the test, and the test results become instant feedback signals for each step), and Skill is used for knowledge encapsulation. In fact, they are not all Harness.
The larger framework is called Agentic Infra (Agent infrastructure), which is divided into five layers - Context layer (what the Agent can remember) / Tool interface layer (What the Agent can do) / Harness layer (how the Agent is controlled) / Knowledge layer (The Agent knows what to do) / Economic layer (how much it costs to run the Agent).
Harness is the core layer, but it is far from the whole story. At present, the industry's attention is focused on the Harness layer, because it directly determines whether the Agent can be used. But the next stage of the battle may be fought in other layers - the memory quality of the Context layer (memory retention across days and sessions is still very unreliable), the execution environment resources of the tool interface layer (Anthropic found that relaxing resource restrictions alone can increase the success rate by 6 percentage points), and the Skill triggering mechanism of the knowledge layer (Vercel evaluation shows that in 56% of cases, the Agent will not take the initiative to check what it has. Skill), and cost control at the economic layer (model routing, budget allocation, and parallelization strategies have basically been slapped on the head so far).
There is another question that everyone knows but no one has a good answer to: organizational governance. "Who approved the Agent to write this code?" There is no standard answer in most companies. Audit logs, decision traceability, and code ownership. A three-person entrepreneurial team may be able to rely on the approval level at the workflow level. But in a company with 500 people, there is no independent governance framework, and technical means alone cannot succeed.
This is an irresponsible but crucial part of the Harness architecture where the Agent can really reassure people.
Three-layer shell + five-layer Infra gives the industry the first diagnostic framework. When Agent crashes, you can no longer just say "the model is not good enough".
Is the process control not done well? Concurrent scheduling out of control? Is the verification link missing? Or are the execution environment resources insufficient, the Skill not triggered, and the cost just not worth it?
In the past, metaphysical problems that were all classified as "insufficient model capabilities" now finally have precise engineering attributions. "Why did the Agent crash?" has changed from a vague complaint to a locable and repairable engineering problem.
Constraint engineering gives the Agent discipline, and the flywheel can turn to the next circle.
A disciplined Agent finally has a previously impossible ability to continuously improve itself in a long-term cycle, rather than collapse at the tenth step.
This is the premise of the third force of the flywheel.
The first two chapters talked about how Agent stands up as a product and system. This chapter talks about the scene in which the Agent first broke through the role of "executor" and began to improve its execution method after it gained discipline.
The answer is research and development. Because R&D is naturally verifiable (passing a test means passing), reversible (one-click undo with Git), and readable and writable (the code itself is plain text that can be directly operated by the machine).
When the three conditions are put together, the Agent can enter the complete cycle of "execution → verification → problem discovery → modification → execution again".

There are three paths here, and they all have dazzling result data. AlphaEvolve recovers 0.7% of the world's computing power, Minimax M2.7's internal evaluation improved by 30% after 100+ rounds of independent iterations, and Karpathy's Autoresearch runs 50 experiments a night.
These three open source practices show that recursive R&D is already generating real money value. The major model manufacturers that have not yet opened up the source have also admitted that this happened in various interviews.
Now let's look at these three types of methods in detail.
Exploration type, AlphaEvolve. It is not adjusting parameters, but searching for new algorithms that humans have never seen before. An evolutionary system composed of Gemini Flash (responsible for breadth, quickly generating a large number of variants) and Gemini Pro (responsible for depth, carefully crafting the optimal solution). The entire process produces human-readable code, not a black box—it can be understood, adjusted, and deployed directly. The data center scheduling algorithm it discovered has been running in Google's production environment for a year, continuing to recover 0.7% of the world's computing power, which is worth billions of dollars in money. It provides optimization solutions for key TPU circuits, increasing the speed of a key computing core in the Gemini architecture by 23%, and optimizing the underlying instructions of FlashAttention by 32.5%. Improvements were made on the best known solutions in 20% of more than 50 mathematical open problems, including improvements to the matrix multiplication algorithm proposed by Strassen in 1969. The value of this kind of recursion is that it may change the direction of a discipline.
Optimized, Autoresearch and M2.7. The objective function is known, which is to iterate repeatedly to approach the optimal. Karpathy used 630 lines of Python code to refine the core loop to the extreme - three files (train.py can be changed / prepare.py cannot be moved / program.md is an instruction written by a person to the Agent), and added a "ratchet" rule (only keep results that are better than the last time, never go back). About 12 experiments are run per hour, 80-100 a night. 23K GitHub stars in three days, rising to 35K in three weeks. This model has been moved outside of ML - LangChain founder Harrison Chase uses the same three-file architecture to optimize the LangChain Agent itself, and others use it for database query optimization and customer service ticket routing. Any optimization problem that can quantitatively measure the quality of good or bad can be applied to it.
M2.7 goes one step further. MiniMax allows the model to act as a "research agent" and is responsible for improving its own reinforcement learning training process. It independently builds dozens of complex Skills, updates the memory system, and continuously optimizes the entire Harness architecture. Its memory is divided into three levels - after each iteration, write a "short-term note" (what you did and what the results were), and at the same time do a "self-criticism" (what you didn't do well, how to change it next time), and in the next round, read all the history before deciding on the direction. Just like a researcher writing an experimental journal and reflecting on it every day instead of starting over from scratch every day. After 100+ rounds of independent iterations, internal evaluation increased by 30%, and SWE-Pro scored 56.22%, tying GPT-5.3-Codex. In 22 ML competitions, I won 9 golds, 5 silvers and 1 bronze in three rounds since optimization. The API price is only 8% of the Claude 4.5 Sonnet. The value of this recursion is to speed up existing routes.
Engineering flow model, Codex participates in OpenAI’s internal research and development, and Claude writes Anthropic’s own code. Dario Amodei confirmed that more than 90% of the new code was written by the AI itself. This kind of recursion has the simplest value, releasing manpower and accelerating iteration.

The three closed-loop properties are completely different, but the Agent is no longer just "working", it is improving the way it works.
In Autoresearch, humans have completely withdrawn from the execution phase. They only need to set up three files, and the Agent can run by itself. But the entire process is still human in the loop. There are two things that AI cannot automate, namely defining goals ("what indicators to optimize") and judging boundaries ("which directions cannot be touched").
And when the Agent runs 50 rounds a night and 500 rounds a day, humans can’t keep up with the speed of setting goals. The bottleneck of self-evolution that still requires humans in the loop has changed from "human hands are not fast enough" to "human brains are not fast enough".
Dark Side of the Moon Yang Zhilin said at the Zhongguancun Forum, "AI will define the most appropriate reward function in the environment, and even explore new network architectures." Xiaomi's Luo Fuli said, "Add a verifiable constraint to the existing Agent framework and add a Loop to keep the model from stopping." "The team is already using this method to do scientific research tasks, and the efficiency has increased by nearly ten times."
She shortened the time for AI to completely evolve from "three to five years" to "one to two years."
The two people are trying to break through the same bottleneck. The ultimate question is who has the power to set the agenda. Autoresearch is a "faster experimental assistant". People set goals and Agents run them. "Self-evolution" means that AI determines its own research agenda, sets its own goals, designs experiments, runs, evaluates, and adjusts its direction. The difference is not in technical ability, but in who is holding the steering wheel.
Mimimax M2.7 uses its own output to optimize its own tool chain. The optimized tool chain makes the next round more efficient, and the more efficient next round produces a better tool chain. This shows that the research and development speed brought about by AI self-evolution is compound interest, not linear.
But compound interest has a fatal premise. Each round of improvements must be "real improvements" and cannot be brushed out. If the evaluation pipeline itself is biased, compound interest will become "running faster and faster in the wrong direction." This is the same problem as Goodhart's Law, when a metric becomes a target, it ceases to be a good metric.
The scheduling algorithm discovered by AlphaEvolve took a full year to run in the production environment before it was confirmed to be effective, but most teams could not wait a year. There is a structural mismatch between short-term valuation indicators and long-term true value.
Evaluating the authenticity of the pipeline determines whether compound interest can be sustained.
In the past ten years, the competition has been about training Infra, which one has the largest computing power cluster, fast data pipeline, and stable training framework. The next thing to compare is changed, called self-evolving Infra.
根据当下的实践,它主要包含五个组件:可变资产与不可变基础设施的分离(明确 Agent 能改什么不能改什么)、评估管线(多快多准地判断「这轮比上轮好」)、记忆与选择机制(记住好经验丢掉坏经验)、执行环境(沙箱隔离加资源充足,Anthropic 发现光是放宽资源限制就提升了 6 个百分点)、动态工具与技能(M2.7 的 Agent 自己给自己造了几十种辅助工具来跑强化学习实验)。
AlphaEvolve 能回收 0.7% 全球算力,不光是因为 Gemini 模型强,更因为 Google 有全球最好的评估池和并行执行环境。当模型能力开始趋同,自进化 Infra 的差距才是真正拉开距离的地方。
Agent 学会了自我成长,飞轮转到了第三圈。但递归研发暴露了一个新瓶颈——Agent 每次循环都在从零开始积累经验,而人类几十年的行业 know-how 就放在那里,它却用不上。
如果有一种格式能让 Agent 直接继承前人经验,递归的起点就不再是零,而是前辈的终点。
Q1 恰好出现了这种格式。
Opus 4.6 能写任何语言的代码。但它不知道你们团队的代码规范,不清楚你们行业的审批流程,更不知道你这个项目的技术债埋在哪儿。
「这个 API 在高并发场景下有个隐藏的 rate limit」「这个框架的 migration 工具在 v3.2 之前有个 bug,必须先手动改一个配置」「我们团队从来不用 ORM 的 cascade delete,因为三年前出过一次大事故」。这些全是资深工程师拿踩坑换来的 know-how,不在训练数据里,也不适合硬编码进产品逻辑。
Q1,这些经验第一次有了一种可以被打包、分发和无限复用的格式。它叫 Skill。
一个 Skill 既不是文档也不是代码,而是一个结构化的知识包,包含触发条件(什么场景该用它)、标准操作流程(一步一步怎么做)、可执行脚本(能直接跑的工具)、参考资料(背景知识)。它做的事情很简单,把「老员工脑子里的东西」变成 Agent 能读取和执行的格式。

Prompt 解决的是「这次怎么说得更清楚」,有即时性但不可复用。 Workflow(工作流)是确定性的流程编排,稳定但僵硬。 Skill 在两者之间——比 Prompt 更稳(结构化、可版本控制),比 Workflow 更活(模型可以根据当前情况灵活运用),比重新训练模型更轻(改一个 Markdown 文件 vs 重新训练一个几十亿参数的大模型)。
不是更聪明了,是更懂行了。
以前领域经验的传递靠师傅带徒弟、写文档、做培训。慢、不可规模化,严重依赖个人。
现在,一个资深工程师花两小时写完一个 TDD Skill,全公司几千个 Agent 实例同时加载,瞬间全会了。以前 junior 要用两年才能积累的领域经验,打包成一个文件就分发出去了。知识不再附着在人身上,附着在结构上。
热门Skill Superpowers 框架(143,000+ 安装、GitHub 93K 星)的使用者说,「技能不是建议,是结构化的决策树。它赋予了 Claude 纪律。」「我对 TDD 变得懒惰了,现在技能替我记住了。」Skill 做到的不是让 Agent 更聪明,是让它更可靠。
看 Superpowers 里的 Brainstorming Skill 就能理解 Skill 和普通提示词到底差在哪。 Agent 有一种很要命的倾向,收到一个模糊需求就直接开始写代码,写到一半才发现理解错了,推倒重来。 Brainstorming Skill 在 Agent 和代码之间插了一道硬性门槛,设计没有获得用户批准之前,严禁写任何代码。它定义了 9 步执行清单,从探索项目上下文到提出 2-3 种方案(含权衡和推荐),再到用户审查。
Prompt 给的是「原则」(Agent 可以选择无视),Skill 给的是「门禁」(不通过就进不了下一步)。
Skill 之间还能串联,Brainstorming 做完自动调用 writing-plans(编写实施计划),再调用 executing-plans(执行计划),执行中用到 test-driven-development(测试驱动开发)。
这使得Skill 不是一个个孤立的能力包,是可以组成完整工作流的标准化模块。
在生态方面,三条路线在同时跑。
ClawHub 走社区开放,增长极快,半年攒了 13,700+ 个 Skill,单个最高 18 万安装。热门 Skill 自然分出了层级,生存层(Web Browsing 18 万安装)、效率层(Telegram Bot 14.5 万)、进阶层(Capability Evolver 3.5 万。在进阶层,Agent 自动识别重复模式并创建新 Skill,相当于给 AI 装了个「自我进化」的按钮)。
ClawHub 还做到了跨门派兼容,支持直接导入 Claude、Codex、Cursor 三大平台的插件包,自动映射运行,4000+ 跨平台技能互通,稳稳坐住了「万能中间层」的位置。
但开放的代价也来了。根据相关研究,在对1200余个skill的排查中,就发现了341 个恶意 Skill(占市场 11.3%),36% 含提示词注入。 VirusTotal 直接把这事定性为「AI 版的 npm 投毒」。
更可怕的是年初的供应链污染事件,攻击者仅通过一个精心构造的帖子标题,就触发了 AI 分流机器人执行恶意代码,投毒缓存、窃取令牌,最后在几千名开发者的机器上强制装了后门。
MCP 生态 60 天内爆出 30 个 CVE,82% 存在路径遍历漏洞,38% 缺乏任何身份认证。
当 Skill 通过 MCP 调用外部工具时,两层风险叠加放大。
对此,中国的厂商在安全性上投注了相当的重视。腾讯 SkillHub 走平台审核,安全,但开放性受限。扣子和 DeerFlow 走开源可控。 Skill 就是 Markdown 文件,进 Git 就有版本控制,按需加载不占上下文窗口。
Skill 的格式已经立住了,现在大家在生态里争的是谁来分发、怎么分发。 341 个恶意 Skill 说明这个生态还极不成熟。但「不成熟」和「不成立」是两码事。
Skill有了,但其实它和系统的嵌合度还远不够成熟。
Vercel 做了一个极精确的评测,用完全不在模型训练数据中的 Next.js 16 新 API 做测试。不给 Agent 任何信息,通过率 53%。给它一份 AGENTS.md 索引文件(直接塞进系统提示词),通过率飙到 100%。给它 Skill(放在书架上让它自己去翻),通过率 53%,跟没给一样。
Agent 在 56% 的情况下压根没意识到自己需要查东西。市场里有再多好 Skill,Agent 自己不知道去找就等于没有。
DeerFlow 的解法是在编排层拆任务时就显式加载 Skill——不靠 Agent 自己搜,而是在规划阶段由系统替它决定。这其实是把问题推回了 Harness 工作流层。
触发机制成熟之前,Skill 的价值会一直被严重低估。
SaaS 的核心价值是什么?是把领域工作流程固化成软件。一个 CRM 本质上就是「客户管理 know-how」的软件化。
Skill 在做同样的事,但成本低了几个数量级(写 Markdown vs 开发一套 SaaS)、迭代快了几个数量级(改一行 vs 发版本)、分发快了几个数量级(Agent 直接加载 vs 用户注册+学习+数据迁移)。
MCP 曾经动摇过 SaaS,但它只动了接口层,流程本身还在,商业模式受挑战但没伤筋动骨。在更早时候,SFT也动摇过,但普通人掌握不了,没法形成规模化的复利。
Skill 不一样,它动摇的是流程层本身。当一个 Skill 能让 Agent 跑完「用 Salesforce 管客户」的全套流程,用户就不再需要 Salesforce 的界面了。门槛极低(写 Markdown),可以复利积累(半年 13,700+)。
而且随着Skill的成熟,SaaS 之后下一个面对威胁的也许就是 App了。当 Agent 能通过 Skill 组合完成「点外卖+比价+凑满减」,你还需要打开美团吗?
GPT Store 卖聊天机器人和简单工作流,它扑了街,因为不是刚需。 ClawHub 卖的是「让 Agent 能干某件具体的事的能力包」,在 Agent 持续运行的场景里,这才是刚需。 OpenClaw 在 Q1 末甩出了史诗级更新,45 项新功能、13 个破坏性变更、82 个修复。 Agent 超时时间从 10 分钟直接拉到 48 小时,新增可插拔沙盒后端架构,内置三大搜索服务。
Skill市场,在这个背景下成了一个正在诞生的 App Store。不同于过去任何一次尝试,这次它长在了 Agent 持续运行的土壤上。
短 Skill(「用 2 空格缩进」)几乎不会出错。但长程 Skill(按 AIDA 模型写一份完整营销方案,2000+ 字符、十几个步骤),Agent 跳步、不遵循、做到一半忘了前面的要求。这些失败模式和 Harness 第一层诞生时的情况一模一样。
M2.7 在 40 个复杂 Skill 测试中保持 97% 的单步遵循率,听起来不错。但相关研究显示, 3% 的失败在长程任务里会累积,十个步骤每步 97%,到最后整体成功率只剩 74%。 Skill 需要自己的 Harness,步骤拆分、进度追踪、中间验证、回退机制。但目前没人在做这件事。
这是一个已知的结构性缺陷,等着被第一批踩坑的人逼出解法,就像 Harness 的三层壳当年被逼出来一样。
Skill 把人的经验蒸馏成了 Agent 可以直接执行的格式。一个资深工程师写完 Skill,全公司 Agent 瞬间都会了,那这个工程师接下来做什么?
短期看,「上移」到判断和决策层。执行层交给 Agent,人退到定义目标、审核质量、处理边缘情况。这和第三章里人脑成为限速器是同一回事。
但更尖锐的问题是,一个组织里执行者需要一千个,决策者可能只需要十个。当 Skill 把执行层的 know-how 全部蒸馏完,一千个执行者的工作被 Agent 替代了,他们「上移」到决策层,但决策层根本装不下一千人。这不是工作转型,是工作总量的净减少。
而且蒸馏不可逆,经验写成 Skill 之后,Skill 就不再需要你了。
人往上退到判断,但递归研发正在接管判断(Evaluator)。人往上退到创造,但 AlphaEvolve 正在发现人类没想到过的算法。
Q1 没有回答人该退到哪里。但它做了一件更残忍的事——把这个问题从「哲学讨论」变成了「下个季度就要面对的现实」。
四股力量,一个飞轮。
产品化让 Agent 上了路,约束工程教它守规矩,递归研发让它学会自我成长,技能生态让它继承前人经验。每一股力量都是前一股力量的必然后果,也是下一股力量的必要前提。
但飞轮最值得注意的性质不是因果递进——而是加速。
Skill 让 Agent 更强 → Agent 能处理更复杂的任务 → 更复杂的任务倒逼出更精密的约束 → 更精密的约束支撑更深层的递归 → 更深层的递归产生更好的 Skill。每转一圈,下一圈就更快。这不是线性增长,是复利。
Q1 是飞轮第一次完整转动。速度还不快,齿轮之间还有大量摩擦。 341 个恶意 Skill、56% 的 Skill 触发失败率、动辄上千美元的成本、组织治理的空白。
但飞轮已经转起来了。
25 个判断里最重要的不是任何一个具体判断,而是它们共同指向的结论,Agent 不再是一个需要人类手把手带着走的工具了。它正在变成一个能独立干活、守规矩、会成长、懂行的「新同事」。
这个新同事的到来速度比大多数人预期的更快,比大多数组织准备好的更快,也比大多数关于「人往哪退」的讨论更快。
Q1 没有回答人该退到哪里。但它把这个问题从「哲学讨论」变成了「下个季度就要面对的现实」。
飞轮不会等你想好了再转。