-
Cryptocurrencies
-
Exchanges
-
Media
All languages
Cryptocurrencies
Exchanges
Media
Share
Author: Silicon Valley Alan Walker
The press conference put the spotlight on SWE-bench, but the real signals were hidden in footnotes, introduction blocks, and an inconspicuous auto mode sentence.
On California Avenue in Palo Alto, the light at nine-thirty in the morning swept in from the glass window of the Coupa Café and shone on Alan Walker's half-cup of cold flat white. He had just finished browsing Anthropic's official website, leaned back in his chair, and spoke to Tony, who had just sat down across from him.
"Anthropic released Opus 4.7 this time, and the press conference was quite restrained - the protagonists were the SWE-bench pillars, the customer quote carousel, and a beautiful alignment chart. Most of the technology media copied the press release and left."
"But the real stuff is buried in footnotes, a migration guide, and an understatement like 'auto mode extended to Max users.' You have to read it like a 10-K - The main text is for retail investors, and the notes are for institutions. "
"Before I finish this cup of coffee today, I will take apart eight knives. I will tell you who it cuts with each one."

The press conference briefly mentioned: "In Claude Code, we've raised the default effort level to xhigh for all plans."
When most people see xhigh, they think it has "one more level", just like the iPhone has one more color. wrong. The real signal is the last half sentence - the default gear of all plans in Claude Code is pulled to xhigh.
This is a very Anthropic move: quietly, raise everyone’s baseline by one level, and then keep the computing power bill unchanged. It’s like giving you a smarter colleague, but without a salary increase.

TONY: Wait, doesn’t this mean that Pro users originally spent $20 to get medium, but now they can get xhigh directly?
ALAN: Yes. And please read the Hex quote carefully - "low-effort 4.7 ≈ medium-effort 4.6". The superimposed default level is increased, which is equal to the effective intelligence obtained by ordinary users, and jumps up to two full levels. The press conference did not capitalize this number because they did not want the token consumption page to look ugly.
◆ Landing scene
On Monday morning, you asked Claude Code to modify a 500-line back-end module - originally you had to manually type /effort max before it could run by itself; now you have nothing, the default is xhigh, and you can finish the work after a cup of coffee. The difference is not 10% faster, it's "you don't need to worry about it".
KILL LIST
→ "AI tuning/prompt configuration" SaaS - tools that teach you how to adjust your thinking budget and how to choose effort. The default values are automatically correct and the middle layer has no business
→Junior Engineer Position——xhigh’s default job is already the lower quality limit for engineers with three years of experience
→ Outsourcing code review company——The third knife below will kill this one
—— BLADE NO. 02
Footnote on the third line of the press conference: "Auto mode extended to Max users." Just one sentence.
Original words from Anthropic's official website: "auto mode is a new permissions option where Claude makes decisions on your behalf."——"Making decisions on your behalf."
In the past year, all agent startups have been in two extremes: either skip-all-permissions (the path of Devin and Cognition), or crazy pop-ups of approve/deny (the early days of Cursor). Anthropic took the third path: training the model to judge on its own what to ask and what not to ask, and internalizing this judgment into auto mode.

KAI: Alan, what is the essential difference between this and skip permissions? Aren’t you just letting it run?
ALAN: Big difference. skip, you pulled the safety bolt, you are responsible for anything that happens. Auto means that the model has a set of insurance installed on its own - it will actively stop and ask you for dangerous operations, and handle low-risk operations by itself. The essence is to move the entire "permission UI" layer from the product shell to the model weight.
TONY: So there are a bunch of YC startups doing "agent governance / guardrails"...
ALAN: The product is built into the model. This is a living example of what Andrej said last year, "the model is the product."
KILL LIST
→ Agent guardrails / approval-flow SaaS——Those who are doing "human-machine collaborative approval platform", the entire category has been reduced in dimensionality
→ RPA traditional industries (UiPath / Automation Anywhere) - Their core value is "controllable automation", and now controllability is endogenous
→ The middle and back offices of the BPO outsourcing industry - data entry, customer service dispatch, and invoice reconciliation in the Philippines and India, all in auto mode for a day, the work of a team
—— BLADE NO. 03
Official website wording: "a dedicated review session that reads through changes and flags bugs and design issues that a careful reviewer would catch."
Notice that word - "a careful reviewer". Not junior, not linter, but "careful reviewer". Translated into adult language: senior engineer.
CodeRabbit's David Loker gave a more direct figure: recalls increased by more than 10%, the most difficult bugs were dug out in the most complex PRs, and precision was almost lost. Recall increases, precision does not decrease - in the field of code review, this is holy grail. The last person to get this combination is Google's internal Tricorder, who has been working for ten years.

MARCUS: One of our FAANG staff members earns $800K a year, and reviewing PR accounts for half of the time. If this thing could really hit...
ALAN: Pro and Max users are given three free ultrareviews to give you a try. This is Silicon Valley's usual "freemium poisoning" routine - giving you a taste and then making it impossible for you to go back.
MARCUS: So this is not a tool, it is a stand-in.
ALAN: Not completely. It doesn't replace staff, it replaces the two hours staff spends reviewing ten PRs every afternoon. During the two hours of release, thesenior is the senior, not the human GitHub bot.
◆ Landing scene
In an engineering team of 20 people, the tech lead originally spent three hours a day reviewing PR. Go to /ultrareview, the tech lead only needs to look at the "design issues" marked in red by Claude - three hours becomes twenty minutes, and the time saved is really about architecture. This is not "AI assistance", it is a rewriting of job responsibilities.
KILL LIST
→ All independent AI code review startups—CodeRabbit, Codacy, Qodo, which are now Anthropic features
→ SAST / DAST traditional security scanning tools (Snyk / Checkmarx) - rule-driven static scanning, overwhelmed by "reading code like a human"
→India/Eastern Europe outsourced code review services—This market was valued at billions of dollars in the past ten years and has now evaporated
—— BLADE NO. 04
2,576 pixel vision - Computer-Use from demo to weapon
"Acceptable images are up to 2,576 pixels on the longest side, which is about 3.75 megapixels, more than three times the previous size."
This one is the most underestimated. When most people see it, they think, "Oh, it's more high-definition." How wrong. This is the watershed for the entire computer-use category from demo to production.
The evidence is in the quotation block at the bottom of the release page, what XBOW CEO Oege de Moor said——

54.5% → 98.5%. This is not a gradual improvement, it is a transition from "cannot use" to "cannot use". Opus 4.6 is still trying to guess where the buttons on the screen are, but 4.7 can already read the fine print and nested tables on the dense dashboard.
SARAH: Our enterprise customers have been stuck at this point. 4.6 Let it automatically process the scanned invoice, half the mistake - the boss just said "stop playing".
ALAN: The current figure of 98.5% means that RPA, IT operation and maintenance, reimbursement auditing, old system relocation - all workflows that still rely on human eyes to see the screen, for the first time have an acceptable supporting model.
KAI: computer use is no longer demo video, it is productivity.
ALAN: Yes, and please note - this is a model-level upgrade, not an API parameter. Old users don't change anything and get it automatically. Anthropic is quietly pushing the product capabilities of all integrators to the next level.
KILL LIST
→ OCR / Document Understanding SaaS (Rossum / Hyperscience / Nanonets) - Their moat is originally "visual + structural", and now it is equal to or even exceeded by the general model
→The three traditional RPA giants—UiPath’s screen recognition core technology, its value evaporated by half overnight
→Enterprise application data entry department—medical insurance claims, bank KYC, government form processing, the entire human fleshing assembly line
→ Independent penetration testing/red team industry——Companies like XBOW have reaped dividends, but traditional pentesting consulting services have been broken
—— BLADE NO. 05
File-System Memory——Anthropic chose the simplest path
A footnote from the press conference: "Opus 4.7 is better at using file system-based memory. It remembers important notes across long, multi-session work."
OpenAI uses "embedded memory" - the memory is buried in the model, and you can't see it or change it. Google is doing mysterious infini-attention. Anthropic shows its cards this time:The file system is memory. Claude writes .md notes and reads .md notes. You can cat them out at any time.
This choice may seem low-tech, but it is actually a victory of first principles. The core issue of memory has never been storage, but auditability, editability, and migration. Both vector databases and embedded memory violate these three points.

ERIC: What corporate customers fear most is "I don't know what this AI remembers about me."
ALAN: File system memory directly addresses compliance. GDPR right to erasure? rm. SOC2 audit? cat to the auditor. This is not a technical advantage, it is a legal advantage.
ERIC: So those startups working on "AI memory layer"...
ALAN: Mem0, LangMem, Zep - raised a lot of money this year. What they solve is that "the model itself does not care about memory". Anthropic has written this capability into the model and uses the simplest POSIX file system. Intermediate layers are skipped.
KILL LIST
→ AI Memory Infrastructure Startup (Mem0 / LangMem / Zep) - Value proposition is internalized into the model
→ Agentic memory usage scenario of some vector databases——A main narrative of Pinecone and Weaviate is affected
→ AI enhancement layer of enterprise knowledge management SaaS - No need for third-party middleware, Claude directly reads and writes project files
"Giving developers a way to guide Claude's token spend so it can prioritize work across longer runs." (public beta)
This was missed by all the media, but it is the most important engineering breakthrough of the long-range agent this year.
In the past year, all agent companies have been struggling with the same devil: tokens for long tasks have gone out of control. Give Devin or Cursor a complex task, and it will run on its own for two hours and come back to tell you that you burned $800 and only half of the work was done. The boss's eyes turned green when he saw the bill.
The design of Task budget is very clever - it is not a simple token upper limit, but allows the model to see the budget countdown by itself, and decide by itself which steps to skip and how to complete the task to the most critical degree.

CLAIRE: Isn't this the "minimum deliverable" thinking of engineering project management?
ALAN: Yes. Anthropic has trained the PM skill of scope-cutting into the model. Give you a budget of $10 to run the agent, and it will decide which function can be accepted if it achieves 80%, and which one must achieve 100%.
TONY: So Notion’s quote——"implicit-need tests" is the first to pass——
ALAN: That’s right. The model begins to have "resource awareness" and can guess what you didn't say but expected, and give priority to keeping it within the budget. This is to train "senior engineer judgment".
KILL LIST
→ AI cost-control / LLM observable entrepreneurship (Helicone / Langfuse cost module) - core functions are nativeized
→Agent orchestration framework (part of LangGraph / CrewAI usage) - The model can plan its own budget and does not require external scheduling
→ The project management part of the traditional consulting industry - the intelligence of "resource allocation + delivery tailoring" has been overtaken by models
—— BLADE NO. 07
Joe Haddad, Distinguished Eng at Vercel: "It even does proofs on systems code before starting work, which is new behavior we haven't seen from earlier Claude models."
This sentence was buried in more than 20 quotes, and no one amplified it. But the old OG put down his coffee immediately after reading this.
"proofs on systems code" - Before writing system-level code, the model will first do its own mathematical/formal proofs. This doesn’t mean smarter,This is the model starting to verify its own code in the same way as a PhD verification paper.

MARCUS: This behavior appears in the training data, indicating that Anthropic clearly rewards "prove first and then write code" in the RL stage.
ALAN: Yes, this is consciously trained. Combine Vercel's piece with Genspark's "loop resistance", and Hex's "correctly reports when data is missing instead of plausible-but-incorrect fallbacks" - and what you see is a complete taste training project:Let the model start to work like an engineer who is not easy to fool.
MARCUS: Not easy to deceive - meaning not to deceive oneself.
ALAN: Yes. Opus 4.7 no longer gives you a plan that looks workable in order to complete the task. This is a reflection of alignment’s practical implementation at the product level.
KILL LIST
→ Formal verification tool market segment (part) - Coq/Lean/TLA+ Some entry-level scenarios for these high-threshold tools, the model will help you do it yourself
→ High-frequency trading/blockchain security audit industry—The core work of auditors ("reading code to find invariant violations") has been model-collaborated, and the unit price of audits has been suppressed
→ Operating system kernel/embedded outsourcing - those segments that require proof-based reasoning, the threshold has been leveled
—— BLADE NO. 08
"During its training we experimented with efforts to differentially reduce these capabilities."
The most outrageous operation is here. Anthropic admitted that it actively reduced the network offensive and defensive capabilities of Opus 4.7 during the training process because the stronger Mythos Preview behind it was not released. Then——
Then they opened a Cyber Verification Program, which allows legitimate security researchers, pentesters, and red teams to unlock higher permissions after certification.

ERIC: Isn't this...a model version of export control?
ALAN: More precisely, "Capability KYC". The model has three levels of ability gates. Only by proving your identity can you unlock the corresponding level. For the first time, the window for regulatory arbitrage has been clearly priced by AI companies themselves.
ERIC: What does it mean for startups?
ALAN: First, for general "AI + security" entrepreneurship, if you want to do high-end scenarios, you must first obtain Anthropic certification, and the supply chain itself will be managed. Second, a whole new category will emerge: consulting services to help you pass Anthropic certification—just like the companies that help you pass SOC2 today. Third, this is the way Anthropic will release all frontier models in the future. The release of Mythos will only be more strict.
TONY: So companies like Palantir and Booz Allen with government compliance...
ALAN: Pick up a layer of moat for nothing. Theyalready had liquidation-level status, and now they naturally unlock the top model.
◆ Landing scene
A YC entrepreneur who wants to do AI pen testing, starting from Q2 of 2026, must answer "Have you obtained Anthropic Cyber Verification" on the first page of the business plan. No? VCs don’t invest. Got it? Multiply the valuation by 2. A certification, a watershed moment in the capital market.
KILL LIST & NEW TRACK
→ General Network Security Entrepreneurship SaaS - Without Anthropic certification, you cannot obtain upper-layer model capabilities, and the ceiling is locked
→ The new track of "AI model capability compliance consulting" is born—a number of intermediaries will emerge in the next 12 months to help enterprises with frontier model certification
→ Traditional military industry and government system integrators (Palantir / Booz Allen) - natural benefits, the threshold becomes a moat
→ Open source / local deployment camp - Llama, Qwen, DeepSeek routes have benefited instead, and "can be used without certification" has become a core selling point
The sun outside the California Avenue window has climbed over the roof of the Palo Alto Creamery, and its slanting light hits the glass.
"Eight knives, cutting in eight directions. Some tracks begin to die today, and some begin to live today."
"Every generation of frontier model is released, the real thing is not written on the Headline." He said to Tony, "The press conference is for analysts to see. The numbers in the footnotes and quotes are for us to see."
"Don't watch the excitement."