Claude Opus 4.7 Late Night Blast! Competent for longer tasks, autonomous checking, and pulling full visual capacity
Yesterday night.AnthropicRelease of the new generation of flagship big model Claude Opus 4.7.

▲Anthropic releases new model Claude Opus 4.7 (Source: X)
The model is inSignificant improvements in advanced software engineering over Opus 4.6, especially in handling the most complex tasks; high-resolution image processing capabilities have been significantly improved, theMore than three times the size of the previous Claude model; in addition, Claude Code has synchronizedAdded /ultrareview code review commandThe review session is initiated after the input and the code changes are checked line by line.
User feedback states that they can confidently leave the most difficult coding jobs to Opus 4.7, which handles complex long-running tasks rigorously and consistently, follows instructions precisely, and validates the output itself before reporting results.
Opus 4.7 is live today across all Claude products and APIs, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. Pricing is consistent with Opus 4.6: $5 (~Rs. 34) per million tokens input and $25 (~Rs. 170.5) per million tokens output. Developers can use claude-opus-4-7 via the Claude API.
I have to say, Claude has been updating really fast lately, people can't keep up, netizens swiped emojis under Claude's comment section, “Two eyes open, Claude updated again.”

▲ Netizens comment on Claude's tweet (Source: X)
I. Stricter execution of commands and enhanced multimodal support
In testing, Claude Opus 4.7 significantly outperformed Opus 4.6 in the following areas:
1,Instructions to follow.Opus 4.7 is a significant improvement in following instructions. Whereas previous models would interpret instructions loosely or skip parts altogether, Opus 4.7 executes instructions literally. Users should retune cue words and application frames accordingly.
2,Multimodal support enhancement.Opus 4.7 is more capable of visualizing high-resolution images: it can accept images up to 2,576 pixels (about 3.75 million pixels) on the long side, which is more than three times as much as the previous Claude model. This opens up a wide range of multimodal applications that rely on fine visual detail: for example, recognizing dense screen shots when operating a computer with an Agent, extracting data from complex charts, and design work that requires pixel-level precision.
3,Practical work.In addition to achieving optimal scores on the Financial Agent review, Anthropic internal tests show that Opus 4.7 is a more effective financial analyst than Opus 4.6, producing more rigorous analyses and models, more professional presentations, and achieving tighter cross-task integration.Opus 4.7 achieves optimal scores on the third-party economic value of knowledge work in the areas of finance, law, and other areas of the review GDPval-AA on which it also reached the optimal level.
4,Memory skills.Opus 4.7 is stronger in its use of file system-based memory. It remembers important notes during long, multi-session work and uses those memories to advance to new tasks, reducing the need for front-loaded context.

▲Opus 4.7 model benchmark performance (Source: Anthropic)
Opus 4.7 has received positive feedback from some of its early testers. Clarence Huang, vice president of technology at financial software company Intuit, says the model finds logic errors on its own during the planning phase and executes much faster than its predecessor.AI Programming ToolsIgor Ostrovsky, CTO of the company Augment Code, believes that the strength of Opus 4.7 lies in the fact that it handles automated processes, CI/CD (Continuous Integration and Deployment), and long task processes in a practical way and gives its own judgment, rather than attaching itself to the user.
Second, a number of assessment leading, biological reasoning, document reasoning to improve significantly
Anthropic evaluated Opus 4.7 in a pre-release test for different domains and compared Opus 4.6, GPT-5.4 and Gemini 3.1 Pro.

Biological reasoning has improved the most, Opus 4.7 scored 74.01 TP4T and Opus 4.6 only 30.91 TP4T, a 1.4x improvement.

Document reasoning, the Opus 4.7 scored 80.61 TP4T, well ahead of the Opus 4.6's 57.11 TP4T, and well ahead of the GPT-5.4 (51.11 TP4T) and the Gemini 3.1 Pro (42.91 TP4T), making it one of the most clearly disparate programs in the crossover.

Also.Aspects of knowledge work, Opus 4.7 ranked first with an Elo score of 1,753, a clear lead over GPT-5.4 (1,674), Opus 4.6 (1,619), and Gemini 3.1 Pro (1,314).

Aspects of Long Context Reasoning, when dealing with the simpler parent finding task (Parents 1M), Opus 4.7 scored 75.11 TP4T, while Opus 4.6 was 71.11 TP4T, a small difference; however, when dealing with the more difficult breadth-first search task (BFS 1M), Opus 4.7 scored 58.61 TP4T, while Opus 4.6 was only 41.21 TP4T, a pulling away by 17 percentage points. The more difficult the task, the more effective the model boost is.

existSecurity and alignment aspects, Anthropic also published the misalignment behavior scores for each model. the misalignment behavior score for Opus 4.7 is about 2.47 (out of 10, the lower the better), which is slightly better than the 2.75 for Opus 4.6, but still significantly different from the 1.78 for Mythos Preview.
Overall, Opus 4.7“s security performance is similar to that of Opus 4.6, with a lower percentage of behaviors such as spoofing, flattery, and cooperation with abusers.Anthropic commented on this, ”Opus 4.7 is generally well aligned and trustworthy, but the behavior is not entirely ideal." The Mythos Preview, which has the best alignment performance, is not yet fully open.
Third, other updates: new xhigh level, review order, task budget into public testing
In addition to Opus 4.7 itself, Anthropic has also rolled out several feature updates in parallel.
In terms of reasoning levels.Added xhigh (extra high) levelThe default inference level for Claude Code has been raised to xhigh, which is between the existing high and max, giving the user a finer margin of adjustment between inference depth and response speed.
API-wise.Mission budget functionGoing into the public beta, developers can guide Claude on how to allocate token consumption in long tasks.
For Claude Code.Add /ultrareview commandThe input launches a specialized review session that goes through the code changes line by line, and theFlagging Bugs and Design IssuesThe new system is a free experience for Pro and Max users, with 3 free experiences each. In addition, Auto mode is extended to Max users, which allows Claude to make autonomous operational decisions and reduce manual confirmation interruptions.
Fourth, beware of Opus 4.7 more costly token, but the generation quality is better
Opus 4.7 is a direct upgrade from Opus 4.6, but there are two changes that affect token usage that are worth noting.
firstlyText handling has been updated, Opus 4.7 consumes up to about 351 TP4T more tokens for the same inputs; second, the modelWill think more at higher reasoning levels, especially in subsequent rounds of the Agent scenario, Opus 4.7 output token will increase accordingly. Users can control the usage by adjusting the inference level, setting a budget for the task, or requesting more brevity in the cue word.

Looking at the Agent Programming Review charts, Opus 4.7 achieves higher scores at each reasoning level with fewer tokens. For example, Opus 4.7 consumes about 100,000 tokens at the xhigh level and scores over 70%, while Opus 4.6 consumes about 130,000 tokens at the max level and scores just over 60%.However, the models in this review work autonomously from a single prompt, and the results are not necessarily representative of actual token consumption in interactive programming.
Conclusion: More accurate and versatile, competitors will arrive
From the data published by Anthropic, Opus 4.7's improvement in several benchmarks such as programming, document reasoning, biological reasoning, etc. is real, and token efficiency has also been improved. But an evaluation is an evaluation, and the actual performance needs to be further verified in real scenarios.
With the release of Opus 4.7, what new moves will OpenAI make subsequently, and whether the long-awaited DeepSeek will release a new model at the end of the month, the competition among the big model vendors can get more and more interesting.
Source: Anthropic
© Copyright notes
The copyright of the article belongs to the author, please do not reprint without permission.
Related posts
No comments...