Anthropic has made Claude Opus 4.7 generally available today. Anthropic has pushed the update to public users. Anthropic says Opus 4.7 reasons more deeply, writes better code and accepts much larger images — up to 2,576 pixels on the long edge.

What Opus 4.7 brings

Anthropic has released Claude Opus 4.7 to the general public through its Claude AI interface, the Claude API and partner platforms including Microsoft Foundry, according to the company announcement and coverage by Timothy Beck Werth, Tech Editor at Mashable. Anthropic positions Opus 4.7 as the most capable Opus model available to the public; Mythos remains a more powerful, partner-only release. Anthropic says Opus 4.7 improves multi-step reasoning, advanced coding and visual comprehension compared with Opus 4.6.

You'll probably notice it's blunter: Opus 4.7 is more literal and follows instructions more precisely than 4.6.

There's a new 'xhigh' effort level between 'high' and 'max' — it spends more compute on reasoning but runs faster than 'max'.

Anthropic has also added public beta task budgets: developers can set token ceilings for jobs so a particularly deep or protracted run won't blow past a cost limit.

Anthropic admits there's a trade-off: deeper reasoning uses more output tokens, so teams will need to manage budgets. The company published guidance on optimising token usage for teams moving from 4.6 to 4.7.

Vision, memory and long tasks

Opus 4.7 raises the bar for visual intelligence. The model can now ingest images with a longest edge up to 2,576 pixels — roughly three times the resolution supported by earlier Opus releases. That should help users who feed large, detailed schematics, high-resolution interface mock-ups or product photos into the model during workflows.

Opus 4.7 improves file-based memory — it keeps context across multi-hour sessions, which helps with long coding or agent runs. If Opus is writing to a scratchpad or keeping notes in a document during a prolonged coding task, it's less likely to lose track of earlier steps. Developers who rely on sustained, stateful runs — such as autonomous agents or long-running code synthesis jobs — will find those improvements useful.

Benchmarks and third-party tests

Independent evaluations reported by Decrypt back up Anthropic's performance claims. On SWE-bench Multilingual, which measures coding capability, Opus 4.7 scored 80.5% against 4.6's 77.8%. That's a clear uplift, though not a wholesale leap.

On GDPVal-AA — a third-party measure of economically valuable knowledge work across finance and legal tasks — Opus 4.7 registered 1,753 Elo, outpacing GPT-5.4 at 1,674. Document reasoning saw the largest gains: OfficeQA Pro results climbed to 80.6% for 4.7 versus 57.1% for 4.6. Tests focused on long-context coherence, such as Vending-Bench 2, showed a higher sustained performance for 4.7 too; Decrypt reported a $10,937 simulated money-balance figure for Opus 4.7 compared with $8,018 for 4.6.

In short, the benchmarks show 4.7 wins when you hand it complex, multi-step tasks that play out over long periods — not short one-off queries.

Safety, cyber limits and access for specialists

Anthropic says it intentionally limited Opus 4.7's offensive cyber capabilities for public release and added automated safeguards to block high-risk requests. Opus 4.7 ships with automated safeguards that detect and block high-risk cybersecurity requests. The firm said it experimented with reducing cyber capabilities during training to lower the abuse surface.

Security teams and researchers can apply to a Cyber Verification Program for access to additional capabilities that aren't enabled by default. Anthropic describes that initiative as a test run for the kinds of verification and oversight it will need when rolling out more powerful models to trusted partners.

Anthropic also continues to draw a line between Opus and Mythos. Claude Mythos is a more powerful model that Anthropic judged too risky for broad release. Opus 4.7 sits below Mythos by design — a pragmatic choice aimed at giving organisations stronger tools while keeping certain capabilities gated.

How to try it and what it costs

Opus 4.7 is available now to anyone who uses Claude AI, to developers via the Claude API and through Anthropic partners such as Microsoft Foundry. Anthropic has kept pricing for 4.7 the same as Opus 4.6, according to the company's release and subsequent reporting.

The new default behaviour of spending more token budget on deeper reasoning means that teams will want to rethink how they control costs. Anthropic's task budgets let a developer cap token usage per job. If a complex task threatens to run away with costs, the system will stop it before it exceeds the limit.

Early user reaction and developer notes

The reaction from developers was immediate. A string of complaints about Opus 4.6's perceived regression earlier in the year helped set expectations for 4.7. Some on X (formerly Twitter) said 4.7 felt like "early Opus 4.6" — a return to form for users who felt 4.6 had been quietly adjusted down. One prominent user, Dev Ed (@developedbyed), posted a sarcastic welcome-back message when the announcement appeared.

Timothy Beck Werth at Mashable noted that Opus 4.7 handles complex, long-running tasks with more rigor and consistency than earlier Opus releases. That echoes Anthropic's own claims that users can now hand off more of their hardest coding work to the model with less supervision. But Anthropic also warned that prompts tuned for older models may require small adjustments to get the best results from 4.7.

Where Opus 4.7 fits in the market

Opus 4.7 arrives in a crowded field of large models. Decrypt's benchmarking showed Opus 4.7 leading several comparable models on particular tests, notably document reasoning and long-context tasks. That may give Anthropic an edge in enterprise workflows that need deep, sustained reasoning rather than flashier single-turn outputs.

At the same time, the decision to throttle or gate cyber capabilities and to keep Mythos restricted shows the industry's tightrope: firms want to ship capability but also limit misuse. Anthropic's approach is to iterate publicly while reserving a higher tier for partners and vetted researchers.

For teams building with Claude, the practical question is whether the gains in correctness, memory and vision outweigh higher token consumption. For some projects they will. But for others, especially at scale, costs will need close management.

Related Articles

Decrypt's tests show Opus 4.7 scored 80.6% on OfficeQA Pro, compared with 57.1% for Opus 4.6.