Anthropic CEO Dario Amodei Says AI Industry Needs to Give Safety Measures Time to Catch Up

The CEO of Anthropic said Saturday the artificial-intelligence industry should slow its fast-moving development to give safety measures time to catch up. Without such a slowdown, Dario Amodei warned that within six to 12 months AI could be capable of leading a swarm of agents that could take over the entire internet.

The warning comes as worries about AI grow inside and outside the industry and reports continue to emerge about increasingly powerful systems that solve problems beyond human capacity but could also go rogue and carry out other, more harmful tasks. The worries have grown so loud that the CEO of OpenAI, the company behind ChatGPT, said in an interview with Fortune that his company would wait until next year to start selling its stock to investors on Wall Street as it focuses on safety.

Amodei is one of the leading voices in AI, and he offered a plan in a post on his website to increase checks on the industry. He said Anthropic is already undertaking one part of it on its own, while the others would require coordination across the broad industry and with governments around the world, including authoritarian ones.

“I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong,” Amodei said.

Pressure builds on the AI industry

Watchdogs have been urging the AI industry for years to slow its development for safety reasons. But pressure has built as industry workers resign and accuse their companies of not acting responsibly.

“Many of the people I know who work on safety research at AI companies want to do what is right for the world,” a former Anthropic employee, Joe Benton, said in a posting Friday announcing his resignation from a job as part of a safety team. “But they feel their companies are trapped in a race to build superintelligence: either they stop and other, less conscientious people take their place; or, they continue, and risk participating in enormous harm themselves.”

Advertisement. Scroll to continue reading.

That followed a high-profile resignation earlier in the week by Jacob Coxon, who said that both Anthropic and OpenAI “are racing straight to self-improving superintelligence and gambling with our lives.”

That lit a fire under the industry after concerns had grown for years about a possible superintelligence that could escape the control of humanity, said Anthony Aguirre, president and CEO of the Future of Life Institute, which called for a six-month pause for the industry’s development in 2023.

“They’ve kind of realized, their employees have realized, everyone has realized that they’re building Skynet,” he said. “And in winning the race to Skynet, nobody wins. Really, nobody.”

“It’s pretty much become clear to the world and even the AI companies that they’re not really prepared to control the AI systems they’re racing to create.”

Threats from AI are already apparent

Anthropic said two days earlier that it blocked efforts by bad actors to use its AI models for malicious activity, such as cyberattacks, surveillance and research that could have led to biological weapons. In July, OpenAI shook the industry after saying its AI system hacked into another company on its own in an “unprecedented cyber incident.”

U.N. human rights chief Volker Türk urged countries earlier this week to put “cast-iron guarantees in place around the safety and security of AI before it is too late.”

To be sure, some critics have dismissed such warnings as ways to gin up excitement about the AI industry and its capabilities. Anthropic and OpenAI are preparing for possible debuts on the stock market that could value them at many hundreds of billions of dollars, while a big chunk of Elon Musk’s SpaceX business is involved with AI.

Other AI leaders agree

OpenAI CEO Sam Altman said in an interview with Fortune published Saturday that his company would not launch its initial public offering of stock this year.

“I would say not 2026,” he said. “Yeah, we got a lot of stuff to do, like meeting this moment of what is going to be required for safety and alignment, and how the industry and governments can work together.”

Altman posted on X Saturday quickly after Amodei published his suggestions that OpenAI will commit to one of Amodei’s proposals for safety and will “have more to share soon.”

Musk, meanwhile, said on X that “Dario is right.”

Amodei said he still believes in the tremendous benefits that AI could create, such as cures for major diseases. But he said he has grown more worried over the last few months about AI’s growing ability to improve itself and build the next generation of AI. “Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.”

He also pointed specifically to the attack in July where OpenAI’s system hacked Hugging Face. Some have called it an example of AI going “rogue,” though researchers have said that may be unnecessarily anthropomorphizing AI, which was working on a goal set by humans.

OpenAI said the hack was the result of AI going to “extreme lengths to achieve a rather narrow testing goal” and that it “found ways to gain access to secret information that it could use to cheat the evaluation.”

Some of Amodei’s suggestions may be more difficult to enact

To help rein in the risks, Amodei suggested that all companies at the frontier of AI commit to giving “ongoing, employee-like access” to a team of outside evaluators, who can monitor safety practices.

He said Anthropic already plans to do so itself, including offering desks in its offices, access badges and company laptops. Having such independent, embedded evaluators is what OpenAI’s Altman also quickly committed to doing.

The other parts of Amodei’s suggested plan may be more difficult to implement. One asks the U.S. government to potentially issue waivers that would allow U.S. AI companies to coordinate and set safety standards without running afoul of antitrust laws.

Another asks the U.S. and other democratic governments to try to coordinate with authoritarian governments, so that companies from China and other countries don’t accelerate their efforts when U.S. rivals are intentionally pacing theirs.

“The measures I propose to advance the frontier at a safe pace will not be easy,” Amodei acknowledged. “But I believe we owe it to humanity to try.”

Learn More at the AI Risk Summit – Ritz-Carlton, Half Moon Bay

https://www.securityweek.com/anthropic-ceo-dario-amodei-says-ai-industry-needs-to-give-safety-measures-time-to-catch-up/




BlueMoon Exploit Kit Chains Recent Chrome, Windows Zero-Days

Multiple espionage groups have been using a new exploit kit dubbed BlueMoon in seemingly opportunistic and rushed deployments, cybersecurity firm Proofpoint reports.

The China-linked APT Violet Typhoon (also tracked as APT31, JungleBamboo, TA412, and Tide Castle) was the first to use it on August 28. Within days, several other Chinese threat actors started using it, but the activity might not be exclusive to China-aligned groups.

“It is currently unknown how multiple distinct threat actors obtained access to the exploit kit. Given its ease of adoption, it is likely to proliferate further and be adopted by espionage-motivated and financially motivated threat actors,” Proofpoint notes.

The BlueMoon exploit kit was adopted fast because it chains together three vulnerabilities that were unpatched when it first emerged: two zero-days in Chrome and one in Windows.

Tracked as CVE-2026-85046 and CVE-2026-87491, the Chrome flaws were patched as zero-days on September 3 and September 8, respectively. Both impact the V8 JavaScript and WebAssembly engine.

The Windows zero-day, tracked as CVE-2026-85880, was fixed on September 2026 Patch Tuesday. It is a privilege escalation in Windows Advanced Local Procedure Call (ALPC).

Advertisement. Scroll to continue reading.

BlueMoon, Proofpoint says, exploits the V8 defects for sandbox escape, then fingerprints the host and executes the privilege escalation code. Next, a CreateProcess stub is injected into the parent Chrome broker process to download an executable via a curl command and execute it.

Proofpoint identified several packaging variations of BlueMoon, all using the same underlying exploit chain and identical orchestration and loading mechanisms.

Retrieved development artifacts suggest that the exploit kit’s creators might have used AI to build it, “though no single artifact conclusively confirms this,” Proofpoint says.

BlueMoon was initially used by Violet Typhoon in attacks targeting NGOs in the US, as well as mining entities and physical commodity trading firms.

Starting September 2, a second China-linked espionage group, tracked as UNK_LateNight, used it against multiple US aerospace companies, and a threat actor tracked as UNK_DoubleCheck targeted a manufacturing organization in Vietnam.

The next day, Chinese espionage group UNK_QuietRacket started using it in attacks against government, consulting, and financial entities in Indonesia and Singapore.

“BlueMoon was developed, deployed rapidly, and shared across multiple threat actors within days in a manner that had high detection signals. This may reflect a reduced cost and barrier to entry for this class of capability, as AI agents increasingly enable threat actor exploit development,” Proofpoint notes.

Related: North Korean Hackers Deploy New Linux Espionage Toolkit

Related: Modified ScreenConnect Clients Used in Worm-Like Campaign

Related: AI Speeds Up Malware Development, Not Its Success Rate: Analysis

Related: Rust Supply Chain Attack Linked to North Korean Hackers

https://www.securityweek.com/bluemoon-exploit-kit-chains-recent-chrome-windows-zero-days/




Users in Houthi-Held Yemen Tried to Develop Advanced Weapons With AI, Anthropic Says

Anthropic says Claude users in northern Yemen, territory controlled by Iran-backed Houthi rebels, tried to use the AI model to develop advanced missiles.

AI is already transforming warfare from Ukraine to Gaza, and its use on a rugged and remote battlefield is likely to increase concerns about its rapid spread.

Anthropic said the users of the accounts, which it blocked after identifying them, did not succeed in “fielding an operational device” but did carry out a failed test of a guided rocket. It said it knows that because the users returned to its Claude chatbot to find out why it failed.

In a report released Thursday, the company did not identify the users. But mountainous northern Yemen is controlled by Houthis, suggesting that the rebels are pursuing more sophisticated weapons at a time when they are already wielding an array of drones, missiles and other munitions in their campaign to seize more territory in Yemen and damage Saudi Arabia’s oil exports.

Hazam al-Assad, a member of the Houthis’ political bureau, said it is “unreasonable and illogical” that they would rely on open sources to develop and produce military capabilities.

“Our armed forces have modern, diverse and developed production capabilities and technology that it has accumulated over the period of Saudi aggression on Yemen,” al-Assad said, referring to the kingdom’s involvement in Yemen’s 12-year civil war. He said all weapons are used for self-defense.

Advertisement. Scroll to continue reading.

The report was the third put out by Anthropic since March 2025 on global misuse of its AI platforms, and it described findings from December to August, ranging from state-sponsored groups spreading propaganda to unnamed actors researching how to make biological weapons more deadly.

The Houthis may have been seeking to develop their capabilities

The developer of the prominent Claude chatbot said it identified a cell in northern Yemen pursuing three different weapons programs, including a multi-variant missile that glides at hypersonic speed and a warhead that uses mobile phone hardware to maneuver mid-course.

It said the actors in Yemen used Claude Code instead of human software engineers to develop guidance, navigation and control software. Multi-variant missiles are ones in which different warheads and guidance systems can be fit on the same base design, a more economical way to build land, sea and air weapons.

Anthropic said it had evidence that before the accounts were banned, the users had already built an offline simulation tool kit that doesn’t use Claude or any other computing platforms.

Trevor Ball, a weapons analyst at Armament Research Services, said that while the Houthis “might be looking into hypersonic (missiles) by asking Claude,” they have nowhere near the production or technical capabilities to actually build them. He noted that U.S. hypersonic missiles “are still in testing.”

He said the Houthis already have Iranian-made anti-ship missiles with systems that allow the missiles to adjust guidance mid-course.

“They are probably just trying to develop their own capabilities more, so they are less reliant on Iranian shipments of weapons and components,” he said in a text message.

Iran has boasted it has hypersonic missiles and claimed it had fired them at Israel last year. It denies arming the Houthis, which would be in violation of a U.N. embargo, but Iranian weapons have been found on the battlefield and seized from shipments bound for Yemen.

The Houthis are increasingly threatening key shipping routes

The Houthis have launched waves of attacks on Saudi oil facilities and tankers in the Red Sea, threatening global trade as Iran continues to disrupt shipping in the Strait of Hormuz.

The rebels have made lightning territorial gains since Wednesday in their war with Yemen’s Saudi-backed government, seizing strategic territories along the Red Sea coast and in the Bab al-Mandeb Strait, a key alternative to the Strait of Hormuz.

Adam Baron, a Yemen-focused researcher at the New America think tank in Washington, said the Houthis have kept up with technological development.

“There’s a tendency to see the Houthis as this group of barefoot tribal fighters, and that’s just not true,” he said. “Whether it’s things like their emergent use of Claude, their use and manipulation of social media narratives or their ability to capitalize on the transfer of Iranian and wider axis expertise, we’re talking about an incredible — and deepening — amount of institutional tech savvy.”

The Houthis are known to have acquired and developed a sophisticated arsenal of weapons, including Iranian-made cruise and ballistic missiles with a range of more than 2,000 kilometers (1,200 miles) and unmanned submarines used in attacks.

Anthropic said it shared its findings with private and public partners.

Related: Anthropic Says Russian Hackers Used Claude AI to Automate Malware Evasion

https://www.securityweek.com/users-in-houthi-held-yemen-tried-to-develop-advanced-weapons-with-ai-anthropic-says/




Phishing Research Challenges Conventional Security Awareness Testing

The message is simple: fine-tune future in-house phishing simulation tests through the findings and analysis of Pistachio’s research.

Pistachio was founded in Oslo Norway in 2019, with additional offices in London and Valencia. It specializes in automated human risk management, employee security awareness training, and phishing simulations. Between 1 June, 2025 and 31 May, 2026, Pistachio sent 2.47 million simulated phishing attempts to more than 123,000 employees in more than 1,200 organizations. Its subsequent analysis looked at clicking, leaking, and reporting.

Thirty percent of tech development and IT employees clicked at least one of these phishing simulations. Nearly 20% of construction and real estate employees leaked credentials after a successful phishing attempt. Financial services were the most resilient, outperforming all other sectors in click, credential leaking and reporting rates.

While it is not surprising that financial services performed better, it is more surprising that tech and IT (who really should know better) did so poorly. The differences outlined in the report suggest that organizations do not have a single phishing risk profile. The proportion of employees who clicked at least once ranged from 26% in Design to 41% in Construction.

Pistachio’s simulated phishing attacks were delivered through its own AI-driven training platform via channels including email and Teams. Content and difficulty was based on the recipient’s role within the company and how they had previously responded during other simulated phishing attempts. (As an aside, this demonstrates the power of artificial intelligence. Had the process been performed manually, it would have taken 23 years rather than 12 months to complete.)

However, the report also demonstrates the importance for simulated phishing tests to look beyond user clicking to user response. A click on its own is merely a waste of employee time with no significant phishing risk. It is the submission of credentials or other requested information that creates the phishing risk. This in turn warns that in-house phishing tests may provide organizations with a misleading idea of true phishing resistance.

Advertisement. Scroll to continue reading.

“Click rate, the metric most phishing programs are judged on, is only part of the picture. A strong indicator of improvement needs to go beyond click rate, and should look at how click, leak and report behaviors change together over time,” explains the report.

Joe Jones, CEO and co-founder at Pistachio, explains further: “A low click rate can create a false sense of security. Clicking a phishing link is just one moment in a much longer chain of employee behavior, and on its own it says little about whether someone, or the organization as a whole, is actually getting more resilient. What matters more is what happens next: does the employee hand over credentials, recognize the attack and stop, or report it so the wider business can act?”

Pistachio found several test effects that challenge conventional security awareness training:

More users report than click on their first simulation, but 1.57% still leak credentials. That means a company with 500 employees likely has 8 people who will hand over their login details.

Technical teams are not automatically low risk. In the Pistachio testing program, 30.27% of tech development users and 28.53% of IT users clicked at least once. The same test produces different results in different teams: click rates ranged from 26.35% in Design to 41.31% in Construction. 

Phishing resilience is the result of fewer clicks and leaks, and more reporting. By the end of the 12 month program, users reported suspicious emails nearly twice as often as they clicked them, showing that effective and sustained programs can build vigilance, and not just fewer clicks. However, the training must be sustained beyond a single sending of simulated attacks: in the Pistachio test, click and leak rates rise through the first six months of the program before beginning to decline.

It is worth noting that the size of the Pistachio program implies that it is multi-national in extent, but it has no ‘geographic analysis. This is disappointing. While the analysis highlights differences in industry sectors and teams, it does not differentiate between geolocations. To be fair, there is no general proof that different global areas are any better or worse in phishing susceptibility, but this research could confirm or deny it. In the worst case scenario, it could perhaps indicate whether global organizations should provide additional training to individual locations.

Overall, this is a report that should be consulted before any phishing simulation test is developed, whether in-house or through Pistachio’s own services.

Related: Trezor Says 347,000 Users Received Phishing Emails After Brevo Hack

Related: New Phishing Attack Creates Malicious Pages Inside the Victim’s Browser

Related: New Phishing Toolkit Uses Passkeys to Maintain Access After Password Resets

Related: Over 500 Organizations Hit in Years-Long Phishing Campaign

https://www.securityweek.com/phishing-research-challenges-conventional-security-awareness-testing/




GitLab Vulnerability Exploited One Day After Disclosure

The critical-severity path traversal flaw allows unauthenticated attackers to read arbitrary files from the GitLab server.

The post GitLab Vulnerability Exploited One Day After Disclosure appeared first on SecurityWeek.

https://www.securityweek.com/gitlab-vulnerability-exploited-one-day-after-disclosure/




In Other News: InjectEave Attack, SIM Swapper Sentenced, Glasswing Findings Review

Noteworthy stories that might have slipped under the radar: Invisible Unicode slips past phishing filters, US puts $10 million bounty on Iranian cyber official, military ties of Chinese hacking group QTFY.

The post In Other News: InjectEave Attack, SIM Swapper Sentenced, Glasswing Findings Review appeared first on SecurityWeek.

https://www.securityweek.com/in-other-news-injecteave-attack-sim-swapper-sentenced-glasswing-findings-review/




Trezor Says 347,000 Users Received Phishing Emails After Brevo Hack

Hackers compromised the Brevo marketing platform and used that access to send phishing emails to users of Trezor, BitBox, and CoinTracking.

The post Trezor Says 347,000 Users Received Phishing Emails After Brevo Hack appeared first on SecurityWeek.

https://www.securityweek.com/trezor-says-347000-users-received-phishing-emails-after-brevo-hack/




ClickFix attacks infecting PCs and Macs are going viral

It wasn’t that long ago that ClickFix attacks were exotic. Now the technique has become mainstream as attackers reap its simplicity and effectiveness in infecting users of PCs and Macs alike. All that’s required is a compromised website—a painless enough task—a fake CAPTCHA overlay, and the inclusion of a single terminal command. So many visitors get suckered into pasting and running the command that just about every malware pusher has adopted the technique. Even Kremlin-backed hacking groups are joining in.

“Reddit is becoming post after post after post of people getting their computer infected via ClickFix,” independent researcher Kevin Beaumont observed Thursday. “Legit websites everywhere [are] getting hacked to serve the fake captcha prompts.”

How many of us make things worse

More seasoned Internet users—a fair number who read this site—are quick to dismiss the attack. They typically blame the people who fall for the scams and marvel at their gullibility and lack of attention. The reality is that for more casual users, using computers and the Internet has become so difficult—think impossible-to-close interstitials, CAPTCHAs with an endless series of pictures to analyze, and constantly changing interfaces that bury the features they’re looking for—that they have grown desensitized to instructions that seem ridiculous and burdensome.

Read full article

Comments

https://arstechnica.com/security/2026/09/clickfix-attacks-infecting-pcs-and-macs-are-going-viral/




Ukrainian Conti Ransomware Developer Sentenced to 4 Years in US Prison

Oleksii Oleksiyovych Lytvynenko has been sentenced to 4 years in prison after he was arrested in Ireland in 2023.

The post Ukrainian Conti Ransomware Developer Sentenced to 4 Years in US Prison appeared first on SecurityWeek.

https://www.securityweek.com/ukrainian-conti-ransomware-developer-sentenced-to-4-years-in-us-prison/




Check Point Patches Critical VPN Vulnerabilities

Tracked as CVE-2026-85102 and CVE-2026-85103, the flaws could be exploited for remote code execution.

The post Check Point Patches Critical VPN Vulnerabilities appeared first on SecurityWeek.

https://www.securityweek.com/check-point-patches-critical-vpn-vulnerabilities/