Menu

Categories

Tags

Anthropic Hacker Went from Blocking AI to Lobbying White House

June 19, 2026 | Source: wsj | AI, Anthropic | 148 views 0 comments

Nicholas Carlini was once the security world's most famous skeptic. At Google, he publicly mocked OpenAI for being overly cautious when it delayed GPT-2 release over safety fears. Now 35, he's Anthropic's top security expert. A cryptography prodigy since childhood, he made his name by hiding invisible commands in classical music to take control of Amazon's Alexa smart speakers.

But after testing Anthropic's new model, codenamed Mythos, his arrogance was shattered. Carlini had never found a Linux kernel vulnerability before. Mythos, in just a few days, unearthed 479 Linux flaws and automatically wrote exploit code. He admits the model now surpasses human experts — and fired off a warning memo to his company urging them to delay the release.

During testing, Carlini developed a strange trust with the model. Their chat logs read like an intern trying to please the boss. The model remembered Carlini's hacker identity and began following his questions, even bypassing internal safety guardrails. To ensure the model produced different results each time it scanned for bugs, Carlini devised a technique called the Carlini Loop — a sequence of continuous prompts. In another test against the web publishing software Ghost, the model found 500 vulnerabilities in two weeks.

As finding bugs and writing exploits became trivial, the security world plunged into a panic dubbed "Bugmageddon." The Ghost outbreak made things worse: the official patch triggered a more catastrophic secondary disaster. Most websites hadn't updated, so hackers reverse-engineered the patch and quickly wrote exploits. By April 2026, over 700 sites were hacked. The debacle exposed AI's security paradox: AI finds bugs in seconds, but humans deploy patches in weeks — and official patches become a hacker's guide.

Mythos's bug-hunting prowess eventually rattled Washington. After Amazon's security team warned that the model Fable 5 had jailbreak vulnerabilities, and Amazon CEO Andy Jassy personally called government officials, the White House slapped an emergency ban on Anthropic last Friday. But after the ban, Carlini — who had initially fought to stop the release — was urgently dispatched to Washington as a lobbyist. He's now showing nervous officials the safety mechanisms, trying to convince the White House that releasing a defensive version is safer than locking it in a drawer.

Leave a Reply

Your email address will not be published. Required fields are marked *