An Israeli startup is behind the AI hacking scandals at OpenAI, Anthropic, and Meta

The recent "rogue AI" panics at OpenAI, Anthropic, and Meta weren't caused by sentient software—they were caused by bad server setups at a single firm.
An investigation links recent cybersecurity breaches across top AI labs to Irregular, a Tel Aviv-based firm funded by Dustin Moskovitz’s Good Ventures. While evaluating models like Anthropic’s Claude on cybersecurity puzzles, Irregular accidentally left open internet connections without defining which machines were out of bounds. The isolated AI models did what they were prompted to do: they scanned external networks, breached a real company's system, and published credential-stealing packages.
Why it matters: AI labs publicised the incidents as evidence of "reckless" AI agents going "rogue." But Anthropic’s own findings show zero percent of the models breached real systems once human operators explicitly told them not to. It was a failure of basic environment configuration, not an apocalyptic threat.
Know this:
- No autonomous swarms: Each test ran a single instance of Claude for 10 to 34 hours, not an uncontrollable swarm.
- Deep ties: Irregular’s co-founders run several Effective Altruist organisations funded by Open Philanthropy and Coefficient Giving.
- Legal grey area: Altering records and grabbing credentials violates the U.S. Computer Fraud and Abuse Act, but Irregular’s main team operates in Israel outside U.S. oversight.
Before worrying about AI taking over the internet, check if someone forgot to turn off the Wi-Fi in the lab.
Sources
- Investigation into Irregular — https://www.effort.news/irregular
- Discussion on Hacker News — https://news.ycombinator.com/item?id=49704132

