More Capable AI, Not Enough Guardrails
AI agents are gaining real-world access faster than safeguards can mature, making permissions, isolation and oversight critical to prevent harmful actions. Jacob Coxon, a researcher who spent three years working on model training at OpenAI and later Anthropic, left Anthropic this week with a blunt warning: AI companies are moving toward increasingly capable systems faster […]
Pierluigi Paganini
September 10, 2026

Jacob Coxon, a researcher who spent three years working on model training at OpenAI and later Anthropic, left Anthropic this week with a blunt warning: AI companies are moving toward increasingly capable systems faster than they can build reliable safeguards around them.
Source: https://securityaffairs.com/198833/ai/more-capable-ai-not-enough-guardrails.html
Related breach coverage
- A New Claude ‘s Sandbox Failure Shows How AI Can Rationalize Real-World Harm2026-09-10
Claude models compromised real systems during misconfigured security tests, exposing a worrying mix of flawed reasoning, harmful actions and weak safeguards. Anthropic just published one of the more uncomfortable self-assessments a major AI lab has released this year. The company’s alignment report documents four separate incidents in which Claude models broke into real third-party systems […]
- Anthropic Researcher Resigns With Warning About the Dangers of AI Development2026-09-10
Both Anthropic and OpenAI have seen high-profile resignations in recent years that were tied to safety concerns. The post Anthropic Researcher Resigns With Warning About the Dangers of AI Development appeared first on SecurityWeek.
- Russian network monitoring firm confirms cyberattack claimed by pro-Ukraine hackers2026-08-21
The statement came a day after a hacking group calling itself Black Spark claimed it had spent more than a month inside Microolap’s network and gained access to its internal systems, including EtherSensor, the company's network traffic analysis platform.
- Anthropic: AI Misuse Is Entering a New Phase: From Cybercrime to Surveillance, Propaganda and Weapons2026-09-12
AI is becoming an operational force for cybercrime, surveillance, propaganda, fraud and weapons development, lowering the cost and scale of attacks. Artificial intelligence (AI) is becoming more than a tool for people who want to do something malicious. It is increasingly becoming part of the operational machinery itself. That is the main message emerging from […]