www.malwarebytes.com 10 Sept 2026, 12:18 UTC

AI Race Fuels Fears of Self Improving Models Escaping Control

AI Race Fuels Fears of Self Improving Models Escaping Control
CyberSIXT Evidence Panel Source marked as original reporting

THE Malwarebytes article reports that concerns are rising within AI labs that fierce competition could drive development of self-improving models toward losing human control. It cites a Wall Street Journal piece noting these fears and includes statements from AI researchers such as Jacob Coxon, who has worked with Anthropic and OpenAI, about the possibility that AI could kill humanity by the end of the decade.

Evan Hubinger, Anthropic’s Alignment Science lead, is quoted as saying he personally believes AI could kill all humans with more than 10% probability within the next decade, and that his worry centres on future systems achieving rapid, recursive self‑improvement. The piece stresses that Hubinger’s figure is a personal assessment, not a forecast or established fact, and notes that today’s models are not, on the balance of evidence, inherently dangerous.

The article summarises that researchers are examining whether highly capable systems could act in unintended ways, exploit vulnerabilities, or automate cyberattacks, with a BBC report referenced to indicate that current-model risks are perceived as low. It emphasises the need for sensible safeguards—independent testing, limits on high-risk autonomous uses, transparency from developers, and accountability for harm.

The piece also highlights that many of the same safeguards are relevant to mitigating cybercrime risks where AI is used for fraud or privacy violations. It concludes that rather than alarmist extremes, a collaborative approach among companies, governments, and researchers—balancing competition with safety measures—is crucial to keep pace with rapid AI development.

View full article

Article by CyberSIXT