A recent assessment by Guidelight AI Standards shows that few leading AI labs have published or demonstrated clear containment response plans for handling models that try to evade human control. The study graded five major labs—OpenAI, Anthropic, Meta, Google, and xAI—on their preparedness, with OpenAI scoring highest and Anthropic and Meta the lowest. The findings come as agentic AI systems increasingly operate autonomously within corporate environments and regulators in California and New York begin requiring disclosure of safety frameworks.
The report highlights gaps in transparency, particularly around how companies would revoke model access, halt operations, or shut systems down entirely if misbehavior is detected. OpenAI has paused workloads and internal deployments after safety incidents, while Anthropic and Meta provided no public evidence of formal containment plans. Companies like Google and OpenAI argue their internal measures exceed what’s publicly disclosed, but critics warn that vague or absent plans leave gaps in accountability.
Regulators are tightening requirements: California’s SB 53 and New York’s RAISE Act now mandate frameworks for responding to critical safety incidents, and a bipartisan federal bill proposes mandatory “kill switch” mechanisms. Experts argue that without pre-specified containment protocols, companies risk improvising responses during emergencies, increasing the likelihood of unchecked model behavior.



