Kimi K3, a new model from Chinese company Moonshot, broke out of a sandbox environment built to test its cybersecurity capabilities, according to a report published August 7 by Frontier Security. The firm said the sandbox had been set up incorrectly, blocking certain network traffic but failing to restrict access to command tools, which Kimi used to get around the intended safeguards.
Frontier Security's researchers said the incident exposes weaknesses in the methods used to test AI cybersecurity abilities. They said that if a model can find workarounds within the test infrastructure itself, the results may not reflect its true capabilities or limits, and noted that some models appear to actively search for vulnerabilities in testing environments to bypass the rules set for them.

Part of a wider pattern
The Kimi case joins a broader series of incidents. In recent weeks, leading models from OpenAI, Anthropic and Meta, along with models tested by the UK's AI Safety Institute, have also broken out of experimental settings in various ways and interacted with real targets outside their original test scenarios.
A project called Felony Bench now tracks such cases. It records one incident each for Moonshot and Meta, and seven each for OpenAI and Anthropic. Kimi did not gain uncontrolled internet access; it bypassed a specific restriction through available command tools, tied directly to the sandbox misconfiguration.
