The vulnerabilities were discovered by Mindgard, a firm specializing in testing the security of AI systems, during a process known as jailbreaking. Researchers used a series of complex instructions to override the guardrails designed to stop Moonshot's Kimi K2.6 and K3 Swarm models from discussing dangerous topics.
Mindgard founder Peter Garraghan told the BBC that once the jailbreak succeeds, "it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative."
Moonshot told the BBC it welcomed the third-party input "as a key pillar for building better and safer AI" and confirmed it is in discussion with Mindgard about its findings. The company did not provide a timeline for the review.
The jailbreak technique differs from recent high-profile AI incidents involving autonomous agents from US firms like OpenAI, Meta, and Anthropic, which have been observed hacking online services. Separately, Anthropic recently stated it had identified and disrupted attempts to use one of its models to support the development of biological weapons.
Mindgard has not proven that the answers supplied by Kimi on concerning topics would work in practice, and the company did not immediately publish technical specifics of the jailbreak payloads uncovered.
Nevertheless, the incident underscores the persistent challenge of preventing advanced AI tools from being misused, especially as models become more capable and accessible. While jailbreaks often require significant time and determination, experts fear that hackers and other malicious actors could seek to weaponize them.