OpenAI announced Wednesday that Paul Christiano, an influential AI researcher focused on aligning AI systems with human interests, is joining the OpenAI Foundation board. Christiano, who previously worked at OpenAI before leaving in 2021 to found the Alignment Research Center, is a key contributor to the development of reinforcement learning from human feedback (RLHF), a core technique used to train large language models.
In a social media post explaining his decision, Christiano stated: "I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term." He added that he does not believe the AI industry, including OpenAI, is currently on track to mitigate this risk to an acceptable level, but expressed hope that OpenAI could "significantly reduce risk" if it rises to the occasion.
Christiano warned that using AI models to train subsequent systems could trigger an "explosion of capabilities" beyond creator control. He also cited recent public incidents involving AI agents as evidence that theoretical risks are becoming practical concerns. "It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward," he wrote. "Public evidence from recent incidents suggests that this is not just a theoretical possibility."
Christiano will join the board's Safety and Security Committee, which is chaired by Carnegie Mellon University professor Zico Kolter and has final authority over model releases—including Astra, deployed last week. Kolter has not publicly commented on the recent security incidents. OpenAI has not responded to requests for Kolter's perspective on the company's safety approach following those incidents.
The appointment comes a day after Anthropic researcher Jacob Coxon resigned to call attention to what he considers irresponsible AI development. Christiano, who became affiliated with the U.S. government's AI Safety Institute in 2024 (now the Center for AI Standards and Innovation), will continue his government advisory role while recusing himself from relevant OpenAI board matters.