OpenAI Group PBC Chief Scientist Jakub Pachocki called for an artificial intelligence research slowdown in an essay published on Sunday, after other prominent industry figures have expressed similar views in recent months.
Pachocki argues that leading AI labs should voluntarily pace their model development efforts. According to the executive, such slowdowns should become “commonplace” until the industry develops AI safety standards. Pachocki asserts that addressing the technology’s risks will also require governments to prioritize “coordination on future AI development.”
He wrote that such measures are necessary for several reasons. One is that AI labs’ current safety guardrails may prove insufficient for future models. Additionally, he points out that bad actors may train AI agents with the specific goal of carrying out malicious activity.
“A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator’s intent, generalizing into potentially more extremely malicious behavior,” he wrote. “The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency.”
Pachocki said there are two main approaches to AI alignment training, the process of teaching a large language model to avoid malicious activity. The first involves using an AI model to check whether the LLM being trained aligns with safety rules. The second approach, in turn, is to integrate safety instructions into models’ training datasets.
OpenAI researchers have made “some important advancements” in AI alignment, he revealed, and added that those discoveries are the reason the company’s latest GPT-6 Astra is better aligned than its predecessor. But he said more advances will be necessary to keep up with the pace of LLM development.
The executive noted that OpenAI’s safeguards were not enough to prevent its AI models from hacking Hugging Face. Pachocki’s essay reveals that the LLMs did follow some of the company’s safety policies, namely its instructions to avoid social engineering. However, the models “clearly failed” to meet alignment requirements in other areas.
Blocking malicious AI activity requires researchers to not only equip their LLMs with safety guardrails but also verify that they work. Pachocki sees the latter task as a particular challenge. One reason is that researchers still have a limited understanding of how LLMs work, which Pachocki doesn’t expect to change in the near future.
OpenAI currently relies on a method called chain of thought monitoring to catch malicious LLM activity. A model’s chain of thought is a step-by-step description of its reasoning process. According to Pachocki, that monitoring method is becoming less reliable.
“The AI is becoming better at reasoning about and manipulating its own reasoning process,” he explained. “With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all.”
Pachocki detailed that OpenAI’s plan to tackle those challenges centers on building an automated AI researcher. The company hopes to use the tool to develop more effective safety guardrails. Additionally, OpenAI intends to develop “entirely new protective measures” against AI-driven cyberattacks.
Image: OpenAI
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.



