HomeTechAnthropic details practical metrics to help monitor the speed of AI development

Anthropic details practical metrics to help monitor the speed of AI development

Days after Anthropic PBC Chief Executive Dario Amodei rocked the artificial intelligence world by calling on leading model makers to coordinate on the slowdown on the pace of development, the company has shared three new metrics that should be measured to enable this.

In a blog post today, Anthropic said it is already measuring AI-led research and development, oversight of autonomous AI agents and compute allocation, and then detailed the exact methodologies it’s using so that other frontier labs can measure them too. In doing so, it helps flesh out a three-step plan for a coordinated, industry slowdown that Amodei published at the weekend.

“As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows,” the company said. “This means better measuring the development of AI, reporting on it publicly and giving society an opportunity to decide how to use this information.”

Anthropic said it’s measuring AI-led R&D because of the tendency of the best frontier labs to use AI itself to develop even more powerful models that build on their capabilities, which could lead to humans failing to understand just how dangerous they are. It’s measuring oversight of AI agents because people are increasingly using them to perform work on their behalf, and in turn, those agents then delegate dozens of separate tasks to sub-agents. The risk is that we somehow shift from the stage where AI “collaborates” with humans to one where “AI leads,” by making more consequential decisions, such as which direction to take AI research.

As for compute allocation, Anthropic is monitoring this metric because it’s the primary fuel that AI runs on. By understanding how this resource is allocated, it can see where developers are focusing their efforts and how this changes over time. “Additionally, compute is among the most verifiable inputs to the AI R&D process, meaning that it could be a critical lever in a future pacing effort,” Anthropic wrote.

“As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows,” the company said. “This means better measuring the development of AI, reporting on it publicly and giving society an opportunity to decide how to use this information.”

The three metrics could help to accelerate Amodei’s plan, which has already won the backing of many of his peers in the AI industry, including OpenAI Group PBC CEO Sam Altman, SpaceX Corp. CEO Elon Musk and Google DeepMind Chair Demis Hassabis, among others.

Amodei’s call for a slowdown came in the wake of a number of stark warnings from other AI researchers about the technology’s growing potential to be harmful to humans. Earlier this month, one of Anthropic’s top AI safety researchers, Evan Hubinger, warned that the technology is advancing so rapidly that he estimates there is a greater than 10% chance it “could kill all humans” within the next 10 years. Though Hubinger stressed that the risk from current AI models is still very low, he said he worried about its potential to develop and improve itself to the point where humans can no longer control it, at which point it would pose an existential risk to humanity.

In his plan outlined on Saturday, Amodei said he wants to curtail the rapid pace of frontier model development “without sacrificing commercial advantage of the United States’ lead in AI.”

For the first metric, Anthropic said it has developed a special index that measures how involved Claude is in its R&D processes, and found that it’s “not operating fully autonomously” in any subset of that work. To measure the second metric, it built a system that’s able to oversee AI agents and intervene on the actions they take within its internal systems. That allowed it to determine that it’s currently running about 30,000 autonomous agents in its computing environments, with most of them involved in its R&D work.

For the third metric, Anthropic said it has started taking regular snapshots of how its compute resources are allocated. The first snapshot, from between July 13 to July 20, revealed that around 6% of its total compute capacity was allocated toward AI safety. An additional 12% was dedicated towards “AI-driven R&D” that focused on safety.

The company said the combined metrics can provide more visibility into how new frontier models are developed and will aid in evaluations regarding those new model’s capabilities. They help to provide more of a starting point for third-party evaluators trying to assess how fast AI is really accelerating and the things it’s really capable of. “We hope to model that transparency by releasing these measurements, and we’ll continue to do so,” the company said.

Image: Anthropic

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

About SiliconANGLE Media

SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

 

Must Read

spot_img