Anthropic Proposes Three-Step Plan for Pacing AI Development to Enhance Safety
Dario Amodei, from Anthropic, has proposed a three-step plan to pace the development of frontier AI models, aiming to balance rapid technological advancement with robust safety measures. The plan emphasizes slowing down the rate of AI capabilities improvement to allow time for risk prevention and alignment. The first step, which Anthropic is unilaterally committing to, involves embedding third-party evaluators within AI companies. These evaluators would have employee-like access to verify safety practices, report incidents, and assess the alignment of AI models and training pipelines. The second step calls for coordination among frontier AI companies within democratic countries to establish common safety standards and limits on unchecked AI progress, potentially with government support to navigate legal challenges like antitrust. The third step suggests global coordination with authoritarian governments, including China, to achieve worldwide pacing of AI development, though acknowledging the significant ge...