A group of analysts, who previously predicted the extinction of humanity due to uncontrolled artificial intelligence, has presented an unexpected rescue scenario. The document, titled "AI 2040: Plan A", was published by the AI Futures Project, founded by former OpenAI researcher Daniel Kokotajlo.
Kokotajlo left OpenAI in April 2024 due to fundamental disagreements over AI safety issues. By 2025, he launched his own analytical center. His previous report, "AI 2027," painted a grim picture: an arms race between the US and China in the field of superintelligence, with an extinction probability of 10% to 30%. Now, experts are proposing a specific set of steps that they believe could drastically reduce these risks.
How It Should Work
The key idea of "Plan A" is an international agreement between Washington and Beijing, which must be concluded by 2029. Without it, it is claimed, full automation of AI development will occur as early as 2030. Instead, countries agree to develop neural networks gradually, only up to the level of the best human experts. By 2035, a mandatory pause in the race is introduced to maintain human control over the systems. Only by 2040 are the restrictions lifted, and AI reaches the level of superintelligence.
The plan is based on four principles:
- Buying time for safety research;
- Full transparency of all AI developments;
- Distribution of computing power among different companies and countries;
- Maintaining reversibility of the process.
To ensure trust, verification from space is proposed: large data centers are visible from satellites. The first step is public declarations on chip purchases. Then, a temporary pause on training is introduced, monitored by sensors. After trust is confirmed, restrictions are lifted, but full transparency remains.
The most radical part of the plan is a deterrence mechanism reminiscent of the logic of nuclear parity. It proposes building new Chinese data centers in Canada, and US facilities in Mongolia. In the event of a deal breakdown, the host country would attempt to seize the facilities, while the owner would destroy them to prevent them from falling into enemy hands.
Economics and Social Consequences
According to calculations, global computing power will grow from 20 million H100-equivalents in 2026 to 60 billion by 2034. Real US GDP growth in certain periods of the 2030s could reach 50% per year. However, automation will lead to a catastrophic drop in employment: from 62% in 2027 to 12% by 2040. This is proposed to be offset by "citizen dividends"—payments from state revenues from licensing computing and robots. The forecasts are impressive: $45,000 per person in 2032, $1 million by 2035, and $10 million by 2039.
Alternative Scenarios
Experts also modeled four other paths of development:
- Plan B: The US forms a coalition and pressures China, including through cyberattacks. The outcome is loss of control or war.
- Plan C: An attempt to negotiate, but under pressure from companies, the pause is quickly lifted. The risk is a permanent oligarchy of superintelligence owners.
- Plan D: Minimal regulation and a race. Risks include loss of control, concentration of power, and World War III.
- Plan S: A complete indefinite halt. The main danger is the inevitable breakdown of the deal and the resumption of the race in chaos.
My analysis: "Plan A" is undoubtedly an ambitious attempt at a systematic approach to a problem that many prefer to ignore. However, its implementation requires an unprecedented level of global trust and coordination, which, in the current geopolitical climate, seems extremely unlikely. The "mutually assured destruction of computing power" mechanism is an interesting but frightening analogy to nuclear weapons, which rather highlights the fragility of the balance than offers a reliable solution. The market will likely ignore these forecasts, but for serious investors and regulators, this is an important signal about the depth of potential risks.