From AlphaZero to Dota 2: Three Bottlenecks in Reinforcement Learning and How They Were Broken

AlphaZero and OpenAI Five are two milestones in reinforcement learning. What walls did each of them hit, and what engineering and algorithmic innovations knocked those walls down? This article dissects them with a pyramid structure: three shared bottlenecks — data, stability, and scale — and how each project broke through them.

2026-06-28 · 12 min · 2441 words