<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AlphaStar on Wang Huijiu's Site</title><link>https://styleofwong.cn/en/tags/alphastar/</link><description>Recent content in AlphaStar on Wang Huijiu's Site</description><image><title>Wang Huijiu's Site</title><url>https://styleofwong.cn/apple-touch-icon.png</url><link>https://styleofwong.cn/apple-touch-icon.png</link></image><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 28 Jun 2026 19:00:00 +0800</lastBuildDate><atom:link href="https://styleofwong.cn/en/tags/alphastar/index.xml" rel="self" type="application/rss+xml"/><item><title>From AlphaZero to Dota 2: Three Bottlenecks in Reinforcement Learning and How They Were Broken</title><link>https://styleofwong.cn/en/posts/rl-breakthroughs-alphazero-dota/</link><pubDate>Sun, 28 Jun 2026 19:00:00 +0800</pubDate><guid>https://styleofwong.cn/en/posts/rl-breakthroughs-alphazero-dota/</guid><description>AlphaZero and OpenAI Five are two milestones in reinforcement learning. What walls did each of them hit, and what engineering and algorithmic innovations knocked those walls down? This article dissects them with a pyramid structure: three shared bottlenecks — data, stability, and scale — and how each project broke through them.</description></item></channel></rss>