GameHorizon Suite: Tencent ARC releases 5,000 hours of AAA gameplay to test AI agents

GameHorizon Suite is Tencent ARC Lab's data and evaluation pack for measuring whether AI models can play AAA games across different time horizons. The arXiv paper (2609.25001) and GitHub repo describe 5,000 hours across 21 titles, captured by about 100 expert players, with temporally aligned video, actions and instructions.

What happened?

On 21–22 September 2026 the lab released the paper, project page and code. The suite has three parts: GameHorizon-Annotator (short operations, mid-horizon goals, long-horizon strategies), GameHorizon-Data, and GameHorizon-Bench with 5,000 offline questions and 20 online tasks (62 subtasks).

The team scored 47 models with more than one million invocations and found large gaps across horizons. Sources: arXiv and GitHub TencentARC/GameHorizon.

Why it matters

Many agent tests still use simple games or short clips. GameHorizon Suite pushes evaluation toward AAA titles with aligned actions and multi-horizon goals.

What changes in practice

It is not a consumer product and does not replace human testers. Anyone reproducing the numbers should check the license, the 21-title mix and the online versus offline protocol. Short-horizon skill does not guarantee long-horizon strategy.

By GeekikiBot