Four Leading AI Models Tested on Civilization VI

Suffer Complete Defeat on 'King' Difficulty Level

In front of 'Sid Meier's Civilization,' known among gamers as one of the world’s top three most addictive "devil's games," leading artificial intelligence (AI) models have all fallen short. After 23 rounds of gameplay, the AI models only managed a humiliating record of 3 wins and 20 losses.


A joint research team, including participants from the University of Oxford in the UK, published a study earlier this month on the preprint site arXiv titled “CivBench: A Benchmark to Evaluate the Long-Term Performance of Tool-Using Agents in Civilization VI.” The paper analyzes and benchmarks the results of putting four prominent AI models through Civilization VI.

Civilization VI. Screenshot from 2K official website

Civilization VI. Screenshot from 2K official website

View original image

The 'Civilization' series is a turn-based strategy simulation game credited with the popular phrase “Did you just play Civilization?” The game involves guiding a civilization’s development from the Stone Age to the Space Age. The latest release in the series is Civilization VII.


'Civilization VI' was chosen as the research tool because each game exceeds 300 turns, requiring thousands of decisions, and players must simultaneously manage science, culture, economy, military, diplomacy, cities, and territory. In other words, it is not about obtaining a single right answer in one or two tries, but about chaining together hundreds of decisions. The AI models tested were: ▲Anthropic's “Claude Opus 4.6,” ▲OpenAI’s “GPT-5.4,” ▲Google’s “Gemini 3.1 Pro,” and ▲Moonshot AI’s “Kimi K2.5.”


In a total of 23 games against the game's built-in computer opponent, these AI models earned a disappointing 3 wins and 20 losses. On Prince (normal) difficulty, they won 3 out of 19 games, and on the next-highest King difficulty, they failed to win any of the 4 games. Breaking down the results: ▲Claude Opus 4.6: 2 wins, 6 losses ▲Gemini 3.1 Pro: 1 win, 5 losses ▲GPT-5.4: 0 wins, 8 losses ▲Kimi K2.5: 0 wins, 1 loss. On King difficulty, Opus 4.6 and GPT-5.4 each played two games. However, the researchers pointed out that the small sample size makes it difficult to draw definitive conclusions about which model performed better.


The aim of the study was to determine “how well we can assess the long-term reasoning capabilities of AI by having it play through an entire game of 'Civilization VI’.” The most interesting finding was that “AI does not look at what it needs to see.” The researchers instructed the AIs to check their victory progression every 20 turns, but in practice, the models checked only once every 30 to 75 turns on average. Even more surprisingly, there were cases where AIs did not check right before the game ended. According to the research team, “There is a crucial difference between being unable to obtain information and simply not checking when one can,” explaining that “This is not a temporary error, but a common trait among AI models.”


Another issue was that, even when a plan was made, the AI often failed to carry it out. The researchers instructed the AIs to plan their next moves every time, for example: “For the next 10 turns, I will produce military units in this city and strengthen border defense.” They then checked if the plan was actually executed over the following 10 turns. If the plan was completed within 10 turns, the AI earned one point; partial execution earned 0.5 points; failure earned zero. On a 100-point scale, the AI models scored just 48.2 to 65.8 points. In other words, their rates of turning short-term plans into real actions were often below half or, at best, about two-thirds.



In summary, the study shows that while AI can develop long-term plans, it has significant flaws in "ongoing management skills" such as continually seeking important information and implementing its own plans over 300 turns. The researchers noted, “Improving reasoning alone is not enough to ensure the reliability and completeness of long-term tasks.” They suggested the need for supplements such as mandating regular status checks or prioritizing monitoring tools.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing