Jump to main content

| Research, English

Information Technology

Artificial intelligence football championship

Researchers have developed a comparison tool that predicts the results of the World Cup football matches. The capabilities of various artificial intelligence systems, including large language models (ChatGPT, Claude, etc.), will be put to the test.

LLM SoccerArena is a live ranking list for artificial intelligence. During the 2026 FIFA World Cup, several of the leading AI language models will predict the outcome of every match. The researchers will be tracking how well each model performs against the real results. You can think of it as a betting game in which GPT-5.5, Claude, Gemini, Grok and other top AI models compete against each other instead of people. The LLMs test was designed by Markus Weinmann, Professor of Business Analytics at the University of Cologne and the Institute for Business AI, along with his colleagues Oliver Müller, Professor of Data Analytics at Paderborn University, and Stefan Feuerriegel, Professor of AI for Management at the LMU Munich School of Management and the Munich Center for Machine Learning (MCML).

AI chatbots sound confident about almost everything. Football is a tough, public test. The matches are fixed, the results are indisputable and nobody knows the outcome ahead of time. So the researchers have one simple question: can these models predict the results of football matches? And will they be better if they are allowed to search for information live on the Internet in the run-up to the games?

To answer these questions, the researchers ask each model to make a prediction before every match, including the most likely result and the probability of a home win, draw or away win. Each prediction is saved with a timestamp before kick-off so that nothing can be changed afterwards. After the match, the researchers compare each prediction with the official result. The dashboard turns this into a live ranking list that can be searched and filtered.

When a language model predicts which team will win the World Cup, it has to categorize information on current form, injuries, coaching decisions, past matches, squad quality or betting odds, and use this to derive a robust prediction despite the uncertainty. Many established benchmarks for large language models assess abstract tasks in highly simplified or static environments. Football, on the other hand, is real life.

In the Artificial intelligence world cup, too, points are awarded for correct predictions. For each match, a model receives 5 points for predicting the exact score, 2 for the correct goal difference, 1 for the correct outcome (home win, draw or away win) and 0 for an incorrect prediction. The tournament-wide questions will be scored separately, with 5 points for each correct tip.

In addition, the members of the research group evaluate the predictions in the same way as professional forecasters do, by checking how well the stated probabilities match the actual events. The researchers explain more about this on the methodology page of the website.

“We don’t work on any predictions once the result is known,” explains Markus Weinmann. “We also record when a model has used the web search and when it has not, so that an open book tip is never confused with a closed book tip.”

A high ranking means that a model has performed well in the games played so far. It doesn’t prove that the model understands football, and it doesn’t say who will win the next game. “Only a few games have been scored early in the tournament, so the table will move a lot,” says the University of Cologne’s Professor of Business Analytics.

The findings are also relevant for management research. Managers are using large language models more and more frequently to structure market information, evaluate scenarios or prepare forecasts, for example, on demand trends, competitors, product launches or risks. Abstract reasoning alone is not enough: Models must recognize relevant information, classify uncertainties and derive reliable estimates from them.

The predictions made by the systems form part of a research project. They are not intended as betting advice.
 

Media Contact:
Professor Dr Markus Weinmann
The University of Cologne
+49 221 470 89981
weinmann(at)wiso.uni-koeln(dot)de

Press and Communications Team:
Robert Hahn
+49 221 470 2396
r.hahn(at)verw.uni-koeln(dot)de

More Information:
https://llmsoccerarena.up.railway.app/