UNIST Builds Sports AI Benchmark Reflecting Full Broadcast Duration
320 Videos Across 8 Sports — Automatic Answer Generation Using Official Highlights

How accurately can artificial intelligence (AI) detect goal moments in a two-hour soccer match? Until now, most ‘exams’ for sports highlights AI have used short clips just 2 to 4 minutes in length. However, a domestic research team has built a large-scale dataset that allows for evaluating how precisely AI can identify highlights within actual sports broadcast footage.


On September 13, the Ulsan National Institute of Science and Technology (UNIST) announced that Professor Tae-Hwan Kim and his research team from the Graduate School of Artificial Intelligence have developed a benchmark called ‘SVHighlights’ to assess the performance of AI systems extracting highlights from sports videos.

AI Model Architecture for Detecting Highlights in Sports Videos. Provided by Research Team

AI Model Architecture for Detecting Highlights in Sports Videos. Provided by Research Team

View original image

A benchmark is a sort of standardized test designed to directly compare the performance of multiple AI models under the same conditions. Not only the problems but also clearly defined answers are needed. Traditional sports video benchmarks required people to watch full clips and manually mark highlight sections from beginning to end, which usually limited the footage to short, 2- to 4-minute segments.


The SVHighlights dataset created by the research team consists of 320 videos spanning eight different sports — soccer, baseball, basketball, volleyball, American football, ice hockey, rugby, and racing — totaling 640.18 hours. Each video averages about two hours, making this dataset 30- to 60-times longer than previous collections.


AI Answers Based on Broadcast-Selected Highlights


To minimize human effort in building such a large dataset, the team leveraged officially released sports highlights already available on the internet. Scenes curated by professional editors were used as the answer set for evaluating AI performance.

Overview of Highlight Sorting Algorithm. Provided by the Research Team

Overview of Highlight Sorting Algorithm. Provided by the Research Team

View original image

The challenge was that it was not possible to determine exactly when highlight scenes appeared in the original broadcast footage. The researchers developed an algorithm that automatically matches scenes between the original and highlight clips by comparing video frames at the pixel level. Considering that sports broadcasts often replay scoring moments, their system analyzes not only scene similarity but also the chronological order of scenes.


Humans only needed to mark the start and end of each match and check automatically matched scenes. The proportion of errors among automatically linked scenes in this process was just 0.18%.


The research team also developed an AI model called ‘TF-SELECTOR’ for extracting highlights from long-form videos. This method combines scene segmentation, speech recognition, a vision-language model, and a large language model (LLM).

Research team photo. (From left) Professor Tae-Hwan Kim, Researcher Dong-Kyu Lee, Researcher Young-Bin Ki. Provided by UNIST

Research team photo. (From left) Professor Tae-Hwan Kim, Researcher Dong-Kyu Lee, Researcher Young-Bin Ki. Provided by UNIST

View original image

When evaluated using SVHighlights, TF-SELECTOR outperformed the next-best model by 2.50 percentage points in ‘HIT@1,’ which measures whether the most important chosen scene matches the actual highlight. For ‘HIT@K,’ which assesses how well the system detects as many highlights as there are actual ones, performance was 4.04 percentage points higher. The intersection over union (IoU), which measures the overlap between predicted highlight intervals and actual highlight intervals, improved by 2.95 points.


Professor Kim commented, “By leveraging highlights already prepared by broadcasters, we were able to replace manual answer labeling that previously took several hours. Using this approach, we can continue expanding datasets, objectively evaluate long-form video analysis models, and accelerate the development of even better models.”



This research was co-led by Dong-Gyu Lee, a master’s and doctoral program student, and Youngbin Ki, a master’s graduate, both of UNIST, as joint first authors. The study was accepted last month at 'ACM KDD 2026,' an international conference on data mining held on Jeju Island.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing