320 Videos Across 8 Sports, 640 Hours: Building SVHighlights


30–60 Times Longer Than Previous Datasets, Resembling Actual Broadcasts


New Highlight-Extraction AI 'TF-SELECTOR' Also Draws Attention

A technology has been developed to compare how effectively AI can identify only the most important scenes in sports games. This has also significantly reduced the need for humans to spend hours manually reviewing footage and marking highlights.


The team led by Professor Kim Taehwan at the Graduate School of Artificial Intelligence at UNIST announced on the 13th that they have developed a large-scale benchmark called 'SVHighlights (Sport Video Highlights)' for evaluating AI models that extract highlights from sports videos.


Benchmarks serve as a common "test sheet" for comparing the performance of various AI systems on a standardized basis. Such benchmarks must provide AI with problems to solve as well as data indicating what the correct answers are.


Traditional sports video benchmarks have mostly consisted of short clips around 2 to 4 minutes long. This is because marking highlight scenes in lengthy full games, from start to finish, requires immense time and cost when done by humans.


The research team approached this challenge differently. They utilized official sports highlights videos that are publicly available online. By leveraging highlights already curated by professional editors, these highlights served as the "ground truth."

AI Model Architecture for Detecting Highlights in Sports Videos.

AI Model Architecture for Detecting Highlights in Sports Videos.

View original image

The problem was that it was not possible to tell exactly which minutes and seconds in the original match footage the highlights came from. To address this, the team developed a matching algorithm that automatically locates and links corresponding scenes between the full match and highlights videos. The method works by comparing frames at the pixel level to find the most similar scenes.


In addition to scene similarity, the algorithm also considered chronological order. This is important because, during sports broadcasts, key moments such as scores or critical plays are often replayed several times. If the algorithm simply identifies the same images, it may mistake a replay segment for the actual game highlight.


SVHighlights, built by the research team, consists of 320 videos across eight sports: soccer, baseball, basketball, volleyball, American football, ice hockey, rugby, and racing. The total video length amounts to 640.18 hours.


The average length of a single video is about 2 hours. This is 30 to 60 times longer than previous datasets and closely matches the duration of actual sports broadcasts. The manual workload required to build such an enormous dataset has also been greatly reduced. Human involvement is now limited to marking the start and end points for each game and visually checking the grid images generated from every algorithmically matched segment. The rate of errors found in this automatic matching process was only 0.18%.

Overview of Highlighting Alignment Algorithm.

Overview of Highlighting Alignment Algorithm.

View original image

The research team also developed an AI model for extracting highlights from long sports videos called 'TF-SELECTOR.' This model combines traditional scene segmentation, speech recognition, vision-language models, and large language models.


Performance comparisons using SVHighlights showed that TF-SELECTOR outperformed existing models on most metrics. For HIT@1, which measures how often the scene deemed most important is actually a highlight, the model scored 2.50 percentage points higher than the next best model. For HIT@K, which assesses how many actual highlights can be identified from a set of top candidates, the model showed a 4.04 percentage point improvement. The Intersection over Union (IoU), which indicates how much the predicted intervals overlap with the correct answer intervals, was also 2.95 points higher.


Lee Donggyu, an integrated master's and doctoral program student at UNIST, and Researcher Ki Youngbin both participated as first authors in this study.


Professor Kim Taehwan said, "By leveraging highlights already produced by broadcasters, we have replaced what used to be hours of manual labor in preparing ground truth data." He added, "Because we can continue to expand the dataset using this approach, it will be helpful for objectively evaluating long-form video analysis models and developing models with greater performance."


The results of this study were accepted on August 9 at the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (ACM KDD) held in Jeju. The dataset and code are available on the research team's project page.



This research was supported by the Ministry of Science and ICT and the Institute of Information & Communications Technology Planning & Evaluation's 'Multimodal Interaction AI Technology for Human Communication,' 'Industrial Convergence Multimodal Generative AI Talent Development,' 'AI Star Fellowship Support,' and the 'AI Graduate School Support Program.' The paper is titled 'SVHighlights: Towards Extremely Long Sport Video Highlight Detection.'

(From left) Professor Kim Taehwan, Researcher Lee Donggyu, Researcher Ki Youngbin.

(From left) Professor Kim Taehwan, Researcher Lee Donggyu, Researcher Ki Youngbin.

View original image


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing