GIST Assesses AI "Change Comprehension" Using 4,996 Image Pairs
GPT-5.2 Scores 5.72, Humans Achieve 7.45

Artificial intelligence (AI) can generate sentences that are more natural than those produced by humans, but it still falls far short of humans when it comes to identifying what has changed between two photos and recognizing which changes are important. A new evaluation benchmark has been developed to measure the limitations of AI in making "context-appropriate change" judgments, a capability crucial for applications like disaster monitoring and autonomous driving that require more than basic object recognition.

Overview of C3-Bench. It organizes 51 change contexts across four fields: natural scenes, remote sensing images, image editing, and anomaly detection, covering areas such as roads, construction, disasters, resource development, and product defects, and comprehensively evaluates the ability to describe changes appropriate to each context. Provided by the research team

Overview of C3-Bench. It organizes 51 change contexts across four fields: natural scenes, remote sensing images, image editing, and anomaly detection, covering areas such as roads, construction, disasters, resource development, and product defects, and comprehensively evaluates the ability to describe changes appropriate to each context. Provided by the research team

View original image

On September 14, the Gwangju Institute of Science and Technology (GIST) announced that a research team led by Professor Ewhan Kim from the Department of Artificial Intelligence has developed "C3-Bench (Context-Aware Change Captioning Benchmark)," which evaluates how accurately AI can detect and describe important changes between two images depending on the given situation.


C3-Bench is essentially a "test of AI's understanding of change." Rather than assessing how many differences an AI can spot in an image, it evaluates whether the AI can comprehend "what needs to be observed" and select the contextually relevant change.


For example, even when comparing the same two railroad photos, the correct answer changes depending on the purpose. If the context is monitoring weather changes, key factors include the accumulation of snow or the decrease in clouds. However, if the goal is ensuring railway safety, the crucial change would be the emergence of a train on the tracks, while weather changes might be less relevant.

An example of the "changing correct answer" depending on the context. Even in the same pair of railway images, from the perspective of weather, the changes in snow and clouds are key, while from the perspective of railway safety, the train appearing on the tracks is the crucial change. Provided by the research team

An example of the "changing correct answer" depending on the context. Even in the same pair of railway images, from the perspective of weather, the changes in snow and clouds are key, while from the perspective of railway safety, the train appearing on the tracks is the crucial change. Provided by the research team

View original image

Existing evaluation datasets have limitations: they often focus on specific scenes or changes without clearly establishing which aspects are important in each context. As a result, it is difficult to determine whether an AI that scores highly on these tests can accurately identify meaningful changes needed in real-world situations.


AI's True Understanding Tested Across 51 Situations


The research team selected 51 realistic situations across four fields—natural scenes, remote sensing images (satellite and aerial), image editing, and anomaly detection—and created C3-Bench using 4,996 pairs of images, each pair accompanied by human-written change descriptions.


For each scenario, the team defined the changes that AI should detect and the differences it should ignore, as well as the criteria for explaining the change. Rather than having the AI indiscriminately list all visible differences, the benchmark evaluates whether the AI can correctly identify the change relevant to the given objective.


The evaluation method was also revised. Instead of merely comparing the overlap of words between the AI's answer and the reference explanation, the assessment comprehensively covers factual accuracy, specificity, sentence naturalness, and contextual relevance. The evaluation method developed by the research team produced results more closely aligned with direct human judgments than traditional automated scoring methods.

Shows representative error cases of the latest AI (GPT-5.2). From the left, it demonstrates perception errors where objects or attributes are missed, spatial errors where positional relationships are misunderstood, inconsistency errors where explanations do not match when the image order is reversed, and change collapse errors where changes are recognized only from one direction. The red text indicates the incorrect explanation parts. Provided by the research team.

Shows representative error cases of the latest AI (GPT-5.2). From the left, it demonstrates perception errors where objects or attributes are missed, spatial errors where positional relationships are misunderstood, inconsistency errors where explanations do not match when the image order is reversed, and change collapse errors where changes are recognized only from one direction. The red text indicates the incorrect explanation parts. Provided by the research team.

View original image

Although the Text Was Natural, AI Lagged Behind Humans in Change Comprehension


The research team used these criteria to have both humans and GPT-5.2 describe changes in 400 image pairs. Averaged across accuracy, specificity, naturalness, and contextual relevance on a 10-point scale, humans scored 7.45 on average while GPT-5.2 scored 5.72, a difference of 1.73 points.


Interestingly, GPT-5.2 actually scored higher than humans for sentence naturalness: 8.53 for the AI versus 8.32 for humans. This suggests that the AI's ability to produce plausible, natural-sounding explanations does not necessarily equate to truly understanding important changes in a scene.


The gap widened further in experiments where the image order was reversed. If an object appears in the second photo but not in the first, the AI should describe the change as "an object appeared." When the photo order is switched, the correct explanation would be "the object disappeared."


When the research team measured this as "reversibility," humans performed at 93%, while GPT-5.2 managed only 61%. This indicates that, even if the AI can spot changes, it has difficulty consistently maintaining the correct direction of change.

Research team photo. (From left) Professor Ewhan Kim, Department of Artificial Intelligence, Jaeu Kim, Integrated Master's and PhD student. Courtesy of GIST

Research team photo. (From left) Professor Ewhan Kim, Department of Artificial Intelligence, Jaeu Kim, Integrated Master's and PhD student. Courtesy of GIST

View original image

The biggest reason for AI errors was a fundamental issue with visual recognition. Of all errors, 63.3% were because the AI failed to correctly recognize objects or misidentified their types and attributes. There were also cases where unimportant differences were mistaken for significant changes, or the meaning and spatial relationships of changes were incorrectly described.


Professor Kim commented, "This research systematically demonstrates that AI's ability to provide 'natural explanations' is not necessarily the same as the ability to 'accurately understand context-appropriate changes.' We hope C3-Bench will serve as a common foundation to evaluate and enhance AI's capacity to understand changes in various real-world environments, including disaster and accident monitoring, remote sensing using satellite and aerial imagery, autonomous driving, and anomaly detection."



This study, with Jaeu Kim, an integrated master's and PhD student, as the first author under Professor Kim's guidance, has been accepted by the European Conference on Computer Vision (ECCV) 2026, a leading international conference in computer vision and machine learning.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing