Reasoning Performance Across Text and Images Improved Up to Twofold
Expected Applications in Scientific Paper Analysis, AI Assistants, and More
Presented at CVPR 2026

Generative artificial intelligence (AI) technology capable of connecting clues in both text and images and reasoning step by step like a human has been developed by a domestic research team. By significantly enhancing the ability to synthesize multiple sources of information beyond simply understanding sentences or images individually, this technology is expected to be utilized in applications such as scientific paper analysis, AI assistants, and complex document search.


On August 6, Korea University announced that the research team led by Professor Hongseok Suh of the Department of Computer Science had developed CRIT (Cross-modal Multi-hop Reasoning over Interleaved Image-Text), a training and evaluation system designed to improve the multi-step reasoning abilities of AI that utilizes both text and images in conjunction.

Research paper image. An image showing the step-by-step process of connecting clues scattered in text and images to derive the correct answer. Explains the cross-modal multi-hop reasoning structure and evidence-based answer generation through an example asking the color of the laptop. Provided by the research team.

Research paper image. An image showing the step-by-step process of connecting clues scattered in text and images to derive the correct answer. Explains the cross-modal multi-hop reasoning structure and evidence-based answer generation through an example asking the color of the laptop. Provided by the research team.

View original image

Existing AI systems excel at understanding either text or images individually, but have shown limitations in connecting information of different modalities to perform multi-step reasoning tasks. There have also been criticisms that some evaluation methods allow answers to be found based on specific information alone, making it difficult to verify true reasoning ability.


Enhanced Complex Reasoning through Integrated Understanding of Text and Images


To address these issues, the research team developed CRIT, which automatically generates questions that require the use of both text and images to find an answer. The system is characterized by its design, which enables the analysis of both textual and visual information in diverse materials such as natural images, videos, and scientific papers, and guides the AI to reach conclusions through multiple steps.


In particular, the research team increased the consistency and reliability of the questions by structuring entities, their attributes, and the relationships between entities as graphs, then generating questions and answers based on this structure. This allows for a more precise assessment of whether an AI is truly capable of integrating various sources of information and reasoning step by step.


As a result of training AI with CRIT, the research team found that performance in reasoning tasks—ranging from everyday image and video analysis to the interpretation of scientific papers and tables—improved up to twofold compared to previous models. The researchers explained that these results are significant, given that even the latest AI models generally demonstrate low performance on complex reasoning tasks.

Research team photo. (From left) Hongseok Seo, Professor of Computer Science at Korea University (corresponding author), Junyoung Sung (senior, first author), Minjun Kim (senior, co-author), Seungwoo Yoo (senior, co-author), Sumin Ahn, integrated master's and doctoral program (co-author). Provided by Korea University

Research team photo. (From left) Hongseok Seo, Professor of Computer Science at Korea University (corresponding author), Junyoung Sung (senior, first author), Minjun Kim (senior, co-author), Seungwoo Yoo (senior, co-author), Sumin Ahn, integrated master's and doctoral program (co-author). Provided by Korea University

View original image

Furthermore, models trained with CRIT showed not only specialization for a specific dataset, but also an overall enhancement in their ability to understand and reason with complex information across various environments.


The research team anticipates that this technology will be applicable in a range of fields, including scientific literature analysis, educational content comprehension, complex document search, and vision-and-language-based AI assistants.



This research was presented at CVPR 2026, the world’s leading international academic conference in the field of AI, and was made more meaningful by the participation of undergraduate students as first author and co-authors.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing