From Human Voices to Machine Sounds
The Race to Advance Voice Recognition Technology

As artificial intelligence (AI) evolves from chatbots to agents, competition is intensifying to advance voice recognition technology. The competitive landscape is expanding beyond voice AI that simply understands human speech, to technologies capable of converting machine sounds and ambient noises into data.

After Text and Image Comes Voice: AI Is Listening View original image

According to international media outlets such as TechCrunch on August 19, U.S. voice AI startup Whisper has raised USD 280 million (approximately KRW 400 billion) in funding. The company's valuation soared to USD 2 billion, nearly tripling in just nine months from USD 700 million in November of the previous year.


Whisper offers a voice AI service that allows users to input text into computers and smartphones by speaking instead of using a keyboard. The AI not only transcribes user speech but also removes unnecessary expressions and refines sentences for context. The company aims to establish a 'voice operating system,' positioning voice as a fundamental input method alongside the keyboard and mouse.


This shift is underpinned by the expectation that interactions between people and AI will transition from text to voice. There is growing anticipation that users will be able to delegate tasks with a single spoken command, rather than navigating multiple menus or inputting manual commands.


Major AI companies are also advancing real-time voice technology. Whereas earlier voice AI focused on converting speech into text, recent advancements achieve natural interactions that understand contextual conversation. Last month, OpenAI released its real-time voice model, GPT Live. Unlike previous systems that waited for users to finish speaking before responding, it enabled simultaneous listening and speaking during conversation.


In Korea, audio AI company Deeply analyzes nonverbal sounds—such as machine and impact noises generated on manufacturing floors—using AI to convert them into data. Recently, the company has focused on building AI solutions applicable to manufacturing environments and establishing standardized enterprise models. Another voice AI company, ElevenLabs, is applying technology across industries to separate and restore a specific speaker's voice, even from old or low-quality archive recordings. With only small amounts of voice data, the system learns the speaker's tone and speech patterns, then uses the restored voice to produce new content.


Fortune Business Insights projects that the global voice recognition market will grow more than fourfold: from USD 23.7 billion this year to USD 104.05 billion by 2034. The market is expected to expand by over 20% annually between 2026 and 2034.



Park Jin-ho, Professor at Dongguk University's Department of Computer Science and AI, stated, "As AI evolves into agents capable of performing multi-step tasks, interfaces employing human senses—such as sight and hearing, in addition to voice—will become more prominent." He added, "To improve the accuracy of voice recognition in real-world environments, the quality of input data must be enhanced, and technologies for contextual inference, noise removal, and filtering also need to progress."


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing