Kanana-o Omni AI Model Enables Control of Tone, Speed, and Intonation via Natural Language Commands
Faster and More Efficient... Rivals Global Models in Benchmark Testing

Kakao announced on August 4 that it has advanced the voice generation technology of its self-developed omni artificial intelligence (AI) model, 'Kanana-o.'


According to the information released by Kakao through its tech blog on the same day, Kanana-o has enhanced its voice generation so that users can control the tone, emotion, and intonation of the AI using only natural language instructions. In addition, the company has increased both the speed and efficiency of voice generation through its self-developed voice tokenizer.


Kakao announced on the 4th that it has advanced the voice generation technology of its self-developed omni artificial intelligence (AI) model 'Kanana-o'. Kakao

Kakao announced on the 4th that it has advanced the voice generation technology of its self-developed omni artificial intelligence (AI) model 'Kanana-o'. Kakao

View original image

The improved Kanana-o allows users to directly control the actual vocal expression based on their instructions. If a user requests in natural language, for example, "Read very quickly," "Read in a low voice," "Read in a sad voice," or "Read in a Gyeongsang Province dialect," the AI will generate speech that reflects the requested speed, volume, pitch, emotion, intonation, and strength.


Furthermore, it can handle not only role-based instructions such as "Like a sports commentator," "Like an announcer," or "Like reading a storybook," but also complex instructions that incorporate multiple elements at once, such as "Lower the tone, read quickly in a sad voice." Kakao explained that the same types of directions can also be applied to English voice generation.


Kanana-o recorded a score of 94.50 on 'InstructTTSEval,' a Korean-language benchmark for evaluating the ability to follow voice instructions. This score surpasses GPT-4o-mini-tts (91.10 points) and is similar to Google's Gemini-2.5-Flash-Preview-tts (95.38 points).


Kakao improved generation speed and efficiency by applying its self-developed voice tokenizer 'LM-SPT (LM-aligned SPeech Tokenizer).' LM-SPT enables AI to represent speech with fewer tokens by compressing the information, reducing the data volume the AI must process and allowing for faster and more efficient voice generation.


Kakao plans to further develop the Kanana-o voice technology in the future. The company is focusing on advancing its voice processing capabilities and providing a natural and seamless voice experience through research that integrates both speech understanding and generation in a unified architecture.



Noh Byungseok, Performance Lead of the Kakao Unified Foundation Model Team, stated, "The latest advancement in Kanana-o's voice technology aims not only for human-like natural speech generation, but also emphasizes the ability to express users' desired tone, emotion, and intonation according to natural language instructions. Moving forward, we will apply the Kanana-o model to various services to deliver an even more natural and convenient AI voice experience."


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing