On July 9, OpenAI announced the release of a more natural ChatGPT Voice experience, based on its next-generation voice model, GPT Live.

GPT Live-Based ChatGPT Voice Screen. OpenAI.

GPT Live-Based ChatGPT Voice Screen. OpenAI.

View original image

Previously, voice AI systems would convert a user's speech into text, have a large language model (LLM) generate a response, and then read the answer out loud. This process required a separate step to determine whether the user had finished speaking, resulting in interrupted conversations or delayed responses.


GPT Live goes beyond this method by continuously processing voice input and understanding conversational context in real time. It can respond naturally even while the user is still speaking, or handle situations where the user pauses and then asks another question.


GPT Live was designed by separating the voice conversation function, which reacts swiftly to user speech, from the inference function, which processes complex requests.


ChatGPT Voice now also naturally supports real-time interpretation. With this update, it can listen to and process the user's speech, translating in line with the flow of conversation.



OpenAI stated, "GPT Live is an important step toward advancing both the naturalness and intelligence of voice AI," adding, "We plan to continue developing technology so that tasks users can perform with text, and requests to AI agents, can be carried out just as naturally in voice conversations."


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing