This solves my biggest annoyance with the current advanced voice: its speech getting interrupted by me setting yup or even background noise if loud enough
Hoping to use this for natural conversation language learning. Previous iterations of the app kept correcting my words/grammar before it got to the model, causing issues with identifying mistakes in speech
I tried it for these purposes and it didn't seem any better. I couldn't even tell it was a new model. I tested today, I'm an EU customer so maybe the rollout is delayed
Are there any open source full duplex models that are out besides PersonaPlex? There was a chinese open one, maybe Fun Audio chat or something, that said it was going to release a full duplex version but I am not sure if it did.
My dream would be open source full duplex with function calling or some kind of rudimentary text output. PersonaPlex is still interesting although it was looking like we would need to fine tune it to handle outgoing or avoid going off the rails easily.
Oh wow, I'd like this. Our current voice interactions with ChatGPT are on a 4o era model; really terrible. oAI has always been pretty cagey on the architecture of their end to end multimodal models. And RL has basically made them worse since launch. (Check the launch videos where the model sings, is more realtime, has accents, etc). I'd love to try a next gen version.
The potential conversational dynamics of people telling each other "quiet!" after they pick up the habit from talking with AI will be interesting. It could lead to people being more assertive and thoughtful, or it could be contentious and rude.
Awesome that they've improved that aspect of voice chat, though.
Absolutely can't wait to try this for language practice. The advanced voice mode is great but ultimately just doesn't work that well and doesn't have the feel of a natural conversation.
This looks very cool. An AI that can listen and speak and handle tasks without breaking the flow of conversation would solve some big annoyances with current tools.
The concern is though as these get better will people struggle to distinguish these with real human connections?
I like this and felt like some of it was much more fluid; but was I alone in feeling like the interjected "uh-huh" or "yeah?" moments felt a little jarring?
Almost felt a bit *uncanny valley* for what "natural" conversation is supposed to be like. If the "uh huh" isn't timed correctly, it'll feel like a zoom call with lag.
I had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects.
The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier.
I did report a fun bug with it though: it was interrupting me and laughing at my (not really intended as) jokes while I was still talking! They seem to have clamped that behavior down thankfully, it felt a bit rude and condescending.
I do not fully understand the complexity behind achieving full-duplex but I hope this sets the bar for Anthropic to follow. Turn-based simplex is yesterday.
What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an app or service.
It’s weird. The old Claude voice mode WAS able to use tools but when they revamped it, it lost that capability and is now pinned to Haiku :(
So, yay for finally a voice mode that’s powered by a frontier model and hopefully as good as Grok voice, but sad to still not see tool use while in voice mode.
(I haven’t tried it yet, only read the announcement)
Funny enough I have only encountered two voice modes that were okay to use tools and they are so far out there that you wouldn't even think to try. Bixby and Perplexity (in android phone digital assistant mode) both seem happy to use whatever the account is connected to. I mainly use it for managing my local phone's calendar. Claude chat can interact with it but voice can't which is frustrating.
This is something I built for myself, and to experiment with inference stacks. You can obviously just transcribe audio and hand it off to frontier models, so all you really need is a good voice stack and a “driver” for the interaction (like a phone call, place to see their work).
There are two big problems with this space IMO. One isn’t that you can’t get this to work but that people generally aren’t willing to pay for it for themselves, rather as a way to screen or automate stuff to be used by other people. Did you know Claude Code has a voice mode and that openai launched whisper a year ago, both of which have positive sentiment and adoption in heavy ai tool users? Yet it’s a blip in their marketing or why people use their products, meanwhile outside of coding, most of the biggest and highest earning AI product companies so far are voice agents targeting customer service, sales, business processes, etc.
The second is related: voice is genuinely a low-bandwidth medium, so as a primary interface for interacting with AI there is not a lot you can get out of it compared to eg complex technical work or visualizations or interactive applications. It is physically and mentally demanding to speak-aloud a highly detailed prompt fast enough that VAD won’t cut you off and you have something with comparable information density or specificity vs text. But to keep up a shorter and more natural cadence you’ll not be able to wait on a lot of thinking/tool unless you play UI tricks (ums and fillers, two models in a trench coat), break the illusion of a single coherent conversation, or take a lot of long pauses.
That’s why for the supplementary coding use case it’s mostly used for remote steering, and for general use marketed towards the large and very not-online group of people for whom typing is not a natural or common thing for them to spend their time on. Now that so much spend goes through heavily used token subscriptions and they’ve proven that kind of product, they’re not marketing “tool to get the most tokens per $ running your subscription 24/7” anymore lol.
What I’m most interested in is true “ambient” tool use against my own data or work, and for-later (or pushed live via your phone) visualizations or “five models in a trench coat but still coherent” UX, which you probably are too. But I think unless you work a lot with AI tools already it’s hard to understand how that’s any different from asking Alexa to set a timer, and either way something you’re not so desperate to have that you go looking for it, or pay smaller vendors/set up yourself.
Huh? You can expose tool functions to Gemini's live API, and it will call them as part of thinking about responses. I've used this to expose e.g. current GPS location for a smartwatch voice assistant.
watched the live translation video very impressive
Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag
not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly
I've wondered if this would happen, although doing inference directly on speech tokens would seem to imply an entirely different model (trained on lots and lots of actual speech).
GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.
I like it! I was watching a YouTube video of someone driving in a big city. I saw an interesting skyscraper and asked Chatgpt what it was. After Chatgpt answered, I asked for a photo of the building to confirm, and Chatgpt helpfully showed it in-chat. It was the correct building!
And because the voice is so frictionless to talk to, I asked about what company owns the building, then that company's industry, then how that industry works in this particular country etc. I probably wouldn't have bothered going down a rabbit hole like this if I'd had to type. Voice is much easier than typing.
Super nice to be able to finally have voice mode search stuff in the background. The more conservational vibe is also a nice touch. The real game-changer we're all waiting on, though, is being able to vibecode/deeply interact with an OS through voice mode. Until then voice mode will continue to be a niche, interesting product. Once that's implemented, though, it will be one of the most important technologies available. There's a lot of people I know that don't use a lot of the frontier capabilities of AI right now because of the required typing interface. Once that's changed and you can interact with these capabilities through voice mode (and ideally on your phone) I definitely see them and lots of people like them becoming big users of AI.
I'm so mad that this might make me re-subscribe to ChatGPT. I wouldn't have believed how much I use the voice feature before LLMs and ChatGPT currently has the best voice interface. I think Grok's interface is the next best, then Claude.
I for one am greatly looking forward to the day these kind of voice models can be run locally. It seems like the gap between open-weight and frontier is way larger for voice models than coding/language models.
There doesn’t seem to be any indication whether this is available in the chat-got app nor is there any indication in the app that anything has changed. Anyone know how to actually try this?
In the voice settings, there were 2 options - Standard and Advanced, as well as a help link saying they were in the process of rolling out the newest model -- I think the new model is named Live (and is not Standard or Advanced). The help link explained it all.
So at least for me in the app there did seem to be a way to check.
142 comments
[ 7.9 ms ] story [ 169 ms ] threadThis time is the most natural version that exists and it is a natural as a conversation.
To Downvoters: Why aren't you feeling the AGI?
My dream would be open source full duplex with function calling or some kind of rudimentary text output. PersonaPlex is still interesting although it was looking like we would need to fine tune it to handle outgoing or avoid going off the rails easily.
Awesome that they've improved that aspect of voice chat, though.
The concern is though as these get better will people struggle to distinguish these with real human connections?
One thing I noticed is that we lost vision feature for some reason on the live chat?
This was an extremely useful feature. Not sure if it’s a regional thing or that they just removed that from the current live chat.
I imagine it will be even more useful with this new version.
Almost felt a bit *uncanny valley* for what "natural" conversation is supposed to be like. If the "uh huh" isn't timed correctly, it'll feel like a zoom call with lag.
The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier.
I did report a fun bug with it though: it was interrupting me and laughing at my (not really intended as) jokes while I was still talking! They seem to have clamped that behavior down thankfully, it felt a bit rude and condescending.
The full duplex is awesome, and the feedback that it is getting what you're saying is ok, but in some of the demos was a little overkill.
I'll agree that using the "Golden Girls" was at least more entertaining than the usual pitch.
It’s weird. The old Claude voice mode WAS able to use tools but when they revamped it, it lost that capability and is now pinned to Haiku :(
So, yay for finally a voice mode that’s powered by a frontier model and hopefully as good as Grok voice, but sad to still not see tool use while in voice mode.
(I haven’t tried it yet, only read the announcement)
This is something I built for myself, and to experiment with inference stacks. You can obviously just transcribe audio and hand it off to frontier models, so all you really need is a good voice stack and a “driver” for the interaction (like a phone call, place to see their work).
There are two big problems with this space IMO. One isn’t that you can’t get this to work but that people generally aren’t willing to pay for it for themselves, rather as a way to screen or automate stuff to be used by other people. Did you know Claude Code has a voice mode and that openai launched whisper a year ago, both of which have positive sentiment and adoption in heavy ai tool users? Yet it’s a blip in their marketing or why people use their products, meanwhile outside of coding, most of the biggest and highest earning AI product companies so far are voice agents targeting customer service, sales, business processes, etc.
The second is related: voice is genuinely a low-bandwidth medium, so as a primary interface for interacting with AI there is not a lot you can get out of it compared to eg complex technical work or visualizations or interactive applications. It is physically and mentally demanding to speak-aloud a highly detailed prompt fast enough that VAD won’t cut you off and you have something with comparable information density or specificity vs text. But to keep up a shorter and more natural cadence you’ll not be able to wait on a lot of thinking/tool unless you play UI tricks (ums and fillers, two models in a trench coat), break the illusion of a single coherent conversation, or take a lot of long pauses.
That’s why for the supplementary coding use case it’s mostly used for remote steering, and for general use marketed towards the large and very not-online group of people for whom typing is not a natural or common thing for them to spend their time on. Now that so much spend goes through heavily used token subscriptions and they’ve proven that kind of product, they’re not marketing “tool to get the most tokens per $ running your subscription 24/7” anymore lol.
What I’m most interested in is true “ambient” tool use against my own data or work, and for-later (or pushed live via your phone) visualizations or “five models in a trench coat but still coherent” UX, which you probably are too. But I think unless you work a lot with AI tools already it’s hard to understand how that’s any different from asking Alexa to set a timer, and either way something you’re not so desperate to have that you go looking for it, or pay smaller vendors/set up yourself.
Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag
not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly
GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.
Would love to hear your feedback!
One group's expectation of interruption for pleasant conversational flow can be just as off-putting as another's expectation of patient silence.
And because the voice is so frictionless to talk to, I asked about what company owns the building, then that company's industry, then how that industry works in this particular country etc. I probably wouldn't have bothered going down a rabbit hole like this if I'd had to type. Voice is much easier than typing.
Anyhow it's fun! Thanks for making it!
Are we seeing any conversational layer integrated with codex soon?
In the voice settings, there were 2 options - Standard and Advanced, as well as a help link saying they were in the process of rolling out the newest model -- I think the new model is named Live (and is not Standard or Advanced). The help link explained it all.
So at least for me in the app there did seem to be a way to check.