2 comments

[ 5.3 ms ] story [ 18.5 ms ] thread
The end-to-end speech-to-speech claim is interesting, especially avoiding the ASR→LLM→TTS pipeline, which is where most latency and error compounding happens.