I don't know what to say, your demo is really bad. I cannot understand what the person was speaking after you altered the sound, this approach doesn't make sense to me.
Transcription can happen from audio on the system channel without a bot in the call, so if the remote user can hear it, a model can convert it to text.
The noises confuse the ASR models, I presume.
Someone could still record audio and manually transcribe it or take notes with a pen and pencil... just sayin'. Some of this just comes down to having to trust the other party.
I do take some issue with my voice almost certainly being used to train models when using certain services.
3 comments
[ 2.4 ms ] story [ 15.6 ms ] threadthe major call providers have settings one can use as well
The noises confuse the ASR models, I presume.
Someone could still record audio and manually transcribe it or take notes with a pen and pencil... just sayin'. Some of this just comes down to having to trust the other party.
I do take some issue with my voice almost certainly being used to train models when using certain services.