4 comments

[ 5.6 ms ] story [ 29.0 ms ] thread
cool, whats the datasets and teacher models?
datasets: CREMA-D / RAVDESS / TESS / JL Corpus models: NVIDIA A2F - 3D, LAM audio2 expression, check the HF page for links
anywhere we can see it in action?