[–] tbruckner 9mo ago ↗ A simple cue like asking the model to 'see' or 'hear' can push a purely text-trained language model towards the representations of purely image-trained or purely-audio trained encoders.
1 comment
[ 0.22 ms ] story [ 9.4 ms ] thread