I kind of assumed the model would process the text 'directly', from what I understand, wouldn't this be biasing the input based on how you tokenise as it's lossy?
I assume this tradeoff is purely for speed/compression. Or am I missing what's going on here?
5 comments
[ 4.8 ms ] story [ 53.0 ms ] threadI assume this tradeoff is purely for speed/compression. Or am I missing what's going on here?
Cool project nonetheless, I will go through the code later tomorrow