2 comments

[ 4.4 ms ] story [ 17.6 ms ] thread
As usual LSTMs are shit at generating text. But attention models (like BERT) are a whole different game.
The results shouldn't be that bad. There may be a bug in the implementation.