Ok, I've read the paper and now I wonder, why did they stop at the most interesting part? They did all that work to figure out that learning "base conversion" is the difficult thing for transformers. Great! But then why…
> The difference between them is only one digit (i.e., the last number). Therefore, it's not possible to tell if either value is greater or lesser by just looking at their values without knowing more information about…
> Parse the chess board: Could it be that the actual issue has to do with it having trouble with small tokens (letters, numbers)? Does it give a different result if you ask it to answer in a format like this? > Please…
what ground truth model? is there some paper i can read? can you post a link to it?
Ok, I've read the paper and now I wonder, why did they stop at the most interesting part? They did all that work to figure out that learning "base conversion" is the difficult thing for transformers. Great! But then why…
> The difference between them is only one digit (i.e., the last number). Therefore, it's not possible to tell if either value is greater or lesser by just looking at their values without knowing more information about…
> Parse the chess board: Could it be that the actual issue has to do with it having trouble with small tokens (letters, numbers)? Does it give a different result if you ask it to answer in a format like this? > Please…
what ground truth model? is there some paper i can read? can you post a link to it?