Converting In-Context Learning to Weights in Linearized-Attention Transformers (arxiv.org) 4 points by PaulHoule 2y ago ↗ HN
[–] jonrouach 2y ago ↗ is that a method to save context tokens by "baking" the attended in-context-learned into the weights? Im missing one step of so-what in the paper abstract..
1 comment
[ 3.7 ms ] story [ 17.2 ms ] thread