Zoology 1: Measuring and Improving Recall in Efficient Language Models (hazyresearch.stanford.edu) 2 points by convexstrictly 2y ago ↗ HN
[–] convexstrictly 2y ago ↗ Research suggesting that much of the power of the transformer architecture comes from associative recall over long sequences that does not require scaling model dimensions. They design state space models that narrow the gap.Overview https://hazyresearch.stanford.edu/blog/2023-12-11-zoology0-i...Zoology 2 https://hazyresearch.stanford.edu/blog/2023-12-11-zoology2-b...Monarchs and Butterflies: Towards Sub-Quadratic Scaling in Model Dimension. Scaling in model dimension as opposed to sequence dimension scaling in the previous posts. https://hazyresearch.stanford.edu/blog/2023-12-11-truly-subq...
1 comment
[ 11.8 ms ] story [ 37.5 ms ] threadOverview https://hazyresearch.stanford.edu/blog/2023-12-11-zoology0-i...
Zoology 2 https://hazyresearch.stanford.edu/blog/2023-12-11-zoology2-b...
Monarchs and Butterflies: Towards Sub-Quadratic Scaling in Model Dimension. Scaling in model dimension as opposed to sequence dimension scaling in the previous posts. https://hazyresearch.stanford.edu/blog/2023-12-11-truly-subq...