Retentive Network: A Successor to Transformer for Large Language Models (arxiv.org) 10 points by vagabund 3y ago ↗ HN
[–] vagabund 3y ago ↗ Brief twitter thread with performance metrics. Looks big if results hold up.https://twitter.com/arankomatsuzaki/status/16811139775001845...
[–] fgfm 3y ago ↗ The repo the paper is pointing to indicates that the code will be released within ~1 week! If the sheer difference in VRAM requirements and latency holds up, this will seriously be a major breakthrough for LLM architectures!Can't wait to try it out [–] deimos0x02 3y ago ↗ https://github.com/Jamie-Stirling/RetNet
3 comments
[ 4.4 ms ] story [ 31.4 ms ] threadhttps://twitter.com/arankomatsuzaki/status/16811139775001845...
Can't wait to try it out