Google breaks the trillion-parameter ceiling with the Switch Transformer (arxiv.org) 48 points by groar 5y ago ↗ HN
[–] mensetmanusman 5y ago ↗ It’s an interesting thought experiment to consider models that have more parameters than there are data points being analyzed.What does that mean? [–] rdlecler1 5y ago ↗ All networks are like this. In a network of N nodes you have N^2 potential relationships. That’s just simple paired relationships. You can still go to higher order groups. [–] igorkraw 5y ago ↗ It basically means you need to learn about double descent:https://twitter.com/hippopedoid/status/1243229024085835779and regularize your models well, implicitly or explicitly. Your model might otherwise memorize the data instead of learning features.
[–] rdlecler1 5y ago ↗ All networks are like this. In a network of N nodes you have N^2 potential relationships. That’s just simple paired relationships. You can still go to higher order groups.
[–] igorkraw 5y ago ↗ It basically means you need to learn about double descent:https://twitter.com/hippopedoid/status/1243229024085835779and regularize your models well, implicitly or explicitly. Your model might otherwise memorize the data instead of learning features.
[–] panpanna 5y ago ↗ This is impressive but also requires a lot of power to train.If this trend continues ML will soon surpass bitcoin as the worst polluter.
4 comments
[ 1.4 ms ] story [ 23.3 ms ] threadWhat does that mean?
https://twitter.com/hippopedoid/status/1243229024085835779
and regularize your models well, implicitly or explicitly. Your model might otherwise memorize the data instead of learning features.
If this trend continues ML will soon surpass bitcoin as the worst polluter.