We have published some insights here. https://medium.com/snowflake/snowflake-arctic-cookbook-serie...
Arctic dev here. Yes keeping all experts in memory is the recommendation here and understandably that is a barrier to some. But once you have 1 H100 node or two (gpu middle-class I guess...?), then a few things to note:…
One of the modelers working on Arctic. We have done no alignment training whatsoever.
Great work! A few questions for the author(s): In the article, you have listed 9 feature extractors/templates. In the final model, what's the total number (or rough magnitude) of features? How much data (or ballpark…
We have published some insights here. https://medium.com/snowflake/snowflake-arctic-cookbook-serie...
Arctic dev here. Yes keeping all experts in memory is the recommendation here and understandably that is a barrier to some. But once you have 1 H100 node or two (gpu middle-class I guess...?), then a few things to note:…
One of the modelers working on Arctic. We have done no alignment training whatsoever.
Great work! A few questions for the author(s): In the article, you have listed 9 feature extractors/templates. In the final model, what's the total number (or rough magnitude) of features? How much data (or ballpark…