My small org has definitely had internal discussions around self-hosting gitlab. We'll see what happens.
Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit? Is this the exact same model just with less VRAM allocated for context window?
I generally disagree with the claim that distilling a model is equivalent to training on freely available internet content – mostly due to investment required to turn it into a model – but piracy is another story.…
Similar to balls to the walls (or similar). Coming from the analog aviation controls I believe.
Great historical science book on the steam engine here: https://www.amazon.com/Most-Powerful-Idea-World-Invention/dp... One of my favorites. Good read.
My small org has definitely had internal discussions around self-hosting gitlab. We'll see what happens.
Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit? Is this the exact same model just with less VRAM allocated for context window?
I generally disagree with the claim that distilling a model is equivalent to training on freely available internet content – mostly due to investment required to turn it into a model – but piracy is another story.…
Similar to balls to the walls (or similar). Coming from the analog aviation controls I believe.
Great historical science book on the steam engine here: https://www.amazon.com/Most-Powerful-Idea-World-Invention/dp... One of my favorites. Good read.