This is so interesting, thanks for sharing. I hope I'll find the time to really dig into this someday!
These token numbers look off for a 5090. For comparison, an rtx 5000 Blackwell SFF with just 470gb/s bandwidth gets me 30 tok/s on gemma4-31B-QAT. Almost 60 tok/s with MTP enabled. A 5090 should get you much more than…
"Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead." It just baffles me that this even works at all. Doesn't it just become a token soup? Is attention that…
I'm no expert analyst on this topic, but I'm worried this time might be different for them. China finally has the technological capability to challenge this time around and they will gobble up the unmet demand, allowing…
I think you are spot on and a lot of other comments sharing "I'm also so precise, and people don't get it and it's frustrating" are in fact the problem. It's arrogant to think you're that eloquent that there is not…
This is so interesting, thanks for sharing. I hope I'll find the time to really dig into this someday!
These token numbers look off for a 5090. For comparison, an rtx 5000 Blackwell SFF with just 470gb/s bandwidth gets me 30 tok/s on gemma4-31B-QAT. Almost 60 tok/s with MTP enabled. A 5090 should get you much more than…
"Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead." It just baffles me that this even works at all. Doesn't it just become a token soup? Is attention that…
I'm no expert analyst on this topic, but I'm worried this time might be different for them. China finally has the technological capability to challenge this time around and they will gobble up the unmet demand, allowing…
I think you are spot on and a lot of other comments sharing "I'm also so precise, and people don't get it and it's frustrating" are in fact the problem. It's arrogant to think you're that eloquent that there is not…