They are moving to 22.04 and there's a push for 24.04. See https://github.com/tenstorrent/tt-metal/issues/10469
Don't we already do that with forwarding?
Yes. "Coarse-grain dataflow processor". It's like an FPGA. But LUTs are replaced with RISC-V cores.
Grayskull is make before LLMs being a thing. And their plan is like Groq, to distribute the compute graph across multiple processors to get higher effective memory and throughput by pipelining. But better-ish by having…
Not fully, 8 bits has 256 values. It's easy to keep a look up table in the L1 cache of any CPU and constant cache of any GPU. For ASICs and FPGAs, it's a simple 256-value LUT. It's not ideal, yes, but not a deal…
I've wanted to play with large VLIW processors (Itanium is a disaster what we don't talk about. And there's litte support anyway). It's almost my dream. I hope this is real!
They are moving to 22.04 and there's a push for 24.04. See https://github.com/tenstorrent/tt-metal/issues/10469
Don't we already do that with forwarding?
Yes. "Coarse-grain dataflow processor". It's like an FPGA. But LUTs are replaced with RISC-V cores.
Grayskull is make before LLMs being a thing. And their plan is like Groq, to distribute the compute graph across multiple processors to get higher effective memory and throughput by pipelining. But better-ish by having…
Not fully, 8 bits has 256 values. It's easy to keep a look up table in the L1 cache of any CPU and constant cache of any GPU. For ASICs and FPGAs, it's a simple 256-value LUT. It's not ideal, yes, but not a deal…
I've wanted to play with large VLIW processors (Itanium is a disaster what we don't talk about. And there's litte support anyway). It's almost my dream. I hope this is real!