I think the key application here is spec decoding. This solves the deprecation problem pretty nicely, and will likely have huge performance benefits
AMD actually also does inference on PL, with reasonable commercial success, actually. Have a look at FINN.
I think the key application here is spec decoding. This solves the deprecation problem pretty nicely, and will likely have huge performance benefits
AMD actually also does inference on PL, with reasonable commercial success, actually. Have a look at FINN.