You can amortize memory loading with large continuous batching. I imagine more compute would help the problem for certain workloads like speculative decoding
You can amortize memory loading with large continuous batching. I imagine more compute would help the problem for certain workloads like speculative decoding