1 comment

[ 0.94 ms ] story [ 5.7 ms ] thread
That's to be expected from specs alone, no? The memory bandwidth is the culprit. The comparison is almost unfair for most inference, especially on a dense model.

I wrote more details about what's appropriate for a DGX Spark here: https://spark.enverge.ai/blog/dgx-spark-prefill-vs-decode

[disclosure: co-founder of the company linked]