I think if you could somehow start computing the resultant matrix elements as soon as you read a row/column from the input ones, you could reach their "physically possible" limit. A couch expert on computer architecture…
For GPUs, it's actually much faster than O(n^3) because computing each entry in the result matrix is independent. Hence, the problem is embarrassingly parallel in a way. I don't know how to use O() notation for GPUs but…
I think what he is trying to say is that the metrics were meant be a check, a measure of knowing, for whether we are going to meet our vision/target. Just yesterday, at our dev team at a trading firm, we went through…
I think if you could somehow start computing the resultant matrix elements as soon as you read a row/column from the input ones, you could reach their "physically possible" limit. A couch expert on computer architecture…
For GPUs, it's actually much faster than O(n^3) because computing each entry in the result matrix is independent. Hence, the problem is embarrassingly parallel in a way. I don't know how to use O() notation for GPUs but…
I think what he is trying to say is that the metrics were meant be a check, a measure of knowing, for whether we are going to meet our vision/target. Just yesterday, at our dev team at a trading firm, we went through…