Cerebras Inference now runs Llama 3.1-70B at 2100 tokens/s (cerebras.ai) 6 points by cs-fan-101 1y ago ↗ HN
0 comments
[ 3.8 ms ] story [ 7.3 ms ] threadNo comments yet.