[–] mezark 12mo ago ↗ We look at how comparative advantage from economics applies to LLM inference - some GPUs are relatively better at FLOPs, others at memory bandwidth. What happens if you let each do what it’s best at?
1 comment
[ 4.1 ms ] story [ 34.6 ms ] thread