Cutting LLM Batch Inference Time by Half with Dynamic Prefix Bucketing (daft.ai) 2 points by DISCURSIVE 10mo ago ↗ HN
0 comments
[ 3.9 ms ] story [ 22.6 ms ] threadNo comments yet.