Cutting LLM Batch Inference Time in Half: Dynamic Prefix Bucketing at Scale (daft.ai) 5 points by ykev 10mo ago ↗ HN
1 comment
[ 0.26 ms ] story [ 12.3 ms ] thread