Cutting LLM inference costs by 36% with prompt caching (neradot.com) 1 points by lizakatz 15d ago ↗ HN
1 comment
[ 5.0 ms ] story [ 17.6 ms ] thread