[–] nithiink 2mo ago ↗ How do you handle prompt caching? A lot of cost savings for a single model chat come from cache hits on the conversation context, and switching models invalidates that cache — the new model has to reprocess everything at full input price. [–] [dead] FrancescoMassa 2mo ago ↗ [flagged]
4 comments
[ 3.3 ms ] story [ 22.6 ms ] thread