How do you handle cached input or long-context pricing? Some providers change the rate a lot depending on that, so I’m curious how accurate the estimate stays.
Could you add cost estimates too? I usually care about the tradeoff between VRAM/throughput and roughly what the same workload would cost on different providers.
How do you handle cached input or long-context pricing? Some providers change the rate a lot depending on that, so I’m curious how accurate the estimate stays.
Could you add cost estimates too? I usually care about the tradeoff between VRAM/throughput and roughly what the same workload would cost on different providers.