1 comment

[ 0.23 ms ] story [ 5.4 ms ] thread
token-based pricing is much easier to reason about than GPU-time quotas. This feels like the right direction for ollama cloud