6 comments

[ 0.23 ms ] story [ 9.8 ms ] thread
Need one for vision models too tbh. Token/s doesn't really map easily
Yeah, tokens/s varies a lot based on workload. However, I’ve calibrated the estimator against public benchmarks, so it stays within a 30% error margin!

I'll definitely explore how to estimate vision models next!

How about the other way around? I'll tell you my hardware and you tell me the possible stats? I've been having some problems with bigger models, using MoE to make it work for longer/tedious work like nightly sweeps that don't really on speed.
Could you add cost estimates too? I usually care about the tradeoff between VRAM/throughput and roughly what the same workload would cost on different providers.