Yes, we operate on GB200s and GH200s. Usually we are cheaper for many models and can get up to double the TPS.
Yep, we are actively working on getting this down. We can meet SLAs with tuning for the real time vision workloads but trying to get rid of this compromise is our next big development task.
That's really awesome to hear!!
Also curious about this. We have a 30 day content retention policy and have to have access to your fine-tuned model/LoRa if deploying that. If there's anything we can change, happy to hear it out.
our SLA is actually higher and we are lower priced. We are also using this as a step into serving finetuned models for much cheaper than Fireworks/Together and not having the horrible cold starts of Modal. We're…
Hey Oras, thank you for the feedback! I think we definitely could list on OpenRouter but as you point out, our end goal is to host finetuned models for individuals. The IonRouter product is mostly to showcase our…
Yes, we operate on GB200s and GH200s. Usually we are cheaper for many models and can get up to double the TPS.
Yep, we are actively working on getting this down. We can meet SLAs with tuning for the real time vision workloads but trying to get rid of this compromise is our next big development task.
That's really awesome to hear!!
Also curious about this. We have a 30 day content retention policy and have to have access to your fine-tuned model/LoRa if deploying that. If there's anything we can change, happy to hear it out.
our SLA is actually higher and we are lower priced. We are also using this as a step into serving finetuned models for much cheaper than Fireworks/Together and not having the horrible cold starts of Modal. We're…
Hey Oras, thank you for the feedback! I think we definitely could list on OpenRouter but as you point out, our end goal is to host finetuned models for individuals. The IonRouter product is mostly to showcase our…