yep, it will be through enterprise licenses and our own hosted platform built on the repo
For the online signal, we use a LLM judge with a rubric calibrated offline by the user via TUI. UX of the calibration is a major focus area. Semantic caching is interesting, open to supporting it but not currently…
Thanks! We are going to add continual RL via Tinker soon too
Router and model optimization from traffic is the main differentiator
Great question. You need the simulator to be realistic enough that the optimizer cannot reward hack, but it does not need to be perfectly realistic.
yep, it will be through enterprise licenses and our own hosted platform built on the repo
For the online signal, we use a LLM judge with a rubric calibrated offline by the user via TUI. UX of the calibration is a major focus area. Semantic caching is interesting, open to supporting it but not currently…
Thanks! We are going to add continual RL via Tinker soon too
Router and model optimization from traffic is the main differentiator
Great question. You need the simulator to be realistic enough that the optimizer cannot reward hack, but it does not need to be perfectly realistic.