> The enable_thinking option is worth mentioning. Without it, the model can spend a surprising amount of the completion budget reasoning before it returns a visible answer.
Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
Good to see more developments in this space. I quite like this service, which is a little further than Hetzner and has several models to choose from: https://www.infomaniak.com/en/hosting/ai-services
This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.
Interesting. I could see them perhaps coming in competitive for models that fit into single cards? Less so playing in the big model serving league...climbing into that esp right now would be madness
Whats definitely missing: a solid (non Mistral) GDPR compliant coding plan / subscription. All offerings are either US or China based. With the newest open weights models this became really interesting imo.
Very soon we will be having reseller programs for inference, this will be just like web hosting reseller. After big players, small players will also start entering in this field.
I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.
Hetzner entering LLM inference is the cloud provider equivalent of your landlord also offering to cook you dinner. the margins on compute and the margins on food both rely on you not reading the invoice too carefully
16 comments
[ 1.5 ms ] story [ 24.0 ms ] threadStraight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
> For now, the API is fast, free, and fun to try. The next hardware announcement will tell us much more than another small model would.
Will this be the new division of labor?
Americans - best proprietary models
Chinese - best open weight models
Europeans - best / most efficient inference service
I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.