This must be satire? Makes no sense. OpenAI isn't worth anywhere close to $800B. The gap to open models is shrinking fast, and it seems that many companies will have models with around the same capabilities. And have…
If the benchmarks are a real indication, we now have a local model that is runnable on a high-end personal PC that trades blows with the leading model Claude Opus 4.6 Max from half a year ago. Insane if that is the…
llama.cpp has router mode these days, swapping out models for you, you no longer need llama-swap.
But also remember that while the API cost might be >50x the subscription cost, the actual cost to run it is much much cheaper. It will only become cheaper as models become more efficient.
That is the most recent Qwen and Google models, there is no newer version, yet. Qwen3.8 27B might come in a couple of days tho, if it's launched alongside the large one when the Qwen3.8 countdown reaches zero.
Well, I can run some models that are better than some of the weaker and cheaper Anthropic models locally, like Haiku 4.5, and solve tasks that would cost ~4500$ every day in tokens, so yeah, they are definitely…
You can get easily over 100 tok/s on the Gemma 4 31B QAT if you enable MTP. Same goes for the 27B. I'm getting 880 tok/s of throughout on a single RTX 5090 for batch tasks.
While most don't, there are developers who actually do understand systems from end to end. Most developers are terrible developers, compared to the really good ones.
I've been running a small game dev studio for ~20 years, and the one change I think must be made, is to ban the usage of "buy" when it comes to games. Games are licensed, not bought, and that should be crystal clear to…
A common extreme misconception is that inference is expensive and that providers are loosing a lot of money. Inference is extremely lucrative and profitable.
But you are talking about Europe. If the US and Europe were to cut all ties they would both face some serious consequences, and the US wouldn't be able to do anything about that, not strategically or military. The US…
Hey! We Norwegians have on average more ownership in US tech companies than Americans! Everything is going according to plan. When we have majority, we'll move the companies to Norway, but don't tell anyone, this is…
To be fair, it would be a bigger issue for the US. No country is more economically dependant on the rest of the world than the US. The US is living on the USD, and if others stop using it the US would have to do extreme…
The price, processed tokens, and output can be anything, it just depends on what GPU it is. Nvidia GPUs are much more efficient than Apple hardware for inference(and training).
Cause I'm not trying to fool you in order to make money off you, and it should not be hard to understand that it is impossible to create a usable reply to a text someone sends you if you can't see the text.
The proper way to implement it is to issue digital IDs and use ZK proofs to verify the age. That way the service doesn't know anything other that the fact that you have an official digital ID and that you are at least a…
A dark, but not totally unfair take: It makes it easier for Apple to take payment for the models others provide, and even allows Apple, if they want to, to use the data to build a dataset for training their own models…
> The company reiterated that Apple Intelligence relies on on-device processing and Private Cloud Compute, with a promise that user data is only used to execute the immediate request and is not accessible to Apple or…
Most people seems to have switched to Valkey, and it's backed by the Linux foundation.
Hacker news doesn't generate much traffic, despite what people are saying. The host here has a limit of 160000 files served each day. That is extremely low. If the site has an icon, css, a js file and a few images it's…
While Google does a good job with language support in their models, GPT-5.5 can't write proper Norwegian. It's even making up words that does not exist.
Considering the fact that the US is complaining about Norway putting too much money into the US market, imagine what would happen if all that money was spent in Norway. It would be chaos.
Noctua wants their fans to last for many years, spinning at 2K rpm, with heat. Being able to produce something with lower tolerance is one thing. Making it work long term at ~10 m/s and ~200G is another thing. Have you…
A simple LAMP stack can give you <1ms page loads for most applications. Even 15 years ago you could serve over 100k/rps of lightweight PHP pages from a low end server.
Just to point it out, Cursor has not made any good models themselves. Composer 2 is Kimi K2.5, and they tried to pass it as their own until people noticed that the api specified it as Kimi.
This must be satire? Makes no sense. OpenAI isn't worth anywhere close to $800B. The gap to open models is shrinking fast, and it seems that many companies will have models with around the same capabilities. And have…
If the benchmarks are a real indication, we now have a local model that is runnable on a high-end personal PC that trades blows with the leading model Claude Opus 4.6 Max from half a year ago. Insane if that is the…
llama.cpp has router mode these days, swapping out models for you, you no longer need llama-swap.
But also remember that while the API cost might be >50x the subscription cost, the actual cost to run it is much much cheaper. It will only become cheaper as models become more efficient.
That is the most recent Qwen and Google models, there is no newer version, yet. Qwen3.8 27B might come in a couple of days tho, if it's launched alongside the large one when the Qwen3.8 countdown reaches zero.
Well, I can run some models that are better than some of the weaker and cheaper Anthropic models locally, like Haiku 4.5, and solve tasks that would cost ~4500$ every day in tokens, so yeah, they are definitely…
You can get easily over 100 tok/s on the Gemma 4 31B QAT if you enable MTP. Same goes for the 27B. I'm getting 880 tok/s of throughout on a single RTX 5090 for batch tasks.
While most don't, there are developers who actually do understand systems from end to end. Most developers are terrible developers, compared to the really good ones.
I've been running a small game dev studio for ~20 years, and the one change I think must be made, is to ban the usage of "buy" when it comes to games. Games are licensed, not bought, and that should be crystal clear to…
A common extreme misconception is that inference is expensive and that providers are loosing a lot of money. Inference is extremely lucrative and profitable.
But you are talking about Europe. If the US and Europe were to cut all ties they would both face some serious consequences, and the US wouldn't be able to do anything about that, not strategically or military. The US…
Hey! We Norwegians have on average more ownership in US tech companies than Americans! Everything is going according to plan. When we have majority, we'll move the companies to Norway, but don't tell anyone, this is…
To be fair, it would be a bigger issue for the US. No country is more economically dependant on the rest of the world than the US. The US is living on the USD, and if others stop using it the US would have to do extreme…
The price, processed tokens, and output can be anything, it just depends on what GPU it is. Nvidia GPUs are much more efficient than Apple hardware for inference(and training).
Cause I'm not trying to fool you in order to make money off you, and it should not be hard to understand that it is impossible to create a usable reply to a text someone sends you if you can't see the text.
The proper way to implement it is to issue digital IDs and use ZK proofs to verify the age. That way the service doesn't know anything other that the fact that you have an official digital ID and that you are at least a…
A dark, but not totally unfair take: It makes it easier for Apple to take payment for the models others provide, and even allows Apple, if they want to, to use the data to build a dataset for training their own models…
> The company reiterated that Apple Intelligence relies on on-device processing and Private Cloud Compute, with a promise that user data is only used to execute the immediate request and is not accessible to Apple or…
Most people seems to have switched to Valkey, and it's backed by the Linux foundation.
Hacker news doesn't generate much traffic, despite what people are saying. The host here has a limit of 160000 files served each day. That is extremely low. If the site has an icon, css, a js file and a few images it's…
While Google does a good job with language support in their models, GPT-5.5 can't write proper Norwegian. It's even making up words that does not exist.
Considering the fact that the US is complaining about Norway putting too much money into the US market, imagine what would happen if all that money was spent in Norway. It would be chaos.
Noctua wants their fans to last for many years, spinning at 2K rpm, with heat. Being able to produce something with lower tolerance is one thing. Making it work long term at ~10 m/s and ~200G is another thing. Have you…
A simple LAMP stack can give you <1ms page loads for most applications. Even 15 years ago you could serve over 100k/rps of lightweight PHP pages from a low end server.
Just to point it out, Cursor has not made any good models themselves. Composer 2 is Kimi K2.5, and they tried to pass it as their own until people noticed that the api specified it as Kimi.