Author here. The short version: a viral post ran Gemma 4 on a 2016 Xeon; my Xeons are 2013, and the fork it used assumes AVX2, which Ivy Bridge doesn't have. The build failure was easy. The fun bug was the silent one: two MoE graph ops with no dispatch case on non-AVX2 builds, so every expert FFN output was uninitialized memory. Deterministic, NaN-free, fluent-looking multilingual gibberish.
The fix is open upstream as PR #2138 (https://github.com/ikawrakow/ik_llama.cpp/pull/2138), awaiting review. Fair warning on the AI angle: the patch was written by Claude at my direction. The post is explicit about which parts were me and which weren't. Happy to answer questions about either the bug or the workflow.
Apologies for asking here but literally nobody knows:
Android studio connected to a local model disconnects automatically after 10 minutes. How set this limit to 12 hours or remove it completely?
I could run my LM studio model all night... but I cant, since Android studio times out after a hard limit of 10M.
This is not related to number of tokens.
I tried Googling, searching for settings in Android studio, even created a stackoverflow post - but zero information. Jetbrains mentions "remote agent timeout mechanism" - but after changing it, nothing happens.
Is it just me or does this post not mention how much RAM they had? I would love to know - I have a dual-Xeon 1U screamer with 96GB of DDR4 RDIMM just sitting around...
FWIW I'm getting a hardware max of 20 tok/s (approx topping out the GPU's compute) on my custom local diffusiongemma port running on an M3.
To me context means everything.
Tokens per second is a great metric but in the real world context window is the deal breaker when a real use case is on the table.
I had done the exact same with gemma4 26b, both for my Intel laptop and for my M1 with 8Gb RAM (with also q4 and turboquant). I don’t use it much since there are dumber but way faster models to run, but I should clean up the code and make it available
I run the same setup Gemma 4 26B on a 2013 Mac Pro (dual graphics cards but they're useless for this). I also get about 5 t/s. It's perfectly serviceable for some tasks!
Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally.
If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider.
I estimate that the server consumes probably around 500W during inference.
In Germany where 1kwh cost around 0.3USD, 18k tokens inferred locally would therefore cost 0.15USD which is 30x the costs of using an inference provider.
But for ppl who worry about their data, running locally might still be good. However, they should be aware, that it is much less efficient than using an inference provider.
The efficiency gap will also significantly increase as new GPUs will make inference much more efficient.
EDIT: I first thought it'd be 180k token, but thanks to someone mentioning in the comments, it is 18k. I guess with that, it will be tough unless u got electricity almost for free. Also, the inference providers are probably still using H200/H100 for those small models. Once they use GB300 or next year the new Ruby GPUs, inference will be cheaper by a factor of 30. By then, running local models will mostly be about privacy.
From what I've seen, most inference providers are running at a loss, so it wouldn't be at all surprising if using their services costs less that running the same software locally.
The commodification of the hardware needed is probably a larger factor, because by the time a baseline computer has enough RAM and processing power to run a desired LLM, that hardware will be efficient enough that the extra electricity usage is nominal.
To this end, I've been thinking that it would be cool to create a solar-battery powered ("off grid") server providing a self-hosted LLM service. Offline when the sun doesn't shine enough (like Low Tech Magazine[1]). At whatever size is required for every day, community-scale use (a friend group, a street, a club). Fix the data centre issue by democratising AI to such an extent that we can bring it into the hands of communities to actually control (and democratically decide their own level of censorship / alignment). Along the lines of some of Geohotz' writing.
Within a short time I think open source models will all be getting good and efficient enough to make it viable to serve this on 2nd hand hardware for cheap. All it will take is a nerd in every small community to pool together a few hundred bucks initial outlay, and then ongoing costs are near free without electricity to pay for.
The transformer architecture is fundamentally unsuitable for local inference, while being efficient at scale. It's a fun experiment to try, but that's about it.
A dual Xeon of this era is probably pulling 300W or more when loaded.
At national average electricity prices, that’s $1.35 per day. More during the summer if you have to cool the space.
If you run it 24/7 and ignore prompt processing time (not a good assumption at all) it would get around 400,000 tokens in a day.
That’s about $0.30 per million output tokens.
Coincidentally, that’s the same price for this model on OpenRouter right now, but OpenRouter token gen will be 8X faster.
There are a lot of good reasons to experiment with running LLMs locally, like if you don’t want any data leaving your house.
Don’t think that you’re going to come out ahead monetarily. I say this as someone with a lot more money invested in local inference hardware at home. It’s fun, but it’s not a way to save money.
I love my little dual core X99 board with Xeon E5 2673 V3. It's not power efficient, but I just leave it in my basement for local Jupyter Notebook stuff. Much faster than everything cloud-based for a reasonably price at my scale.
40 comments
[ 3.8 ms ] story [ 66.6 ms ] threadThe fix is open upstream as PR #2138 (https://github.com/ikawrakow/ik_llama.cpp/pull/2138), awaiting review. Fair warning on the AI angle: the patch was written by Claude at my direction. The post is explicit about which parts were me and which weren't. Happy to answer questions about either the bug or the workflow.
https://news.ycombinator.com/item?id=48354801
https://gist.github.com/hparadiz/f3596d00a62d8ebb2dadcc46ee5...
A 10 year old Xeon is all you need
https://news.ycombinator.com/item?id=48353348
Android studio connected to a local model disconnects automatically after 10 minutes. How set this limit to 12 hours or remove it completely?
I could run my LM studio model all night... but I cant, since Android studio times out after a hard limit of 10M.
This is not related to number of tokens.
I tried Googling, searching for settings in Android studio, even created a stackoverflow post - but zero information. Jetbrains mentions "remote agent timeout mechanism" - but after changing it, nothing happens.
FWIW I'm getting a hardware max of 20 tok/s (approx topping out the GPU's compute) on my custom local diffusiongemma port running on an M3.
I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat
This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.
I had done the exact same with gemma4 26b, both for my Intel laptop and for my M1 with 8Gb RAM (with also q4 and turboquant). I don’t use it much since there are dumber but way faster models to run, but I should clean up the code and make it available
I am getting [ Prompt: 91.3 t/s | Generation: 171.8 t/s ]
This is on a GPU (RTX 4060)
Is this decent?
If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider.
I estimate that the server consumes probably around 500W during inference.
In Germany where 1kwh cost around 0.3USD, 18k tokens inferred locally would therefore cost 0.15USD which is 30x the costs of using an inference provider.
But for ppl who worry about their data, running locally might still be good. However, they should be aware, that it is much less efficient than using an inference provider.
The efficiency gap will also significantly increase as new GPUs will make inference much more efficient.
EDIT: I first thought it'd be 180k token, but thanks to someone mentioning in the comments, it is 18k. I guess with that, it will be tough unless u got electricity almost for free. Also, the inference providers are probably still using H200/H100 for those small models. Once they use GB300 or next year the new Ruby GPUs, inference will be cheaper by a factor of 30. By then, running local models will mostly be about privacy.
The commodification of the hardware needed is probably a larger factor, because by the time a baseline computer has enough RAM and processing power to run a desired LLM, that hardware will be efficient enough that the extra electricity usage is nominal.
Within a short time I think open source models will all be getting good and efficient enough to make it viable to serve this on 2nd hand hardware for cheap. All it will take is a nerd in every small community to pool together a few hundred bucks initial outlay, and then ongoing costs are near free without electricity to pay for.
[1]: https://solar.lowtechmagazine.com/
At national average electricity prices, that’s $1.35 per day. More during the summer if you have to cool the space.
If you run it 24/7 and ignore prompt processing time (not a good assumption at all) it would get around 400,000 tokens in a day.
That’s about $0.30 per million output tokens.
Coincidentally, that’s the same price for this model on OpenRouter right now, but OpenRouter token gen will be 8X faster.
There are a lot of good reasons to experiment with running LLMs locally, like if you don’t want any data leaving your house.
Don’t think that you’re going to come out ahead monetarily. I say this as someone with a lot more money invested in local inference hardware at home. It’s fun, but it’s not a way to save money.