The engram stuff is great because RAM is often still cheaper (or at least expandable). My company does currently look into buying some hardware as we handle confidential data and code. Qwen 3.8 Flash is viable on two…
I need to support integrators for a mobile app SDK. Bank stuff. We have one integrator which consistently uses a truly bad AI. My brain skips after two sentences already. There are always +9 questions what should just…
Did you already test TranslateGemma? I use this model for my Android Studio Translation Plugin (https://plugins.jetbrains.com/plugin/30265-localizepipe) and so far it produces great results for its size. If there are…
Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or…
Channels like project farm https://youtube.com/@projectfarm or other reviewers that are not sponsored are truly my main source of information in this age. Some direct reviews between 2 and 4 stars are also sometimes…
If there are no changes to the system message, yes this is possible and also likely done by Anthropic. When there are additional local MCPs, reusability will be lower. I think that Anthropic will bill you in any case :D
right, but the cache retention time is very short for Anthropics LLMs. 5 minutes or 1 hour (with additional costs). So you have to prompt basically non stop to not get a cache eviction. Anthropic even changed this…
The OpenCode CLI does not work as well for me as the PI CLI. I'm a subscriber of OpenCode Go (the sub, good value for me really) but I had not great experiences with OpenCode CLI. It multiple times with different models…
Works on Brave (Chromium) with Android 16
It is quite interesting how this is handled world wide. For me PII is very sensitive and I advice people to be very cautious. Every business in the EU (were I live) also has to be very careful with such data by law.…
Right, I did swap that. Still, you have to pay that 4k then every year and give out the code. I also assume that prices will go up as no AI company (but NVIDIA -> selling shovels) is currently making any money. For some…
With parallelism of 16 you can still get around 25 to 30 tokens per user when all 16 channels are running. Not everyone will use the model at the same time but it certainly will be tight, especially for agentic coding.…
The 5h quota of Codex Pro on GPT 5.4 Medium lasts me for around an hour and a half, maybe 2 hours. And this is already the "savy" setup. Enable GPT 5.5 High fast and you will be beached in 30 minutes with active…
That is crazy. 5 years and they are already shutting down the servers? They should be forced to open up the API when they shut it down. Running a replica yourself should be pretty doable.
RAM + GPU are getting more expensive but mostly for applications that require a lot of it like AI. The hardware cost for regular applications has not vastly increased (especially when factoring in inflation). Spending…
In my experience in software architecture, drawing a diagram often saves you >60 minutes of discussion and potentially multiple meetings. This works even with a badly drawn but truthful one. Use an Ai agent + Mermaid.js…
For programmers maybe. I do this too. But think about all the regular users out there. Your dad and your mum, maybe even your grandparents. This is a huge marked too and for that we can use these special chips at scale.
The cutting edge, max size models will likely stay in the GPU space for a long time. But these models are not needed for most general requests. With a fine tuned 30B quantisized model you can serve a large portion of…
What spec of Framework Desktop do you run this on?
I'm using the Codex Business subscription (about 30€) already for multiple months. Even there they cut back on the quota. A few months back it was hard for me to reach the limit. Now it is easier. Still, in comparison…
Fully agree. I went to school in Germany and many of our textbooks were free there. Sometimes you would get a textbook that is already >= 10 years and out of shape but who cares? Especially the basic knowledge does not…
This seems quite strange to claim. Basically every city in the developed world already has power plants on the outside and a lot of wires to get the electricity in
It exists and does degrade panels but the time horizon is pretty wide. Real world data shows something like 0.5% to 0.7% degregation per year on average. At the start the degregation is higher and but it slows down with…
Especially for hot and sunny areas solar is insane. At mid day, max heat, you get the peak production and can run your AC at full throttle. That enables you to efficiently work at nice temperatures.
The article states the same solar production numbers as your comment. I agree that the headline is overly positive but the ramp up of solar can't really be denied. Change at this scale is sadly slow in this rather…
The engram stuff is great because RAM is often still cheaper (or at least expandable). My company does currently look into buying some hardware as we handle confidential data and code. Qwen 3.8 Flash is viable on two…
I need to support integrators for a mobile app SDK. Bank stuff. We have one integrator which consistently uses a truly bad AI. My brain skips after two sentences already. There are always +9 questions what should just…
Did you already test TranslateGemma? I use this model for my Android Studio Translation Plugin (https://plugins.jetbrains.com/plugin/30265-localizepipe) and so far it produces great results for its size. If there are…
Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or…
Channels like project farm https://youtube.com/@projectfarm or other reviewers that are not sponsored are truly my main source of information in this age. Some direct reviews between 2 and 4 stars are also sometimes…
If there are no changes to the system message, yes this is possible and also likely done by Anthropic. When there are additional local MCPs, reusability will be lower. I think that Anthropic will bill you in any case :D
right, but the cache retention time is very short for Anthropics LLMs. 5 minutes or 1 hour (with additional costs). So you have to prompt basically non stop to not get a cache eviction. Anthropic even changed this…
The OpenCode CLI does not work as well for me as the PI CLI. I'm a subscriber of OpenCode Go (the sub, good value for me really) but I had not great experiences with OpenCode CLI. It multiple times with different models…
Works on Brave (Chromium) with Android 16
It is quite interesting how this is handled world wide. For me PII is very sensitive and I advice people to be very cautious. Every business in the EU (were I live) also has to be very careful with such data by law.…
Right, I did swap that. Still, you have to pay that 4k then every year and give out the code. I also assume that prices will go up as no AI company (but NVIDIA -> selling shovels) is currently making any money. For some…
With parallelism of 16 you can still get around 25 to 30 tokens per user when all 16 channels are running. Not everyone will use the model at the same time but it certainly will be tight, especially for agentic coding.…
The 5h quota of Codex Pro on GPT 5.4 Medium lasts me for around an hour and a half, maybe 2 hours. And this is already the "savy" setup. Enable GPT 5.5 High fast and you will be beached in 30 minutes with active…
That is crazy. 5 years and they are already shutting down the servers? They should be forced to open up the API when they shut it down. Running a replica yourself should be pretty doable.
RAM + GPU are getting more expensive but mostly for applications that require a lot of it like AI. The hardware cost for regular applications has not vastly increased (especially when factoring in inflation). Spending…
In my experience in software architecture, drawing a diagram often saves you >60 minutes of discussion and potentially multiple meetings. This works even with a badly drawn but truthful one. Use an Ai agent + Mermaid.js…
For programmers maybe. I do this too. But think about all the regular users out there. Your dad and your mum, maybe even your grandparents. This is a huge marked too and for that we can use these special chips at scale.
The cutting edge, max size models will likely stay in the GPU space for a long time. But these models are not needed for most general requests. With a fine tuned 30B quantisized model you can serve a large portion of…
What spec of Framework Desktop do you run this on?
I'm using the Codex Business subscription (about 30€) already for multiple months. Even there they cut back on the quota. A few months back it was hard for me to reach the limit. Now it is easier. Still, in comparison…
Fully agree. I went to school in Germany and many of our textbooks were free there. Sometimes you would get a textbook that is already >= 10 years and out of shape but who cares? Especially the basic knowledge does not…
This seems quite strange to claim. Basically every city in the developed world already has power plants on the outside and a lot of wires to get the electricity in
It exists and does degrade panels but the time horizon is pretty wide. Real world data shows something like 0.5% to 0.7% degregation per year on average. At the start the degregation is higher and but it slows down with…
Especially for hot and sunny areas solar is insane. At mid day, max heat, you get the peak production and can run your AC at full throttle. That enables you to efficiently work at nice temperatures.
The article states the same solar production numbers as your comment. I agree that the headline is overly positive but the ramp up of solar can't really be denied. Change at this scale is sadly slow in this rather…