It's important to note that languages evolve much faster without a writing tradition to anchor the language down over time.
Couldn't you put some sort of faraday cage around the antenna?
pi.dev + anything else
That's AMD's fault. RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in…
Yes, it runs in vllm happily. It gets >900TPS output reliably on a single 5090 with the nvfp4 model. It's clearly worse than vanilla 26B-A4B, and lacks some things like structured outputs, and gets some tool calls…
thanks for posting your setup! I think it's smart to set the reasoning effort default to something saner in the base config. Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out: ```…
Yes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both. I gave general numbers of what I'm getting above, the performance…
I think it's a vllm vs llama_cpp performance thing, looking more into it.
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods). Output TPS in vllm for instance: - Gemma4…
Using 98.css would still leave you with the AI slop text wording. The core problem is that some people don't even seem to notice / care.
Laguna XS is MoE, however.
Interesting, what does your tasks & workflow look like with them? I generally found the quality decent (say, similar to other competitors), but the speed of task completion was very slow. I think because they would…
Does anyone here use Manus actively? I found it worse than alternatives in the similar space (Claude, genspark, kagi research, etc.) and slightly baffling they were being acquired at a $2B valuation in the first place.
The value prop really depends on what you're doing. If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month. For some…
I love these, thanks!
Everyone has to find an answer to the meaning if life. For the non religious/escapist, you have to stare into the void and find an answer at some point. Existentialism (Sartre) says you have to find your own meaning.…
The nvidia shield is pretty damn good as well, even if old at this point.
From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from…
The fact that he could only get this story published in the Washington examiner of all places should be a signal that more reputable places don't want to attach their name to this
Cerebras is a whole lot of SRAM, basically a ton more L1/L2 cache, hence increasing throughput. They're pretty supply constrained right now though and their production costs seem prohibitive. The interesting players at…
A lot of benchmarks are setup to not punish false positives (irrelevant answers or extra text) and punish false negatives (missing the snippet being looked for). This leads to answer bloat and/or hallucination if you…
gemma 12B 4bit quant; try something with MTP and an AWQ quant
On a 5090, gemma4 26B runs at 350TPS with the command below [1] and gemma4 31B is around 150TPS with a similar command. I'm really surprised how much slower a DGX spark is for the same price. 1. Here's my command.…
Cerebras are only serving kimi for dedicated endpoint customers; for that you need a >$5m annual deal with them Cerebras also seems to be killing off their regular APIs, they're deprecating models and GLM is still stuck…
Not sure if the M5 is that massively different but I have a M2 max laptop and the screen is noticeably brighter on the Asus.
It's important to note that languages evolve much faster without a writing tradition to anchor the language down over time.
Couldn't you put some sort of faraday cage around the antenna?
pi.dev + anything else
That's AMD's fault. RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in…
Yes, it runs in vllm happily. It gets >900TPS output reliably on a single 5090 with the nvfp4 model. It's clearly worse than vanilla 26B-A4B, and lacks some things like structured outputs, and gets some tool calls…
thanks for posting your setup! I think it's smart to set the reasoning effort default to something saner in the base config. Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out: ```…
Yes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both. I gave general numbers of what I'm getting above, the performance…
I think it's a vllm vs llama_cpp performance thing, looking more into it.
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods). Output TPS in vllm for instance: - Gemma4…
Using 98.css would still leave you with the AI slop text wording. The core problem is that some people don't even seem to notice / care.
Laguna XS is MoE, however.
Interesting, what does your tasks & workflow look like with them? I generally found the quality decent (say, similar to other competitors), but the speed of task completion was very slow. I think because they would…
Does anyone here use Manus actively? I found it worse than alternatives in the similar space (Claude, genspark, kagi research, etc.) and slightly baffling they were being acquired at a $2B valuation in the first place.
The value prop really depends on what you're doing. If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month. For some…
I love these, thanks!
Everyone has to find an answer to the meaning if life. For the non religious/escapist, you have to stare into the void and find an answer at some point. Existentialism (Sartre) says you have to find your own meaning.…
The nvidia shield is pretty damn good as well, even if old at this point.
From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from…
The fact that he could only get this story published in the Washington examiner of all places should be a signal that more reputable places don't want to attach their name to this
Cerebras is a whole lot of SRAM, basically a ton more L1/L2 cache, hence increasing throughput. They're pretty supply constrained right now though and their production costs seem prohibitive. The interesting players at…
A lot of benchmarks are setup to not punish false positives (irrelevant answers or extra text) and punish false negatives (missing the snippet being looked for). This leads to answer bloat and/or hallucination if you…
gemma 12B 4bit quant; try something with MTP and an AWQ quant
On a 5090, gemma4 26B runs at 350TPS with the command below [1] and gemma4 31B is around 150TPS with a similar command. I'm really surprised how much slower a DGX spark is for the same price. 1. Here's my command.…
Cerebras are only serving kimi for dedicated endpoint customers; for that you need a >$5m annual deal with them Cerebras also seems to be killing off their regular APIs, they're deprecating models and GLM is still stuck…
Not sure if the M5 is that massively different but I have a M2 max laptop and the screen is noticeably brighter on the Asus.