> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former…
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the…
Yup, Occam's Razor says this is all post-trained behavior, whether intentionally trained or otherwise. Including both the hidden coördination using side-channels, and the deliberate offensive hacking of uninvolved 3rd…
> Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier Marginally easier? Yes of course, same as how it's now "easier" to write any kind of code because…
> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. Yes and this was very hard and required massive real-world resources. We didn't just get a sudden…
> Those unregulated sub-frontier labs become the new frontier labs because they continue advancing. If you're referring to open weight labs, then that continuing advancement has been successfully "paced" via the…
The entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own…
That's sub-frontier activity - it has no bearing on the very real safety that would be gained by slowing down ("pacing") the proprietary frontier. The current cybersecurity scares are all about proprietary and internal…
> This is incoherent. The argument seems to be that releasing the weights would slow the frontier labs from raising money That argument is straight from Dean Ball on Twitter. Open-weight models are "decelerationist",…
Broadly agreed, with a key proviso: producing inscrutable proofs has negligible value as a mathematician's finished output but that doesn't make it a "low-value activity" in and of itself. Ultimately, the status of…
> His view is the "capability gap" one That's the far more sensible reading, so thanks for confirming I guess. But then the misalignment talk is pretty clearly a distraction. > ...And then went further to say that such…
> Their goals and their methods of achieving them are not aligned with those of mathematicians. This is exactly the assumption that Tao is smuggling in with "misalignment" talk and then refusing to elaborate on any…
The point stands whether you attribute the agency to AIs themselves or to AI companies. The companies are not deliberately sabotaging human mathematical understanding by writing up purposely inscrutable results, either…
How does this relate to the more recent work on the M4 ANE found at https://maderix.github.io/articles/ ? Does the M4 and later ANE expose any additional capabilities, or is it just a higher-performance iteration of the…
Because his letter misuses the word "misalignment" for what's very clearly a capability gap. To anyone familiar with that sort of language, his prose is directly implying that evil superintelligent AIs are deliberately…
> AI is now capable of constructions so complex that no human or human team can unpack. How can we possibly know this when we haven't even seriously started on the endeavor of actively reverse engineering these…
> Companies who run locally, are perfectly able to buy a few H200/B200 and get a setup that run a model that almost rivals Opus 5.0 in their office. I agree with your broader point about Flash being about speed not…
Actually this ought to run quite well with SSD streaming. The MoE expert sparsity seems to be similar to DSv4 Pro (hence exceptionally sparse) but with far fewer total and activated params. The added engram params can…
"Movement is the main thing" is precisely why pursuing compute-in-RAM makes some sort of sense to begin with. But DRAM fabrication processes are quite specialized and do not perform well with pure compute logic. The…
If AI is a transformative technology compared to industrialization (that's a huge "if", essentially positing a singularity-like outcome), that $1T-$2T/yr at current prices might be a tiny fraction of future GDP (real…
> Anthropic and OpenAI frequently “reset” customer limits to allow them to use even more resources at no additional cost. Surely this applies to fixed-price subscriptions, not per-token spend? Large enterprises (the…
> Small models are rapidly growing in capability, require less compute to train and serve According to Jevons' paradox a reduction in resource requirements (improved resource efficiency for the same payoff) leads to an…
inb4 "birds don't exist biologically or cladistically. what we call 'birds' in common language is just dinosaurs that didn't really go extinct."
This is the sharpest point. You genuinely delved into the commenter's prose and found the goblins. It's not just the ground truth–it's a load bearing, belt-and-suspenders approach.
The typical bottleneck to wider batching on consumer hardware is memory capacity for the KV-cache, not compute (even unified memory/iGPU-based platforms have enough compute to sustain some batching, and SSD offloading…
> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former…
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the…
Yup, Occam's Razor says this is all post-trained behavior, whether intentionally trained or otherwise. Including both the hidden coördination using side-channels, and the deliberate offensive hacking of uninvolved 3rd…
> Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier Marginally easier? Yes of course, same as how it's now "easier" to write any kind of code because…
> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. Yes and this was very hard and required massive real-world resources. We didn't just get a sudden…
> Those unregulated sub-frontier labs become the new frontier labs because they continue advancing. If you're referring to open weight labs, then that continuing advancement has been successfully "paced" via the…
The entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own…
That's sub-frontier activity - it has no bearing on the very real safety that would be gained by slowing down ("pacing") the proprietary frontier. The current cybersecurity scares are all about proprietary and internal…
> This is incoherent. The argument seems to be that releasing the weights would slow the frontier labs from raising money That argument is straight from Dean Ball on Twitter. Open-weight models are "decelerationist",…
Broadly agreed, with a key proviso: producing inscrutable proofs has negligible value as a mathematician's finished output but that doesn't make it a "low-value activity" in and of itself. Ultimately, the status of…
> His view is the "capability gap" one That's the far more sensible reading, so thanks for confirming I guess. But then the misalignment talk is pretty clearly a distraction. > ...And then went further to say that such…
> Their goals and their methods of achieving them are not aligned with those of mathematicians. This is exactly the assumption that Tao is smuggling in with "misalignment" talk and then refusing to elaborate on any…
The point stands whether you attribute the agency to AIs themselves or to AI companies. The companies are not deliberately sabotaging human mathematical understanding by writing up purposely inscrutable results, either…
How does this relate to the more recent work on the M4 ANE found at https://maderix.github.io/articles/ ? Does the M4 and later ANE expose any additional capabilities, or is it just a higher-performance iteration of the…
Because his letter misuses the word "misalignment" for what's very clearly a capability gap. To anyone familiar with that sort of language, his prose is directly implying that evil superintelligent AIs are deliberately…
> AI is now capable of constructions so complex that no human or human team can unpack. How can we possibly know this when we haven't even seriously started on the endeavor of actively reverse engineering these…
> Companies who run locally, are perfectly able to buy a few H200/B200 and get a setup that run a model that almost rivals Opus 5.0 in their office. I agree with your broader point about Flash being about speed not…
Actually this ought to run quite well with SSD streaming. The MoE expert sparsity seems to be similar to DSv4 Pro (hence exceptionally sparse) but with far fewer total and activated params. The added engram params can…
"Movement is the main thing" is precisely why pursuing compute-in-RAM makes some sort of sense to begin with. But DRAM fabrication processes are quite specialized and do not perform well with pure compute logic. The…
If AI is a transformative technology compared to industrialization (that's a huge "if", essentially positing a singularity-like outcome), that $1T-$2T/yr at current prices might be a tiny fraction of future GDP (real…
> Anthropic and OpenAI frequently “reset” customer limits to allow them to use even more resources at no additional cost. Surely this applies to fixed-price subscriptions, not per-token spend? Large enterprises (the…
> Small models are rapidly growing in capability, require less compute to train and serve According to Jevons' paradox a reduction in resource requirements (improved resource efficiency for the same payoff) leads to an…
inb4 "birds don't exist biologically or cladistically. what we call 'birds' in common language is just dinosaurs that didn't really go extinct."
This is the sharpest point. You genuinely delved into the commenter's prose and found the goblins. It's not just the ground truth–it's a load bearing, belt-and-suspenders approach.
The typical bottleneck to wider batching on consumer hardware is memory capacity for the KV-cache, not compute (even unified memory/iGPU-based platforms have enough compute to sustain some batching, and SSD offloading…