“A single country” is pretty vague. For some definitions the EU already is one, for other definitions the EU can never be one. What we do need is doing more things on the EU level. E.g. there’s no point in all member…
> It doesn't have to. Of course it has to. The subagent might have misunderstood and spent 250k tokens smoking crack. If that output is blindly trusted the parent agent might go on to burn millions of tokens going in…
CoT is performative and doesn’t reveal how reasoning happens. If you look at those traces locally it’s just gibberish, especially if the model falls into a loop. https://arxiv.org/abs/2605.11746
> that sub-agent can go consume 250k+ context to return an answer that might be a couple of words How can the parent agent verify the answer without reading some of the context of the sub-agent?
Three fifty = 350? It’s not an unusual way to save effort when speaking.
Still waiting for our org to roll out Mythos. I guess it was too expensive so we’re stuck on the previous model until the internal team can figure out self-hosting open models.
> this is cargo-culting the existing ways of working. I think you mean the existing anti-patterns. They don’t call them ivory tower architects for nothing.
> The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and…
Assuming you’re in control of the test data set, you do know if a task is unsolvable. At that point you can reward the model based on how quickly they give up.
Yes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.
It was replaced by other mechanisms. It’s not literally zero any kind of reserves.
> the deal allows the plant to operate beyond 2030 according to the operator. That’s just PR spin. There is no way it would have shut down in 4 years, there has literally been no talk about doing so and all the…
You joke but even though there’s not an app per se, Russia is using the “gig economy” to run sabotage ops in Europe.
> Well, I would care, if they soon would possess the capability to hack into the nuclear arsenal and kill humanity. Or make all autonomous cars crash. Why hasn’t a human already done these things? Why is AI magical?
You’re painting a pretty dystopian picture of a hyper capitalist society where the rich have shaped the law in a way that the poor have no recourse. I think it could happen but it’s completely unrelated to anything AI…
> I would take a human with a 10% mistake rate but who can be held accountable, over an AI with a 0.1% mistake rate that is totally unaccountable for its mistakes. You would take those odds? 100x more deaths? I agree…
Ultimately these are unserious companies ran by unserious people. They don’t even have a business plan, why would they bother with some kind of sensible security policy?
Can you elaborate why a debugging skill would save those tokens?
Why would AI ever be useful in nuclear weapons decisions? There is no need to be faster or more efficient at making that decision since if we need to make the decision all is already lost.
Why do you have a debugging skill? Just tell it to read the docs. Skills are for packaging instructions for how to interact with your organizations homebrew process and tools. By definition skills shouldn’t be useful…
Why would the AI need to be accountable? Make the org that sells the tokens accountable.
Having expensive high tech robots picking up litter while there is an ongoing social crisis like homelessness sounds pretty dystopian.
I think you went from one extreme to another. OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and…
> and it’s often just as good as Sonnet. I think optimizing token usage is a good exercise. We often assume a model will be terrible, when it really isn’t. “Often” doesn’t sound great. If the smaller model fails then I…
I don’t think I understand. Why would faster token generation burn more tokens? The LLM should not be generating anything in between tool calls so the only difference should be that the human waits less between turns.
“A single country” is pretty vague. For some definitions the EU already is one, for other definitions the EU can never be one. What we do need is doing more things on the EU level. E.g. there’s no point in all member…
> It doesn't have to. Of course it has to. The subagent might have misunderstood and spent 250k tokens smoking crack. If that output is blindly trusted the parent agent might go on to burn millions of tokens going in…
CoT is performative and doesn’t reveal how reasoning happens. If you look at those traces locally it’s just gibberish, especially if the model falls into a loop. https://arxiv.org/abs/2605.11746
> that sub-agent can go consume 250k+ context to return an answer that might be a couple of words How can the parent agent verify the answer without reading some of the context of the sub-agent?
Three fifty = 350? It’s not an unusual way to save effort when speaking.
Still waiting for our org to roll out Mythos. I guess it was too expensive so we’re stuck on the previous model until the internal team can figure out self-hosting open models.
> this is cargo-culting the existing ways of working. I think you mean the existing anti-patterns. They don’t call them ivory tower architects for nothing.
> The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and…
Assuming you’re in control of the test data set, you do know if a task is unsolvable. At that point you can reward the model based on how quickly they give up.
Yes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.
It was replaced by other mechanisms. It’s not literally zero any kind of reserves.
> the deal allows the plant to operate beyond 2030 according to the operator. That’s just PR spin. There is no way it would have shut down in 4 years, there has literally been no talk about doing so and all the…
You joke but even though there’s not an app per se, Russia is using the “gig economy” to run sabotage ops in Europe.
> Well, I would care, if they soon would possess the capability to hack into the nuclear arsenal and kill humanity. Or make all autonomous cars crash. Why hasn’t a human already done these things? Why is AI magical?
You’re painting a pretty dystopian picture of a hyper capitalist society where the rich have shaped the law in a way that the poor have no recourse. I think it could happen but it’s completely unrelated to anything AI…
> I would take a human with a 10% mistake rate but who can be held accountable, over an AI with a 0.1% mistake rate that is totally unaccountable for its mistakes. You would take those odds? 100x more deaths? I agree…
Ultimately these are unserious companies ran by unserious people. They don’t even have a business plan, why would they bother with some kind of sensible security policy?
Can you elaborate why a debugging skill would save those tokens?
Why would AI ever be useful in nuclear weapons decisions? There is no need to be faster or more efficient at making that decision since if we need to make the decision all is already lost.
Why do you have a debugging skill? Just tell it to read the docs. Skills are for packaging instructions for how to interact with your organizations homebrew process and tools. By definition skills shouldn’t be useful…
Why would the AI need to be accountable? Make the org that sells the tokens accountable.
Having expensive high tech robots picking up litter while there is an ongoing social crisis like homelessness sounds pretty dystopian.
I think you went from one extreme to another. OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and…
> and it’s often just as good as Sonnet. I think optimizing token usage is a good exercise. We often assume a model will be terrible, when it really isn’t. “Often” doesn’t sound great. If the smaller model fails then I…
I don’t think I understand. Why would faster token generation burn more tokens? The LLM should not be generating anything in between tool calls so the only difference should be that the human waits less between turns.