I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was…
It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing…
This idea of remotely hosting the agent harness is honestly backwards to what I need. In so many cases, all the friction is about how to provision access to local data so the agent can work. So you started with the…
While this result does imply there was some training on the reasoning trace and output of GPT 5.5, it doesn't tell us how much of the source of its training it was (even a small amount of post training could bump up the…
Is this just Google precomputing Alpha genome values - which were already accessible via API and making them available as another API (presumably more broadly)? Or is there actually new information?
wouldn't it be interesting if nVidia buying hugging face was part of hushing up the fallout there On the face of it, they would have very good cause for some action there, assuming they wanted to.
> Sorry, I guess we will put up better guardrails next time Or, if you are Anthropic: > This illustrates the risks posed by open models!
One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a…
Agree. It wasn't as drastic but I think Vue did the 2=>3 update at an unfortunate time. They lost a lot of users along the way but they also confused a lot of the LLM training too, I am sure. Still, I have great success…
Vue had a good tenure as a solid #2 to React, so I think it probably has got a lot of representation in training data. Things really splitered after that but it got a good foothold.
It definitely leaves a bad taste because it is completely transparent their concern is not security here and that means they are lying / misrepresenting this to our faces - which then raises the question of whether you…
I'm curious what your methodology is that results in that? Are you running multiple teams of agents all adversarially reviewing each others code? Lots of different projects in parallel? I've only rarely maxed things out…
Obviously Ed is a special case, but let's be honest, pretty much anybody telling you they can predict the future is selling you BS. And that includes all the people confidently predicting he was wrong. The actual…
not having to implement proper table auditing?
which is again where Anthropic has outplayed them. Because Claude is being happily applied everywhere as a universal term, with "Code" or "Cowork" only appended as necessary.
Yeah ... the simplest answer to what ChatGPT Work is seems to be a panic-clone of Claude Cowork as a hail mary to try and catch up to Anthropic in enterprise. Cowork is like crack cocaine to nearly every exec I've seen…
It's strange to me that there is not a more conscious call out that letting an agent edit its own behavior crosses an explicit risk threshold that requires additional controls. They happily drew the whole loop at the…
The "good enough" concept is interesting because of how systematically people over estimate it. So often, things that are lower quality but thought to be "good enough" turn out to be either not good enough or not worth…
there is an inbetween .... i insist people interactively rebase those commits out. In some contexts it is actually important to have traceability of iterative proof of work towards the final result.
Moves like this are nearly always a precursor to a user hostile change of some kind. I suppose anybody who didn't leave already are likely so rusted on they still won't budge, but people should definitely look for open…
Anyone dealing with sensitive or regulated data can get started instantly with local models where they may need time consuming process or complexities to send sensitive data outside.
The big question / risk is whether there are unknown tipping points where things run away. Many are talked about - ice melting, reducing reflectivity, more heating causes more heating etc. But some fraction of these are…
The real problem for parallel dev is ensuring parallel dev environments can seamlessly co-exist without treading on each other. As soon as one of them wants to open a port, talk to an external database or write to a…
If you initiate a wipe before approaching immigration, which swaps the whole phone contents to an encrypted backup that you physically can't decrypt without a key that (say) a friend knows. Then you would be offering…
> Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image. It's useful but for OCR and a lot of other applications it needs to…
I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was…
It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing…
This idea of remotely hosting the agent harness is honestly backwards to what I need. In so many cases, all the friction is about how to provision access to local data so the agent can work. So you started with the…
While this result does imply there was some training on the reasoning trace and output of GPT 5.5, it doesn't tell us how much of the source of its training it was (even a small amount of post training could bump up the…
Is this just Google precomputing Alpha genome values - which were already accessible via API and making them available as another API (presumably more broadly)? Or is there actually new information?
wouldn't it be interesting if nVidia buying hugging face was part of hushing up the fallout there On the face of it, they would have very good cause for some action there, assuming they wanted to.
> Sorry, I guess we will put up better guardrails next time Or, if you are Anthropic: > This illustrates the risks posed by open models!
One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a…
Agree. It wasn't as drastic but I think Vue did the 2=>3 update at an unfortunate time. They lost a lot of users along the way but they also confused a lot of the LLM training too, I am sure. Still, I have great success…
Vue had a good tenure as a solid #2 to React, so I think it probably has got a lot of representation in training data. Things really splitered after that but it got a good foothold.
It definitely leaves a bad taste because it is completely transparent their concern is not security here and that means they are lying / misrepresenting this to our faces - which then raises the question of whether you…
I'm curious what your methodology is that results in that? Are you running multiple teams of agents all adversarially reviewing each others code? Lots of different projects in parallel? I've only rarely maxed things out…
Obviously Ed is a special case, but let's be honest, pretty much anybody telling you they can predict the future is selling you BS. And that includes all the people confidently predicting he was wrong. The actual…
not having to implement proper table auditing?
which is again where Anthropic has outplayed them. Because Claude is being happily applied everywhere as a universal term, with "Code" or "Cowork" only appended as necessary.
Yeah ... the simplest answer to what ChatGPT Work is seems to be a panic-clone of Claude Cowork as a hail mary to try and catch up to Anthropic in enterprise. Cowork is like crack cocaine to nearly every exec I've seen…
It's strange to me that there is not a more conscious call out that letting an agent edit its own behavior crosses an explicit risk threshold that requires additional controls. They happily drew the whole loop at the…
The "good enough" concept is interesting because of how systematically people over estimate it. So often, things that are lower quality but thought to be "good enough" turn out to be either not good enough or not worth…
there is an inbetween .... i insist people interactively rebase those commits out. In some contexts it is actually important to have traceability of iterative proof of work towards the final result.
Moves like this are nearly always a precursor to a user hostile change of some kind. I suppose anybody who didn't leave already are likely so rusted on they still won't budge, but people should definitely look for open…
Anyone dealing with sensitive or regulated data can get started instantly with local models where they may need time consuming process or complexities to send sensitive data outside.
The big question / risk is whether there are unknown tipping points where things run away. Many are talked about - ice melting, reducing reflectivity, more heating causes more heating etc. But some fraction of these are…
The real problem for parallel dev is ensuring parallel dev environments can seamlessly co-exist without treading on each other. As soon as one of them wants to open a port, talk to an external database or write to a…
If you initiate a wipe before approaching immigration, which swaps the whole phone contents to an encrypted backup that you physically can't decrypt without a key that (say) a friend knows. Then you would be offering…
> Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image. It's useful but for OCR and a lot of other applications it needs to…