There is some sense of rose-tinted glasses of pre-LLM coding. A lot of human written code, particularly at the enterprise level, was of low quality well before AI automated it.
I wanted you to be wrong, and to be able to make this an example of us over-reacting to certain trigger words created by AI, but unfortunately I just scanned the first couple paragraphs with pangram and it reported 100%…
That's a valid take. The issue I'm wrestling with is the inevitable attempts to point a finger at who is responsible when bad things happen. If you claim it is on the tech companies to know if a child is online, then…
I think you may be misunderstanding a bit. There's no forced verification. It'd be more of an RFC that gives parents the ability to communicate their underage child is using the device without revealing or verifying any…
I once heard someone suggest that this should be on the OS level and I'm slowly coming around to the idea. There should be some kind of OS level flag that can easily broadcast to products that a child is using the…
But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like…
Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
That is incredibly dismissive to people undergoing something that is causing them to temporarily have these thoughts and need intervention
LLM's are a funny technology because on the one hand this is all undeniably impressive at the rate of what's changed from them, and yet despite that I find myself disappointed by the lack of breakthroughs for things I…
In what way do they have a moat? A cursory look at https://artificialanalysis.ai/models/gpt-6-astra#intelligenc... it lands at 61, only a single point above glm 5.3 while costing significantly more. The only moat they…
Has anyone been able to get anything substantial done with Fable in the first place? I more or less had totally given up on using it since the alignment checks were so sensitive that it pretty much always threw me back…
There's a couple different ways to look at this. From a charitable point of view to ubisoft, they do not advertise support for Linux. Instead they openly state it's a window's only game. So when a third party (i.e.…
Sometimes I wonder if the anti-immigrant rhetoric that has been growing more popular is a tactical tool used by smart individuals who don't actually believe in it but instead have accepted that they would rather be…
It's becoming increasingly evident this is not how things are going to shake out. Even without frontier models, running Qwen 3.8 27B has demonstrated for me and others a "good enough" competency at general programming.…
Maybe I am misunderstanding but I don't think this is true. Indeed you pay for the inbox, but it's not the primary domain. You can reply from any of the email addresses you created (which doesn't cost extra) and it…
I think of it more as "the automation of the stackoverflow engineer". In enterprise software, there has always just been a non-negotiable large volume of code that was required to be written. This has traditionally been…
> So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621 I don't, we originally had the turing test which was designed to determine human intelligence by its ability to imitate us with natural dialogue,…
> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but…
Fun game. Would be nice if there was a single player mode that simply prints the letters and at the end shows the longest word that was missed. That way you can play it offline and place a phone in the center of the…
The impressive/surprising thing is your premise because we effectively can spin up an army of mathematicians now
> and writes entire operating systems from scratch ehhhh, we're not really there. Not saying it's impossible to reach in the future but large tasks like this are still out of scope for LLM's beyond demoing toy…
I don't feel it's meaningful to berate the point anymore about the hypocrisy of the American labs. Now as the sentiment and effort from them to push for regulation increases so does my perception of how pathetic they…
so they distilled one of the best models in the world AND released it for free to everyone. Where can I send them flowers as a thank you?
Meanwhile non-frontend folks decide to call one thing "threads" and another thing "strings" and have them be completely unrelated to each other.
On the one hand, organizations are without question using LLM's well beyond what is actually necessary, and as reality kicks in they're forced to scale back accordingly. However at the same time, on intervals counted in…
There is some sense of rose-tinted glasses of pre-LLM coding. A lot of human written code, particularly at the enterprise level, was of low quality well before AI automated it.
I wanted you to be wrong, and to be able to make this an example of us over-reacting to certain trigger words created by AI, but unfortunately I just scanned the first couple paragraphs with pangram and it reported 100%…
That's a valid take. The issue I'm wrestling with is the inevitable attempts to point a finger at who is responsible when bad things happen. If you claim it is on the tech companies to know if a child is online, then…
I think you may be misunderstanding a bit. There's no forced verification. It'd be more of an RFC that gives parents the ability to communicate their underage child is using the device without revealing or verifying any…
I once heard someone suggest that this should be on the OS level and I'm slowly coming around to the idea. There should be some kind of OS level flag that can easily broadcast to products that a child is using the…
But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like…
Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
That is incredibly dismissive to people undergoing something that is causing them to temporarily have these thoughts and need intervention
LLM's are a funny technology because on the one hand this is all undeniably impressive at the rate of what's changed from them, and yet despite that I find myself disappointed by the lack of breakthroughs for things I…
In what way do they have a moat? A cursory look at https://artificialanalysis.ai/models/gpt-6-astra#intelligenc... it lands at 61, only a single point above glm 5.3 while costing significantly more. The only moat they…
Has anyone been able to get anything substantial done with Fable in the first place? I more or less had totally given up on using it since the alignment checks were so sensitive that it pretty much always threw me back…
There's a couple different ways to look at this. From a charitable point of view to ubisoft, they do not advertise support for Linux. Instead they openly state it's a window's only game. So when a third party (i.e.…
Sometimes I wonder if the anti-immigrant rhetoric that has been growing more popular is a tactical tool used by smart individuals who don't actually believe in it but instead have accepted that they would rather be…
It's becoming increasingly evident this is not how things are going to shake out. Even without frontier models, running Qwen 3.8 27B has demonstrated for me and others a "good enough" competency at general programming.…
Maybe I am misunderstanding but I don't think this is true. Indeed you pay for the inbox, but it's not the primary domain. You can reply from any of the email addresses you created (which doesn't cost extra) and it…
I think of it more as "the automation of the stackoverflow engineer". In enterprise software, there has always just been a non-negotiable large volume of code that was required to be written. This has traditionally been…
> So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621 I don't, we originally had the turing test which was designed to determine human intelligence by its ability to imitate us with natural dialogue,…
> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but…
Fun game. Would be nice if there was a single player mode that simply prints the letters and at the end shows the longest word that was missed. That way you can play it offline and place a phone in the center of the…
The impressive/surprising thing is your premise because we effectively can spin up an army of mathematicians now
> and writes entire operating systems from scratch ehhhh, we're not really there. Not saying it's impossible to reach in the future but large tasks like this are still out of scope for LLM's beyond demoing toy…
I don't feel it's meaningful to berate the point anymore about the hypocrisy of the American labs. Now as the sentiment and effort from them to push for regulation increases so does my perception of how pathetic they…
so they distilled one of the best models in the world AND released it for free to everyone. Where can I send them flowers as a thank you?
Meanwhile non-frontend folks decide to call one thing "threads" and another thing "strings" and have them be completely unrelated to each other.
On the one hand, organizations are without question using LLM's well beyond what is actually necessary, and as reality kicks in they're forced to scale back accordingly. However at the same time, on intervals counted in…