That does not match my experience at all. I'm sure this did happen, but what I see from agents is obsessively checking `git status` before doing anything git related.
It is a sound argument in the context of trying to estimate what it costs them to generate this specific output. They have training cost eather way.
It sounds like, of the three options given, you are squarely in the “authorities” camp (although it seems to me that you interpret the question as asking what you believe should happen, whereas it appears to be asking…
Why this is preferable? I would rather have the apps I use represented as a list of apps in my system (drawing on a couple of decades of established UX conventions for how they are displayed and how I can interact with…
> From the various reports, including by people aware of gmail's stance, gmail may have had a bug allowing it at some point. I think these are simply people who cannot believe that someone might give out an email…
Wait, what you're describing sounds like Gmail working exactly as documented. If I understand correctly, your name is something like John P. Smith, and there is also someone named John Psmith. You own…
I understand that this isn't entirely clear-cut and that people will interpret it differently, but in my opinion, there is a significant difference between "a webpage that contains text X" and "an HTML document that…
In my mind, the part that is most likely to fail is "return an HTML document that still contains the following text." Nowadays, more and more "HTML documents" do not contain any content at all, just some JavaScript that…
Since this is a long-standing, real-world problem that linters solved, how could it not be relevant to a discussion about the usefulness of linters? And the choice of whether to use tabs or spaces does not affect…
It definitely is relevant to a discussion about linters. And as I said on my comment - it's just an example, I mentioned other types of discussions as well.
Oh, absolutely. I started writing Python professionally in 2010. Whenever a new group of people started working on a Python project, one of the first things they had to decide was tabs versus spaces. Everyone would…
You both seem to be using a different definition of "singularity" from the one I'm familiar with. I've always understood it to mean a rapid feedback loop in which AI creates successive, increasingly capable generations…
> And this thread is seemingly full of people claiming AI can read it while simultaneously sharing that AI could not read the actual message, only the decoy as demonstrated in TFA. That’s 100% on the authors for failing…
I would be very willing to pay more! The choice between “you may get a correct answer, or you may get lied to, without a clear way to distinguish between the two” and “you may get a correct answer, or a clear indication…
The article is about how people decide to share links, nie whether they share links at all.
Maybe that’s because I work with agentic AI in my day job, but this seems utterly obvious to me: no reasonable person would ever claim that LLMs are better at keeping secrets or enforcing rules than human employees.…
> Do you mean things like system prompts or things in Claude.md? All of it - system prompts, user prompts, few-shot examples, Claude.md, things that an agent learned by exploring its environment... > So when I /compact…
I like to dunk on Meta as much as the next guy, but I think this makes sense: deterministic verification like this is not, and should never be, the LLM’s job. The tools it has access to should enforce the permissions…
Is it? Both supervised learning and reinforcement learning are ways of training the model, and the difference between them is not that big. I would say that innate means "in the weights", while non-innate means things…
I think this is exactly it, but let me ask another question (which is not rhetorical, I really don't know). Does the fact that one can describe what consciousness is and where it came from in humans help them to detect…
And also "instilled during their reinforcement training", and we are currently pushing planning hard there, for autonomous agents.
This was said in the context of a person predicting a stock market bust, so of course a stock market index price is the relevant number here.
It’s obviously not a new model capability. But using this well-known, existing capability to solve this particular issue is only obvious after the fact. It’s a useful trick to have in one’s toolbox, and I’m grateful to…
Performing 40 songs in exchange for a property does seem like serious effort...
The top comment categorized scraping as abuse ("abuse such as [...] scraping") - that's precisely why some accuse its author of lack of self awareness.
That does not match my experience at all. I'm sure this did happen, but what I see from agents is obsessively checking `git status` before doing anything git related.
It is a sound argument in the context of trying to estimate what it costs them to generate this specific output. They have training cost eather way.
It sounds like, of the three options given, you are squarely in the “authorities” camp (although it seems to me that you interpret the question as asking what you believe should happen, whereas it appears to be asking…
Why this is preferable? I would rather have the apps I use represented as a list of apps in my system (drawing on a couple of decades of established UX conventions for how they are displayed and how I can interact with…
> From the various reports, including by people aware of gmail's stance, gmail may have had a bug allowing it at some point. I think these are simply people who cannot believe that someone might give out an email…
Wait, what you're describing sounds like Gmail working exactly as documented. If I understand correctly, your name is something like John P. Smith, and there is also someone named John Psmith. You own…
I understand that this isn't entirely clear-cut and that people will interpret it differently, but in my opinion, there is a significant difference between "a webpage that contains text X" and "an HTML document that…
In my mind, the part that is most likely to fail is "return an HTML document that still contains the following text." Nowadays, more and more "HTML documents" do not contain any content at all, just some JavaScript that…
Since this is a long-standing, real-world problem that linters solved, how could it not be relevant to a discussion about the usefulness of linters? And the choice of whether to use tabs or spaces does not affect…
It definitely is relevant to a discussion about linters. And as I said on my comment - it's just an example, I mentioned other types of discussions as well.
Oh, absolutely. I started writing Python professionally in 2010. Whenever a new group of people started working on a Python project, one of the first things they had to decide was tabs versus spaces. Everyone would…
You both seem to be using a different definition of "singularity" from the one I'm familiar with. I've always understood it to mean a rapid feedback loop in which AI creates successive, increasingly capable generations…
> And this thread is seemingly full of people claiming AI can read it while simultaneously sharing that AI could not read the actual message, only the decoy as demonstrated in TFA. That’s 100% on the authors for failing…
I would be very willing to pay more! The choice between “you may get a correct answer, or you may get lied to, without a clear way to distinguish between the two” and “you may get a correct answer, or a clear indication…
The article is about how people decide to share links, nie whether they share links at all.
Maybe that’s because I work with agentic AI in my day job, but this seems utterly obvious to me: no reasonable person would ever claim that LLMs are better at keeping secrets or enforcing rules than human employees.…
> Do you mean things like system prompts or things in Claude.md? All of it - system prompts, user prompts, few-shot examples, Claude.md, things that an agent learned by exploring its environment... > So when I /compact…
I like to dunk on Meta as much as the next guy, but I think this makes sense: deterministic verification like this is not, and should never be, the LLM’s job. The tools it has access to should enforce the permissions…
Is it? Both supervised learning and reinforcement learning are ways of training the model, and the difference between them is not that big. I would say that innate means "in the weights", while non-innate means things…
I think this is exactly it, but let me ask another question (which is not rhetorical, I really don't know). Does the fact that one can describe what consciousness is and where it came from in humans help them to detect…
And also "instilled during their reinforcement training", and we are currently pushing planning hard there, for autonomous agents.
This was said in the context of a person predicting a stock market bust, so of course a stock market index price is the relevant number here.
It’s obviously not a new model capability. But using this well-known, existing capability to solve this particular issue is only obvious after the fact. It’s a useful trick to have in one’s toolbox, and I’m grateful to…
Performing 40 songs in exchange for a property does seem like serious effort...
The top comment categorized scraping as abuse ("abuse such as [...] scraping") - that's precisely why some accuse its author of lack of self awareness.