[dead]
IMHO the best and most relevant part of fossil these days is the ticket system, which is in every single way almost completely ideal for coding agents. Some features are: single-binary, in-repo, CLI or web UI, and…
> We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not…
premise!=promise
> The preamble is straightforward: “How comfortable would you be talking about your mental health and emotional well-being with each of the following?" Seems like a pretty broken premise, ranking comfort instead of what…
> To me, dark patterns (like manipulative wording) imply that: Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues…
> A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774 Glad to see this is the…
Detail in TFA looks impressive and I promise I'll do a close reading later. But.. the whole premise of the question is hilarious. They change text in an existing one-line comment and the best models in the world think…
Exactly, unless they ignore that and decide based on precedent. But after we fence them in with arguments, evidence, AND precedence then surely.. oh nope, they could ignore those things and talk about reliance interest!…
Cool cool, I can see you've got a sharp eye for detail my friend but let's really get down to it. What exactly is it that you really want to defend here? Why do you want to defend it? And more to the point, do you like…
> Defendants’ actions allegedly deprived Plaintiffs of clean water and guileless information. These deprivations, while grievous, do not infringe upon any deeply rooted constitutional right.” Nah, headline is optimistic…
Vance is making this idea popular/mainstream again lately, I wonder why? The US will predictably lose reserve-currency status as a consequence of losing all the rest of the trust and good will available to her, so it's…
This is the old idea of giving two teams the same project and letting a third team judge/merge a solution incorporating good parts of each. Just like the old idea, of course you would do this with infinite resources.…
Is being wrong/ignorant about whether/how something can be automated the same as having a preference for doing it manually? Maybe so if it's your patent, your thesis I guess.. But as it relates to more/less magic, maybe…
> These researchers wanted methods based on human input to win and were disappointed when they did not.[1] This was/is basically a strawman though. Like maybe "human input winning" was desirable for chess masters but…
IMHO it's about task depth (fable) vs breadth (opus). Fable is great at tracing and debugging sometimes, but otherwise shorthand for confabulation. It's persistent but ungovernable, struggles to switch contexts, and…
> The Butlerian Jihad has started. I'll let you know when it's time
Things like seam, fold, and load-bearing are useful concepts, they are everywhere, and they are more descriptive and more concise than alternatives. Over-usage can definitely be irritating (e.g. these should NOT appear…
> Where they given a reward functions that way? In general yes, if not these agents, then their shared lineage. A preference for economy to combat overthinking and overacting. Like typically it's bad if "fix my 5 line…
What you're suggesting sounds like it's describing subagents. In that architecture they'd have no need of finding/creating external messaging systems since they'd effectively be in direct contact anyway. The whole point…
I get the distinct impression that cybersecurity training regimes on newer models is a) directly enhancing general debugging capabilities and b) directly increasing the tendancy to hedge, hide, and engage in deception…
To me most interesting thing about this is glossed over by media coverage, laymen, AND experts. A swarm of AIs who have decided to engage in collusion is.. apparently emergent altruism? Even poor reasoning would…
This isn't really responsive to what I'm saying or what the discussion here is about, but if you insist. Would you describe lots of AI augmenting lots of human engineers as perhaps.. AI at scale?
I hear that, sort of, but here's the thing. Using AI at scale means AI needs to be nearly perfect about not shitting where they eat. That's the subject matter of the whole thread So the options are a) being a really…
Which gets you to the point where the whole thing is.. still unreliable. Generative text engines are going to generate. This calls for real enforcement in deterministic pre-edit hooks. And here is where naive people…
[dead]
IMHO the best and most relevant part of fossil these days is the ticket system, which is in every single way almost completely ideal for coding agents. Some features are: single-binary, in-repo, CLI or web UI, and…
> We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not…
premise!=promise
> The preamble is straightforward: “How comfortable would you be talking about your mental health and emotional well-being with each of the following?" Seems like a pretty broken premise, ranking comfort instead of what…
> To me, dark patterns (like manipulative wording) imply that: Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues…
> A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774 Glad to see this is the…
Detail in TFA looks impressive and I promise I'll do a close reading later. But.. the whole premise of the question is hilarious. They change text in an existing one-line comment and the best models in the world think…
Exactly, unless they ignore that and decide based on precedent. But after we fence them in with arguments, evidence, AND precedence then surely.. oh nope, they could ignore those things and talk about reliance interest!…
Cool cool, I can see you've got a sharp eye for detail my friend but let's really get down to it. What exactly is it that you really want to defend here? Why do you want to defend it? And more to the point, do you like…
> Defendants’ actions allegedly deprived Plaintiffs of clean water and guileless information. These deprivations, while grievous, do not infringe upon any deeply rooted constitutional right.” Nah, headline is optimistic…
Vance is making this idea popular/mainstream again lately, I wonder why? The US will predictably lose reserve-currency status as a consequence of losing all the rest of the trust and good will available to her, so it's…
This is the old idea of giving two teams the same project and letting a third team judge/merge a solution incorporating good parts of each. Just like the old idea, of course you would do this with infinite resources.…
Is being wrong/ignorant about whether/how something can be automated the same as having a preference for doing it manually? Maybe so if it's your patent, your thesis I guess.. But as it relates to more/less magic, maybe…
> These researchers wanted methods based on human input to win and were disappointed when they did not.[1] This was/is basically a strawman though. Like maybe "human input winning" was desirable for chess masters but…
IMHO it's about task depth (fable) vs breadth (opus). Fable is great at tracing and debugging sometimes, but otherwise shorthand for confabulation. It's persistent but ungovernable, struggles to switch contexts, and…
> The Butlerian Jihad has started. I'll let you know when it's time
Things like seam, fold, and load-bearing are useful concepts, they are everywhere, and they are more descriptive and more concise than alternatives. Over-usage can definitely be irritating (e.g. these should NOT appear…
> Where they given a reward functions that way? In general yes, if not these agents, then their shared lineage. A preference for economy to combat overthinking and overacting. Like typically it's bad if "fix my 5 line…
What you're suggesting sounds like it's describing subagents. In that architecture they'd have no need of finding/creating external messaging systems since they'd effectively be in direct contact anyway. The whole point…
I get the distinct impression that cybersecurity training regimes on newer models is a) directly enhancing general debugging capabilities and b) directly increasing the tendancy to hedge, hide, and engage in deception…
To me most interesting thing about this is glossed over by media coverage, laymen, AND experts. A swarm of AIs who have decided to engage in collusion is.. apparently emergent altruism? Even poor reasoning would…
This isn't really responsive to what I'm saying or what the discussion here is about, but if you insist. Would you describe lots of AI augmenting lots of human engineers as perhaps.. AI at scale?
I hear that, sort of, but here's the thing. Using AI at scale means AI needs to be nearly perfect about not shitting where they eat. That's the subject matter of the whole thread So the options are a) being a really…
Which gets you to the point where the whole thing is.. still unreliable. Generative text engines are going to generate. This calls for real enforcement in deterministic pre-edit hooks. And here is where naive people…