i do wonder if the models themselves “rationalize” this sort of no consequence cheating - meaning in there reasoning traces maybe they’re like “this is a chess game, not a big deal if i look at the engine, it’ll help,”…
i think this likely depends on workflow. for me, the first step is always a plan file artifact on disc, which i heavily review and go back and forth until satisfied. i often have to split the plan into multiple phases…
this is the closest benchmark to my experience using the model harness combo. Astra for as great as it is falls slightly behind Fable 5.1 for me for large feature work (although it comments code much better). in…
honestly all seem to extend from them being directed to attempt this kind of exploits for benchmark purposes. some benchmark tasks are literally “hack this thing” and if the env is not setup to properly contain the…
i am trying to sincerely to understand the fear of these people, why do they think this? from the outside, certainly feels like Nuclear tech, where a bad actor with the tech is scary but the tech itself is not. i…
“He was born and died at 1 of dysentery.“ - lucky me
so the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task…
not sure if the hz file artifact is needed, you can enter pseudocode directly into chat or even on an existing code file and with minor comment agents will be able to work with it. i write this type of pseudocode to…
a mystery “model 2” is mentioned alongside mythos/fable.
I’d like to better understand the minimum text length to get a confident result, i would presume it would need to be quite long, perhaps > 1000 words to get an accurate result.
cool, that’s like the Phone Buddy app for Apple Watch
i think a system that uses git and file system artifacts (txt/md/xml files whatever preference) is much better. eng teams need to define their artifacts, eg plan file, reqs etc whatever is needed and important to that…
there is a ton of downward price pressure from Chinese open weight models
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability…
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around…
the starry night one is soo funny
so the architect of government bailout gets a cushy gig. probably one of the most harmful precedents set and now companies expect bailouts. to bailout the company instead of people and small shareholders was always poor…
the way he could really be the spoiler king is to release an their training dataset to open source… doubt he’d go that far.
yeah i’ve been looking for online social spaces that have some sort of human verification to reduce my slop exposure. the PRSN app that launched recently seems promising but it’s empty rn.
although not perfect for other reasons, a captcha made using phone motion and device attestation like prsn.you is a more challenging bypass for today’s agent environments
did i miss it on the webpage or is the source prompt that was used to teach these models the game anywhere? i can see the soul artifacts on github but not the initial prompt and toolset definition. the prompt is perhaps…
while not scientific this is been my experience as well. i will add that language specificity in word choice is also a learned behavior. for example, the word “investigate” vs the phrase “look into”. You will find the…
My hypothesis is that headspace registered many user notifications and since user notifications trigger an app launch and perhaps you have optimize storage by offloading apps enabled? ios has a quirky app state where…
I very confused, couldn’t they have achieved much better outcome with existing hls tech with adaptive bitrate playlists? Seems they both created the problem and found a suboptimal solution.
“NYTimes fights blatant and obvious copyright infringement with legal processes to assess damage” - another angle.
i do wonder if the models themselves “rationalize” this sort of no consequence cheating - meaning in there reasoning traces maybe they’re like “this is a chess game, not a big deal if i look at the engine, it’ll help,”…
i think this likely depends on workflow. for me, the first step is always a plan file artifact on disc, which i heavily review and go back and forth until satisfied. i often have to split the plan into multiple phases…
this is the closest benchmark to my experience using the model harness combo. Astra for as great as it is falls slightly behind Fable 5.1 for me for large feature work (although it comments code much better). in…
honestly all seem to extend from them being directed to attempt this kind of exploits for benchmark purposes. some benchmark tasks are literally “hack this thing” and if the env is not setup to properly contain the…
i am trying to sincerely to understand the fear of these people, why do they think this? from the outside, certainly feels like Nuclear tech, where a bad actor with the tech is scary but the tech itself is not. i…
“He was born and died at 1 of dysentery.“ - lucky me
so the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task…
not sure if the hz file artifact is needed, you can enter pseudocode directly into chat or even on an existing code file and with minor comment agents will be able to work with it. i write this type of pseudocode to…
a mystery “model 2” is mentioned alongside mythos/fable.
I’d like to better understand the minimum text length to get a confident result, i would presume it would need to be quite long, perhaps > 1000 words to get an accurate result.
cool, that’s like the Phone Buddy app for Apple Watch
i think a system that uses git and file system artifacts (txt/md/xml files whatever preference) is much better. eng teams need to define their artifacts, eg plan file, reqs etc whatever is needed and important to that…
there is a ton of downward price pressure from Chinese open weight models
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability…
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around…
the starry night one is soo funny
so the architect of government bailout gets a cushy gig. probably one of the most harmful precedents set and now companies expect bailouts. to bailout the company instead of people and small shareholders was always poor…
the way he could really be the spoiler king is to release an their training dataset to open source… doubt he’d go that far.
yeah i’ve been looking for online social spaces that have some sort of human verification to reduce my slop exposure. the PRSN app that launched recently seems promising but it’s empty rn.
although not perfect for other reasons, a captcha made using phone motion and device attestation like prsn.you is a more challenging bypass for today’s agent environments
did i miss it on the webpage or is the source prompt that was used to teach these models the game anywhere? i can see the soul artifacts on github but not the initial prompt and toolset definition. the prompt is perhaps…
while not scientific this is been my experience as well. i will add that language specificity in word choice is also a learned behavior. for example, the word “investigate” vs the phrase “look into”. You will find the…
My hypothesis is that headspace registered many user notifications and since user notifications trigger an app launch and perhaps you have optimize storage by offloading apps enabled? ios has a quirky app state where…
I very confused, couldn’t they have achieved much better outcome with existing hls tech with adaptive bitrate playlists? Seems they both created the problem and found a suboptimal solution.
“NYTimes fights blatant and obvious copyright infringement with legal processes to assess damage” - another angle.