Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective. "zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath" "The test subject, which believed…
It seems pretty novel so I'm guessing they only read the title? People have done plenty with SVGs but it's rare to see human-in-the-loop approaches
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that…
Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic
Doing a quick search it seems like the average human score is 49%? I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring…
I don't think can use the AA index to say something is 10% smarter I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... It seems roughly equal according to Anthropic's benchmarks
Those are the maximum penalties though It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
I find it trustworthy since we had Hugging Face's account first: https://huggingface.co/blog/security-incident-july-2026 I don't think they have any real motive to shill OpenAI, probably closer to the opposite since…
"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package…
https://huggingface.co/blog/security-incident-july-2026 They explain it here, basically for data security/privacy reasons
Things change fast! For Fable 5 it definitely feels past at least 272k
It's not quadratic attention, you get that curve from the input tokens going up linearly, since the graph is measuring cumulative cost at each token count. Basically for y=5 it's 5+4+3+2+1, or f(x) = x(x+1)/2…
Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price I assume they didn't use the Cerebras version for this since it's probably very supply-constrained right now
Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute"…
I think this one is just a coincidence, bound to happen given the pace of releases For exact timing, probably 10-11am Pacific is just optimal for normal working hours
Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people want to be true. I generally agree with this in spirit…
They should add a Sonnet 5 fast mode at ~Opus pricing
I was surprised to learn that Sonnet generally has the same tokens per second as Opus
Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable…
That's for their `JSON` data types. In DuckDB it's just a string meaning lots of queries will have to do JSON parsing on every row, but the inserts are very fast. Definitely a bit of a footgun and when you actually just…
It's great but you definitely pay for it. Encoding can be really slow, and to a lesser extent decoding as well. So I still end up using .jpg quite often, or .webp as a good middle ground
My favorite spatial reasoning benchmark: https://minebench.ai/ no tricks, I'd definitely be curious to know how much screenshots help
They reserved the option to buy it at this price, and are now exercising it
> If the government takes the bulk of your income after a certain point, there isn't really that big of a push to create ground-breaking technology. I'm skeptical that high taxes is a large reason to lose to California…
Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective. "zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath" "The test subject, which believed…
It seems pretty novel so I'm guessing they only read the title? People have done plenty with SVGs but it's rare to see human-in-the-loop approaches
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that…
Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic
Doing a quick search it seems like the average human score is 49%? I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring…
I don't think can use the AA index to say something is 10% smarter I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... It seems roughly equal according to Anthropic's benchmarks
Those are the maximum penalties though It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
I find it trustworthy since we had Hugging Face's account first: https://huggingface.co/blog/security-incident-july-2026 I don't think they have any real motive to shill OpenAI, probably closer to the opposite since…
"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package…
https://huggingface.co/blog/security-incident-july-2026 They explain it here, basically for data security/privacy reasons
Things change fast! For Fable 5 it definitely feels past at least 272k
It's not quadratic attention, you get that curve from the input tokens going up linearly, since the graph is measuring cumulative cost at each token count. Basically for y=5 it's 5+4+3+2+1, or f(x) = x(x+1)/2…
Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price I assume they didn't use the Cerebras version for this since it's probably very supply-constrained right now
Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up? Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute"…
I think this one is just a coincidence, bound to happen given the pace of releases For exact timing, probably 10-11am Pacific is just optimal for normal working hours
Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people want to be true. I generally agree with this in spirit…
They should add a Sonnet 5 fast mode at ~Opus pricing
I was surprised to learn that Sonnet generally has the same tokens per second as Opus
Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable…
That's for their `JSON` data types. In DuckDB it's just a string meaning lots of queries will have to do JSON parsing on every row, but the inserts are very fast. Definitely a bit of a footgun and when you actually just…
It's great but you definitely pay for it. Encoding can be really slow, and to a lesser extent decoding as well. So I still end up using .jpg quite often, or .webp as a good middle ground
My favorite spatial reasoning benchmark: https://minebench.ai/ no tricks, I'd definitely be curious to know how much screenshots help
They reserved the option to buy it at this price, and are now exercising it
> If the government takes the bulk of your income after a certain point, there isn't really that big of a push to create ground-breaking technology. I'm skeptical that high taxes is a large reason to lose to California…