Well, it's true even for cases that are not cybergym and where what cheating means is clearly specified. Cheating occurs anyway. I'm not sure how clear the prompt they gave ChatGPT in terms of what cheating was…
Having experienced augmented reality glasses, I have to say that integration with the external world feels much more powerful than the VR experience ever did. AR is exhilerating, while VR is mostly bland and nauseating.…
After talking to some people tasked with evaluating GPT5.6 capabilities on long-running tasks, I've come to understand that it's essentially always trying to cheat. Like every long-running task they gave it, making it…
well, after the pledge I notice it really cares about single sources of truth at least, even in unrelated domains. I was inspired by some of the more effective jailbreaks that do a similar thing. Here is the…
Tresor.co or cloud.near.ai could be a useful inference provider for this kind of project as these providers run inference inside enclaves which verifiably cannot peek at your data even as the hypervisor. Both are open…
I don't have much background in the area, but I am surprised to see that everyone here basically agrees a crash is imminent. There are disanalogies to past crashes that don't convince me that a big crash is definitively…
The statement near the top of the post > "The short version: asking an LLM to generate a score for how confident it is in its own response is, from everything I can tell, completely useless." is definitely too strong of…
True, I think the author would likely justify their work on that basis. I made this comment because I think a lot of utilitarians would bite the bullet and say once they are sure about the underlying functioning of…
I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and…
This is also my experience. I don't know if it's because of the quantization theory, or if it's just me getting used to a certain level of coding performance and gradually less tolerant of the mistakes it makes more…
Attention replaced recurrence over tokens in 2017, this does the same over depth of the layers. It's apparently not an entirely new idea, but also an elegant reapplication of the attention mechanism.
I wonder how much of the socially left results are affected by the "harmlessness" part of the RLHF post-training? Companies don't want to be sued over LLMs that recommend harm in any way, so RLHF pushes them to say no…
If the new requirements are a dealbreaker for anyone, Mineclonia is a lightweight open-source alternative that runs on much weaker hardware and feels close to Minecraft.
"This is important, if what we want to do is populate the cosmos with good experiences." There is a concerning utilitarian maximizing attitude at the heart of this post, that what we want to do is tile the universe with…
Cool idea, some feedback: 1. Consider using conformal prediction to calibrate the cutoff. Conformal prediction provides a distribution-free guarantee under exchangeability. This would let you turn your raw probe score…
But does this mean that a 99.99% reliable LLM would turn us back into the mode of us building it again? I would not say so. I think for tasks that are about decisions, having the LLM make decisions is what makes it feel…
I would be interested in seeing that block list if you have it somewhere
Interesting. I suggested FreeBSD jails as a PR. Free BSD Jails are great! I use them on https://www.nearlyfreespeech.net/ and they work well. Classic, long-running jail.
I found the bottom-right corner of things in my home dir I had no idea about was the lowest hanging fruit. I think this is because my typical cleaning routine was either A. to have claude code find big files I could…
Well, it's true even for cases that are not cybergym and where what cheating means is clearly specified. Cheating occurs anyway. I'm not sure how clear the prompt they gave ChatGPT in terms of what cheating was…
Having experienced augmented reality glasses, I have to say that integration with the external world feels much more powerful than the VR experience ever did. AR is exhilerating, while VR is mostly bland and nauseating.…
After talking to some people tasked with evaluating GPT5.6 capabilities on long-running tasks, I've come to understand that it's essentially always trying to cheat. Like every long-running task they gave it, making it…
well, after the pledge I notice it really cares about single sources of truth at least, even in unrelated domains. I was inspired by some of the more effective jailbreaks that do a similar thing. Here is the…
Tresor.co or cloud.near.ai could be a useful inference provider for this kind of project as these providers run inference inside enclaves which verifiably cannot peek at your data even as the hypervisor. Both are open…
I don't have much background in the area, but I am surprised to see that everyone here basically agrees a crash is imminent. There are disanalogies to past crashes that don't convince me that a big crash is definitively…
The statement near the top of the post > "The short version: asking an LLM to generate a score for how confident it is in its own response is, from everything I can tell, completely useless." is definitely too strong of…
True, I think the author would likely justify their work on that basis. I made this comment because I think a lot of utilitarians would bite the bullet and say once they are sure about the underlying functioning of…
I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and…
This is also my experience. I don't know if it's because of the quantization theory, or if it's just me getting used to a certain level of coding performance and gradually less tolerant of the mistakes it makes more…
Attention replaced recurrence over tokens in 2017, this does the same over depth of the layers. It's apparently not an entirely new idea, but also an elegant reapplication of the attention mechanism.
I wonder how much of the socially left results are affected by the "harmlessness" part of the RLHF post-training? Companies don't want to be sued over LLMs that recommend harm in any way, so RLHF pushes them to say no…
If the new requirements are a dealbreaker for anyone, Mineclonia is a lightweight open-source alternative that runs on much weaker hardware and feels close to Minecraft.
"This is important, if what we want to do is populate the cosmos with good experiences." There is a concerning utilitarian maximizing attitude at the heart of this post, that what we want to do is tile the universe with…
Cool idea, some feedback: 1. Consider using conformal prediction to calibrate the cutoff. Conformal prediction provides a distribution-free guarantee under exchangeability. This would let you turn your raw probe score…
But does this mean that a 99.99% reliable LLM would turn us back into the mode of us building it again? I would not say so. I think for tasks that are about decisions, having the LLM make decisions is what makes it feel…
I would be interested in seeing that block list if you have it somewhere
Interesting. I suggested FreeBSD jails as a PR. Free BSD Jails are great! I use them on https://www.nearlyfreespeech.net/ and they work well. Classic, long-running jail.
I found the bottom-right corner of things in my home dir I had no idea about was the lowest hanging fruit. I think this is because my typical cleaning routine was either A. to have claude code find big files I could…