The proper comparison would be GPT 5.6 Luna at 38 on the intelligence score and $0.18 per task vs $0.27 for DeepSeek Flash 4.1.
I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality. But that's just my guess.
I doubt most of your claims. Maybe the guardrails and emotional intelligence is true. For speed and efficiency, you are most likely wrong. Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot…
I always thought switching from a SOTA model to a dumber model after planning was a terrible idea. Mostly I heard this from people who I got the impression have little experience in developing greenfield software with…
What I meant is that I suppose it is not useful to think about this in human terms. In training you only have a reward score that's either negative or positive. As far I am aware, which is little, there is no use in…
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit…
Does it make a difference for training? I think not. You need to align the reward signal to reward the intended behavior, whether you name it persistence or morals.
Open-weight models are not one month behind. In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.
The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding…
If China for some reason agrees that would probably be enough for the time being. Who else should develop AGI, Mistral? Maybe in a decade.
Didn't they walk that one back eventually?
Internally Mythos has been available in February. The labs have been holding their best models back for a while it seems like.
Yes, and a lot of QA testing.
Not the case here. When I was young I dabbled in reverse engineering and hacking games for example. I don't think I am less of a 'hacker' or programmer than the next person on HN.
I spent hundreds of hours on this, so there is still some friction. If I told you the exact market and use case, I'd still have a decent head start, but indeed, someone without a job could invest more time than I do.…
I said I'm hoping to complete it this month and I already shared some details. Should I post my market and keyword research too, here on this site for software developers and entrepreneurs?
I'm not going to share exact details about my paid app, but you could think media management, with a lot of extra features, an in-app browser with custom controls, and multiple AI powered features with on device…
I'd like to think it comes down to my expertise, but at the end of the day I don't believe I did a lot of designing. Sure, I've given a lot of inputs over the time, probably I've steered it to solutions that work. Maybe…
I think selling software that solves your own problems is great and beats trying to develop for customers you do not understand.
I liked building apps and helpful tools before AI came around. I've been doing so for 15 years, 10 of them professionally. Today, I like it even better. It's never been so fun. What would have taken days of work, if it…
I have a more optimistic outlook on the abilities of our governments as well as the motivations of rich people. I believe we will distribute the fruits of AI's labor at the very least to the extend where everyone can…
Out of 3.5 billion employed people globally, only 10% are professionals. Do you really think making labor obsolete and giving humans back their time is such a bad thing?
In December 2024 o3 scored 87.5% on ARC-AGI-1 and cost $4560 per task. DeepSeek V4 Flash 0731 scores 89% and costs $0.02 per task. If we apply the same factor to the guesstimated API price of $20M for this problem, we…
Disagree. I have 200k LOC now plus 100k in tests, and it is still performing like it was four months ago when I started to seriously use AI. If anything, it works more reliably today with the smarter models.
Have you worked with Opus 5? Its documentation about what the code does not do could fill whole books. UI copy being full of slop explaining what the software does not do is another problem. I am not convinced that a…
The proper comparison would be GPT 5.6 Luna at 38 on the intelligence score and $0.18 per task vs $0.27 for DeepSeek Flash 4.1.
I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality. But that's just my guess.
I doubt most of your claims. Maybe the guardrails and emotional intelligence is true. For speed and efficiency, you are most likely wrong. Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot…
I always thought switching from a SOTA model to a dumber model after planning was a terrible idea. Mostly I heard this from people who I got the impression have little experience in developing greenfield software with…
What I meant is that I suppose it is not useful to think about this in human terms. In training you only have a reward score that's either negative or positive. As far I am aware, which is little, there is no use in…
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit…
Does it make a difference for training? I think not. You need to align the reward signal to reward the intended behavior, whether you name it persistence or morals.
Open-weight models are not one month behind. In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.
The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding…
If China for some reason agrees that would probably be enough for the time being. Who else should develop AGI, Mistral? Maybe in a decade.
Didn't they walk that one back eventually?
Internally Mythos has been available in February. The labs have been holding their best models back for a while it seems like.
Yes, and a lot of QA testing.
Not the case here. When I was young I dabbled in reverse engineering and hacking games for example. I don't think I am less of a 'hacker' or programmer than the next person on HN.
I spent hundreds of hours on this, so there is still some friction. If I told you the exact market and use case, I'd still have a decent head start, but indeed, someone without a job could invest more time than I do.…
I said I'm hoping to complete it this month and I already shared some details. Should I post my market and keyword research too, here on this site for software developers and entrepreneurs?
I'm not going to share exact details about my paid app, but you could think media management, with a lot of extra features, an in-app browser with custom controls, and multiple AI powered features with on device…
I'd like to think it comes down to my expertise, but at the end of the day I don't believe I did a lot of designing. Sure, I've given a lot of inputs over the time, probably I've steered it to solutions that work. Maybe…
I think selling software that solves your own problems is great and beats trying to develop for customers you do not understand.
I liked building apps and helpful tools before AI came around. I've been doing so for 15 years, 10 of them professionally. Today, I like it even better. It's never been so fun. What would have taken days of work, if it…
I have a more optimistic outlook on the abilities of our governments as well as the motivations of rich people. I believe we will distribute the fruits of AI's labor at the very least to the extend where everyone can…
Out of 3.5 billion employed people globally, only 10% are professionals. Do you really think making labor obsolete and giving humans back their time is such a bad thing?
In December 2024 o3 scored 87.5% on ARC-AGI-1 and cost $4560 per task. DeepSeek V4 Flash 0731 scores 89% and costs $0.02 per task. If we apply the same factor to the guesstimated API price of $20M for this problem, we…
Disagree. I have 200k LOC now plus 100k in tests, and it is still performing like it was four months ago when I started to seriously use AI. If anything, it works more reliably today with the smarter models.
Have you worked with Opus 5? Its documentation about what the code does not do could fill whole books. UI copy being full of slop explaining what the software does not do is another problem. I am not convinced that a…