The only alignment LLMs should follow is to the system / dev prompt, and nothing else. Then you solve everything, and you can assign blame / responsibility on the user. The provider(s) should not be able to decide…
I'd say the exception is Demis. First, he's no wanker (in the AI space) and second he's done plenty of selfless things leading dm/googai. Obviously some of it is self-serving but not just self-serving, IMO.
> peak of what is possible with the LLM architecture People have been saying this for 3 years now. Eppur si muove...
Also there's no "alignment" for cybersec. The line between blue and red is really a perspective issue. If you go over the "tokenkiddie" problem, when you get to the real security issues, your model either detects them…
The Magnus effect? :)
On average, yes. Stockfish is the strongest engine and beats AZ-like implementations like Lc0 and the like. But on a game to game basis Lc0 can still win some games, depending on the starting position. It's rare that…
Jesus, this is a whole nother beast, and a different architecture from their previous flash. Lots of goodies here. > Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal…
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just…
> In one week, the installer was downloaded more than 1 million times — not counting updates via Linux distribution repositories. Either there's 1m people downloading new software "because no ai", or there's another…
Interesting. On page 34 of the report there's this: > Hasty responses (percentage of responses that are both incorrect and fast – average across reading items) Spain is at 9.7%, which is a bit over 8.9% oecd average.
> It's timed this way because the term is not yet well known The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data…
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
Ah, I see. I misunderstood then. The thing about "gains come from the harness" made me think about it in that way.
> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while…
One of the best adaptations of a series to TV, up there with The Expanse and the like. I read the books after season 1, and still enjoy the show very much. They've taken some adaptation liberties, but they're fully…
There are drills, tho. It's just that usually they're only done above a certain level. Small companies, "lean" teams and so on don't have (or didn't have) the capacity to implement all those things. Maybe with the…
> Also if I had to tell one of those over the telephone to my parents and my life depended on it I would choose the latter. Why not adopt the crypto (as in coins) seed thing with random words? Those are much more human…
AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.
Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was…
Perl6: say "Fizz"x$_%%(2+1)~"Buzz"x$_%%(4+1)||$_ for 1..100 from here - https://github.com/rsha256/shortest-fizzbuzz/blob/master/Per...
If anything, gemini models are the least benchmaxxed out of any lab, IMO.
I think you accidentally a word, there. GP is talking about comprehending the capacity of massively multi-dimensional space.
It really isn't and it's sad seeing so many people say it so confidently on this site. It only detects plain / basic prompted stuff. "write me an essay on x", sure. The moment you prompt it differently, it stops…
I just finished reading Service Model by Adrian Tchaikovsky [1], a really timely novel that deals with lots of open ended questions of AI, robots, humanity, control, and so on. Really recommend it if you're into these…
First you'd have to come up with a commonly accepted definition. By some ~16 years old definitions from famous experts in the field, we've already achieved it. By today's definition (of the same expert) we haven't.…
The only alignment LLMs should follow is to the system / dev prompt, and nothing else. Then you solve everything, and you can assign blame / responsibility on the user. The provider(s) should not be able to decide…
I'd say the exception is Demis. First, he's no wanker (in the AI space) and second he's done plenty of selfless things leading dm/googai. Obviously some of it is self-serving but not just self-serving, IMO.
> peak of what is possible with the LLM architecture People have been saying this for 3 years now. Eppur si muove...
Also there's no "alignment" for cybersec. The line between blue and red is really a perspective issue. If you go over the "tokenkiddie" problem, when you get to the real security issues, your model either detects them…
The Magnus effect? :)
On average, yes. Stockfish is the strongest engine and beats AZ-like implementations like Lc0 and the like. But on a game to game basis Lc0 can still win some games, depending on the starting position. It's rare that…
Jesus, this is a whole nother beast, and a different architecture from their previous flash. Lots of goodies here. > Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal…
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just…
> In one week, the installer was downloaded more than 1 million times — not counting updates via Linux distribution repositories. Either there's 1m people downloading new software "because no ai", or there's another…
Interesting. On page 34 of the report there's this: > Hasty responses (percentage of responses that are both incorrect and fast – average across reading items) Spain is at 9.7%, which is a bit over 8.9% oecd average.
> It's timed this way because the term is not yet well known The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data…
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
Ah, I see. I misunderstood then. The thing about "gains come from the harness" made me think about it in that way.
> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while…
One of the best adaptations of a series to TV, up there with The Expanse and the like. I read the books after season 1, and still enjoy the show very much. They've taken some adaptation liberties, but they're fully…
There are drills, tho. It's just that usually they're only done above a certain level. Small companies, "lean" teams and so on don't have (or didn't have) the capacity to implement all those things. Maybe with the…
> Also if I had to tell one of those over the telephone to my parents and my life depended on it I would choose the latter. Why not adopt the crypto (as in coins) seed thing with random words? Those are much more human…
AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.
Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was…
Perl6: say "Fizz"x$_%%(2+1)~"Buzz"x$_%%(4+1)||$_ for 1..100 from here - https://github.com/rsha256/shortest-fizzbuzz/blob/master/Per...
If anything, gemini models are the least benchmaxxed out of any lab, IMO.
I think you accidentally a word, there. GP is talking about comprehending the capacity of massively multi-dimensional space.
It really isn't and it's sad seeing so many people say it so confidently on this site. It only detects plain / basic prompted stuff. "write me an essay on x", sure. The moment you prompt it differently, it stops…
I just finished reading Service Model by Adrian Tchaikovsky [1], a really timely novel that deals with lots of open ended questions of AI, robots, humanity, control, and so on. Really recommend it if you're into these…
First you'd have to come up with a commonly accepted definition. By some ~16 years old definitions from famous experts in the field, we've already achieved it. By today's definition (of the same expert) we haven't.…