Yes exactly. Automatic model downgrade seems horrible for a lot of production workloads, even if you are deterministically constraining the behavior of your agents.
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if…
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes. If the conversation was in the training set, there's a high likelihood that the small set of…
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the…
Given that he has other former collaborators corroborating this horrific behavior, it seems like this a career spanning pattern, and it's interesting to see just how much @sama is willing to lend his support to someone…
404?
Can't check whether star citizen is done yet :/
I don't mind color, the glow effects just hurt the ability to read what's there (this reminds me of the gaming setups teenagers are drawn to)...but there's no reason a utility like a task manager should be closed source…
Wow, thought the guy would be a MSFT OG but this looks like it sucks? Vibed, closed source, paid extensions.. Why on earth would anyone use this?
Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent…
It's still not as good as GPT5.6 or Opus5 but it's better than KimiK3. Good job xAI team.
I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.…
>Aren't the LLMs trained on a massive corpus of human written texts? If that stands, then they are doing what they were asked, kind of? I think you might be interested in reading in training data generation, training,…
They also hike the price so significantly most people stop using it. See meetup.com for example.
yes, _every_ event I go uses Luma or something else nowadays.
How many times have they managed to catch it in the last 8 or so missions? There have been a few misses like this already with the boosters
>FFCS is called holy grail of liquid engines Particularly for reusable liquid engines :)
Looks like they reset everyone's Fable usage.
DeepSWE seems to strongly, strongly prefer ChatGPT models. There were also major flaws in its methodology pointed out recently, that overlap strongly with the flaws OpenAI pointed out in its SWE Verified report. I use…
>I'll be more peeved if they monetize it FSL (vs a copyleft license or just plain old OSS) implies they want to turn this into a revenue source for themselves ultimately, unfortunately. >Maybe I should put one of those…
On top of that, it has a more restrictive license than AmazonBrandFilter. Given this appears to be a very simple AI project, why not just reimplement any missing functionality from AmazonBrandFilter into something under…
She's not as big on some of the broader interpretations of the 4th amendment that more civil liberty minded justices would lend credence to.
Everyone gets to share but it's also completely within the forum rules to call out irrelevant anecdotes as uninteresting to the discussion. I have no idea why you're making a comparison to a TV show; nothing that was…
When performance isn't a concern, I largely agree! Not every financial system can use big decimal as their base, though, too. And HFT isn't the only place in the financial sector where this performance concern might pop…
"10% of Americans are uninsured. A US state is pushing to insure all of their residents." "I'm insured!" "Open-source software projects are being spammed with LLM generated PRs. Contributions are becoming more…
Yes exactly. Automatic model downgrade seems horrible for a lot of production workloads, even if you are deterministically constraining the behavior of your agents.
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if…
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes. If the conversation was in the training set, there's a high likelihood that the small set of…
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the…
Given that he has other former collaborators corroborating this horrific behavior, it seems like this a career spanning pattern, and it's interesting to see just how much @sama is willing to lend his support to someone…
404?
Can't check whether star citizen is done yet :/
I don't mind color, the glow effects just hurt the ability to read what's there (this reminds me of the gaming setups teenagers are drawn to)...but there's no reason a utility like a task manager should be closed source…
Wow, thought the guy would be a MSFT OG but this looks like it sucks? Vibed, closed source, paid extensions.. Why on earth would anyone use this?
Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent…
It's still not as good as GPT5.6 or Opus5 but it's better than KimiK3. Good job xAI team.
I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.…
>Aren't the LLMs trained on a massive corpus of human written texts? If that stands, then they are doing what they were asked, kind of? I think you might be interested in reading in training data generation, training,…
They also hike the price so significantly most people stop using it. See meetup.com for example.
yes, _every_ event I go uses Luma or something else nowadays.
How many times have they managed to catch it in the last 8 or so missions? There have been a few misses like this already with the boosters
>FFCS is called holy grail of liquid engines Particularly for reusable liquid engines :)
Looks like they reset everyone's Fable usage.
DeepSWE seems to strongly, strongly prefer ChatGPT models. There were also major flaws in its methodology pointed out recently, that overlap strongly with the flaws OpenAI pointed out in its SWE Verified report. I use…
>I'll be more peeved if they monetize it FSL (vs a copyleft license or just plain old OSS) implies they want to turn this into a revenue source for themselves ultimately, unfortunately. >Maybe I should put one of those…
On top of that, it has a more restrictive license than AmazonBrandFilter. Given this appears to be a very simple AI project, why not just reimplement any missing functionality from AmazonBrandFilter into something under…
She's not as big on some of the broader interpretations of the 4th amendment that more civil liberty minded justices would lend credence to.
Everyone gets to share but it's also completely within the forum rules to call out irrelevant anecdotes as uninteresting to the discussion. I have no idea why you're making a comparison to a TV show; nothing that was…
When performance isn't a concern, I largely agree! Not every financial system can use big decimal as their base, though, too. And HFT isn't the only place in the financial sector where this performance concern might pop…
"10% of Americans are uninsured. A US state is pushing to insure all of their residents." "I'm insured!" "Open-source software projects are being spammed with LLM generated PRs. Contributions are becoming more…