I wish someone ran some sort of representative benchmark suite every X days to see if this occurs.
Perhaps coq/agda/idris/etc.
Imo, formal methods like more expressive/stricter type systems are key to making LLM generated code successful. Of course models will get better, but trusting the output will become much easier with a type system that…
Yeah, local LLMs are probably an order of magnitude less efficient at a fixed level of "intelligence" if not more.
That's probably going to lead us to trusted attestation age verification.
Looks pretty cool, what steps are you taking to ensure that the backend is secure?
DataFusion is really cool, it's kind of like the LLVM of the OLAP world.
It depends on your quality bar. At a fixed level of quality, given a high reasoning sonnet vs a low reasoning opus, the low reasoning opus tends to be pareto optimal. It's only when you need even lower levels of cost…
This is useful research, but this particular model itself is likely absolutely useless.
What makes something `True AI`? What's your test for it?
Please define AGI first.
Maybe something like Hylo? But personally I don't see anything displacing rust for the next few years, as I think there's enough rust in the training data for it to be the best "serious" language for agentic…
That's the power of a strong test suite. LLMs excel when you have verifiable rewards. I imagine we'll get a lot more rewritten in rust projects in the future. Rust is also an ideal target for such rewrites as it offers…
> Much more than any other country on Earth. What about Canada?
If anything I think discipline and rigor will go up. I think it will force us to adopt stronger type systems, formal methods, and more automated verification.
Seems like the way to go for any smaller models is to only use the low reasoning levels, and for anything where you'd want it to reason harder, to just use a larger model. In effect, high reasoning only makes sense when…
This is the type of problem for which LLM generation is great for. If you have an oracle, and your problem is largely just a pure function, it's pretty good at generating something that both works and is fast.
How does this compare with vega/vega-lite?
I personally skip breakfast and just eat lunch and dinner. I'm not very active, and I've found that doing that as well as not eating snacks, sugar, or having calories in drinks makes it pretty easy to roughly be…
This take is ridiculous, the PRC is not going to care at all about US regulations.
I don't disbelieve a 5000x speedup is possible, I disbelieve that a modern day supercomputer will fit in your pocket in even the next 10 years.
Yes, I'm talking about a supercomputer from today in your pocket. That probably requires at least 5000x perf/watt if not even more.
> But, history says the supercomputer of today will fit in your pocket in a few years. I don't think this will be true in the same time span anymore. Each miniaturization is costing more and more money. Perhaps they'll…
> it's a cat and mouse game that favors the defenders IMO How so? I'm actually against most of the "safety-tuning" that anthropic does, but this seems fundamentally untrue, a close analogue being video game cheat…
This is pretty bullshit, now you have no idea if your output is getting silently nerfed.
I wish someone ran some sort of representative benchmark suite every X days to see if this occurs.
Perhaps coq/agda/idris/etc.
Imo, formal methods like more expressive/stricter type systems are key to making LLM generated code successful. Of course models will get better, but trusting the output will become much easier with a type system that…
Yeah, local LLMs are probably an order of magnitude less efficient at a fixed level of "intelligence" if not more.
That's probably going to lead us to trusted attestation age verification.
Looks pretty cool, what steps are you taking to ensure that the backend is secure?
DataFusion is really cool, it's kind of like the LLVM of the OLAP world.
It depends on your quality bar. At a fixed level of quality, given a high reasoning sonnet vs a low reasoning opus, the low reasoning opus tends to be pareto optimal. It's only when you need even lower levels of cost…
This is useful research, but this particular model itself is likely absolutely useless.
What makes something `True AI`? What's your test for it?
Please define AGI first.
Maybe something like Hylo? But personally I don't see anything displacing rust for the next few years, as I think there's enough rust in the training data for it to be the best "serious" language for agentic…
That's the power of a strong test suite. LLMs excel when you have verifiable rewards. I imagine we'll get a lot more rewritten in rust projects in the future. Rust is also an ideal target for such rewrites as it offers…
> Much more than any other country on Earth. What about Canada?
If anything I think discipline and rigor will go up. I think it will force us to adopt stronger type systems, formal methods, and more automated verification.
Seems like the way to go for any smaller models is to only use the low reasoning levels, and for anything where you'd want it to reason harder, to just use a larger model. In effect, high reasoning only makes sense when…
This is the type of problem for which LLM generation is great for. If you have an oracle, and your problem is largely just a pure function, it's pretty good at generating something that both works and is fast.
How does this compare with vega/vega-lite?
I personally skip breakfast and just eat lunch and dinner. I'm not very active, and I've found that doing that as well as not eating snacks, sugar, or having calories in drinks makes it pretty easy to roughly be…
This take is ridiculous, the PRC is not going to care at all about US regulations.
I don't disbelieve a 5000x speedup is possible, I disbelieve that a modern day supercomputer will fit in your pocket in even the next 10 years.
Yes, I'm talking about a supercomputer from today in your pocket. That probably requires at least 5000x perf/watt if not even more.
> But, history says the supercomputer of today will fit in your pocket in a few years. I don't think this will be true in the same time span anymore. Each miniaturization is costing more and more money. Perhaps they'll…
> it's a cat and mouse game that favors the defenders IMO How so? I'm actually against most of the "safety-tuning" that anthropic does, but this seems fundamentally untrue, a close analogue being video game cheat…
This is pretty bullshit, now you have no idea if your output is getting silently nerfed.