Sigh... I'm willing to bet (and this is not a hard thing to bet on) that even bb(50) is independent of ZFC and similar, and Omega(1KB) is > 748 , which is at this point known to be independent of ZFC.... Please do not…
Citadel already is. They are happy to do it - bedrock is the host of choice. They are less happy about cost - but that is a question of value :)
One - and I do not mean to be snarky - you can literally ask Gpt 5.6 Sol this - and if you want to see cool stuff - Fable running in their app (not website) has a view thinking button that is actually a good way to…
Glm 5.2 nvfp4 on 4 b300 with dpattn 4 and ram will get you about 20 users live at 400k context - and 60 easily if you give that server 2tb of ram and 4 nvme 8 tb drives. There are some sglang patches needed - but we…
Or more colloquially : paperclip maximization . From OpenAI - you know, the guys who _really_ know this... Sigh... Did they finish the prompt with "And do whatever you can to get this done!" ? Cause that's the only…
This is an _amazing_ typo :) Thank you, thank you :)
3T at nxfp4 (which is most of it) is only 1.5TB of vram - so 8x288GB B300 or MI355 will do it if you are careful with context - maybe dp-attn? Certainly not TP. 2 of those together can easily serve it. The new AMD MI400…
Don't forget that you are not really seeing the thinking tokens used - so non-trivial to count them.
Yeah, if you have a fixed llm topology, you can just effectively burns 2 top layers of the chip as Rom (model weights) - which has a per area density even better than dram - so it’s just attention and kv streaming that…
Well, for a lot of agentic stuff nowadays, having 250k-500K context is where things live - and the benchmarks don't really show that unfortunately - but they could :)
Agreed - there was always a set of things I wanted to do that I knew the magic core for, but wanted a team of implementers for the curft, the 100k of actual testing harnesses, hyperparameter exploration, etc.. . I now…
We are there already pretty much - if I understand your point (“How the models are wielded”) refers to the harness - which is part of model training already. Fable was trained to use Claude code harnesses effectively to…
[dead]
Well - there is a giant push to allow non-qualified investors to invest their 401k (and roth and whatever) into the private equities market - pre-IPO companies and such. I can't shake the feeling of a grand fleecing…
A lot - and over the coming 2 years, even more. Utilization rates are under 50% across the board, and special and cheaper chips are coming out all the time for inference. And a truckload of research - TurboQuant, HC…
Imagine an agent shadowing all your terminals, providing ideas and asking to run commands that will let it verify the hypotheses it comes up with, while at the same time doing research on vendor docs, etc... Quite safe,…
Minor nit re[2]: for agentic workloads that are actually worth money - i.e., claude code and similar, things are either prefill-bound - which this does not help - or more importantly tps/user bound (at 150k+ context…
Kindof yeah - predictivity is a question though for larger layers - when trying to scale this up. But yeah, this is a "95% predictor in latent space is a 7x improvement in speed if done right" approach.
Yeah, forgot about them - 100%.
I kindof agree that it is unattractive - but the regulators are perfectly happy with "EOD also introduces credit risk on the clearing house/bilateral." if it allows them to protect retail and institutional investors.…
Note that _passenger aviation_ is commercially non-competitive. The big 4 US airlines make money on credit cards, not airfare : they lose money on airfare. So, most people who are trying to make money will not use them…
"who do carry liability when things go wrong" -> unless one pierces the corporate veil, it's just money. Not even their money. HIPAA - unless basically stealing data - will not generate personal liability. And even for…
We are multiple orders of magnitude away from Landauer limits - so next big thing in matmul could be photonic multipliers - there’s a bunch of them coming up in the next 3? years. So that’s a 2-4 order of magnitude…
I think the one thing you are not taking into account is that the investors on average fundamentally don’t care. Scale arbitrage means that small companies are fundamentally about velocity - and if they get sued due to…
But it does allow these investors to participate in the markets without losing their shirts - and the lack of such liquidity would impact the market more so than the cost of the risk mitigation - which as you completely…
Sigh... I'm willing to bet (and this is not a hard thing to bet on) that even bb(50) is independent of ZFC and similar, and Omega(1KB) is > 748 , which is at this point known to be independent of ZFC.... Please do not…
Citadel already is. They are happy to do it - bedrock is the host of choice. They are less happy about cost - but that is a question of value :)
One - and I do not mean to be snarky - you can literally ask Gpt 5.6 Sol this - and if you want to see cool stuff - Fable running in their app (not website) has a view thinking button that is actually a good way to…
Glm 5.2 nvfp4 on 4 b300 with dpattn 4 and ram will get you about 20 users live at 400k context - and 60 easily if you give that server 2tb of ram and 4 nvme 8 tb drives. There are some sglang patches needed - but we…
Or more colloquially : paperclip maximization . From OpenAI - you know, the guys who _really_ know this... Sigh... Did they finish the prompt with "And do whatever you can to get this done!" ? Cause that's the only…
This is an _amazing_ typo :) Thank you, thank you :)
3T at nxfp4 (which is most of it) is only 1.5TB of vram - so 8x288GB B300 or MI355 will do it if you are careful with context - maybe dp-attn? Certainly not TP. 2 of those together can easily serve it. The new AMD MI400…
Don't forget that you are not really seeing the thinking tokens used - so non-trivial to count them.
Yeah, if you have a fixed llm topology, you can just effectively burns 2 top layers of the chip as Rom (model weights) - which has a per area density even better than dram - so it’s just attention and kv streaming that…
Well, for a lot of agentic stuff nowadays, having 250k-500K context is where things live - and the benchmarks don't really show that unfortunately - but they could :)
Agreed - there was always a set of things I wanted to do that I knew the magic core for, but wanted a team of implementers for the curft, the 100k of actual testing harnesses, hyperparameter exploration, etc.. . I now…
We are there already pretty much - if I understand your point (“How the models are wielded”) refers to the harness - which is part of model training already. Fable was trained to use Claude code harnesses effectively to…
[dead]
Well - there is a giant push to allow non-qualified investors to invest their 401k (and roth and whatever) into the private equities market - pre-IPO companies and such. I can't shake the feeling of a grand fleecing…
A lot - and over the coming 2 years, even more. Utilization rates are under 50% across the board, and special and cheaper chips are coming out all the time for inference. And a truckload of research - TurboQuant, HC…
Imagine an agent shadowing all your terminals, providing ideas and asking to run commands that will let it verify the hypotheses it comes up with, while at the same time doing research on vendor docs, etc... Quite safe,…
Minor nit re[2]: for agentic workloads that are actually worth money - i.e., claude code and similar, things are either prefill-bound - which this does not help - or more importantly tps/user bound (at 150k+ context…
Kindof yeah - predictivity is a question though for larger layers - when trying to scale this up. But yeah, this is a "95% predictor in latent space is a 7x improvement in speed if done right" approach.
Yeah, forgot about them - 100%.
I kindof agree that it is unattractive - but the regulators are perfectly happy with "EOD also introduces credit risk on the clearing house/bilateral." if it allows them to protect retail and institutional investors.…
Note that _passenger aviation_ is commercially non-competitive. The big 4 US airlines make money on credit cards, not airfare : they lose money on airfare. So, most people who are trying to make money will not use them…
"who do carry liability when things go wrong" -> unless one pierces the corporate veil, it's just money. Not even their money. HIPAA - unless basically stealing data - will not generate personal liability. And even for…
We are multiple orders of magnitude away from Landauer limits - so next big thing in matmul could be photonic multipliers - there’s a bunch of them coming up in the next 3? years. So that’s a 2-4 order of magnitude…
I think the one thing you are not taking into account is that the investors on average fundamentally don’t care. Scale arbitrage means that small companies are fundamentally about velocity - and if they get sued due to…
But it does allow these investors to participate in the markets without losing their shirts - and the lack of such liquidity would impact the market more so than the cost of the risk mitigation - which as you completely…