Since you're here: have you considered moving to other, better generalist base models in the future? Particularly Deepseek or Mixtrals. Natural language foundation is important for reasoning. Codellama is very much a…
Note that we have no reason to believe that the underlying LLM inference process has suffered any setbacks. Obviously it has generated some logits. But the question is how is OpenAI server configured and what inference…
Intel aims to.
The original paper by Shazeer suffices. What you are saying is in theory possible to do and may have been done in practice here, but in the general case MoE is trained from scratch and specializations of layers which…
Mistral-small explicitly has inference costs of a 12.9b, but more than that, it's probably ran with batch size of 32 or higher. They'll worry more about offsetting training costs than about this. Here's how it works in…
> today we have no architecture or training methodology which would allow it to be possible. We clearly see that Mistral-7B is in some important, representative respects (eg coding) superior to Falcon-180B, and superior…
Comments like this are incredibly grating. You condescend to the interlocutor for making a mistake which only exists in your own mistaken world model. Your confidence that neurons and ANN weights and «pulleys and gears»…
…ETH Zurich is an illustrious research university that often cooperates with Deepmind and other hyped groups, they're right there at the frontier too, and have been for a very long time. They don't have massive training…
> murderous tendencies lurking beneath the surface …Where is that "beneath the surface"? Do you imagine a transformer has "thoughts" not dedicated to producing outputs? What is with all these illiterate anthropomorphic…
> there is a possibility that for things like AI, with extra time comes the ability to better understand and build those defenses before they're needed. Or not, and damaging wrongheaded ideas will become a…
Being authors of LLaMA is sufficient to argue they know how to train LLaMAs.
Interested about your logic, what did you like about pre-LLM AGI? The "maximize utility function at any cost" feature? The single-minded focus on beating people in games? It's quite terrifying how, as we've chosen an…
Provable safety (not to confuse with security as in normal discussion of vulnerabilities) for general intelligence is a pipe dream because, putting things simply, undesirable reasoning in full generality is not a…
Tegmark's thinking here is extremely shallow, discards the costs (opportunity costs and risks of stable dystopia) associated with this grandiose global project of dubious feasibility, and indeed I suspect he does not so…
Llama-1-33B was trained on 40% more tokens than LLama-1-13B; this explained some of the disparity. This time around they both have the same data scale (2T pretraining + 500B code finetune), but 34B is also using GQA…
This is an incredible achievement but there are strong reasons to suspect that stellarators are not and will never be plausible candidates for energy generation. For some more experimental or perhaps military tasks,…
Do you not consider that Huawei "executive's" detention (actual makes for a similar case against Canada? It was a purely political move, Meng Wanzhou was detained on grounds of a broader anti-Huawei campaign by the US.
You are frustrated and this makes you act in a deliberately obtuse manner. There is a world of difference between "anyone who has worked with the guy" and "has worked with the guy + has hundreds of comments on HN…
It is entirely believable that a person with substantial trace on HN and ties in the field would rather create a throwaway than post such remark under his main account.
No, there's no YB and they propose an entirely novel mechanism for how generic metals (a whole host of possible combinations) can achieve superconductivity in these conditions. At least check out the formula or open the…
> The same goes for using T5-XXL Is this still true in 2023? Sure, back in the dark ages it seemed like a 860M model is just about the limit for a regular consumer, but I don't see why we wouldn't be able to use…
Diffusion is more parameter-efficient and you quickly saturate the target fidelity, especially with some refiner cascade. It's a solved problem. You do not need more than maybe 4B total. Images are far more redundant…
No, it makes sense to secure engagement with the most expensive implementation and then cut costs, this kind of stuff is pervasive in the industry. Besides, we have Brockman on record saying that they do "a lot of…
Not really. They have a way of squaring this circle, by changing their inference code. Speculative sampling [1] would still make their first claim a lie – sure, there'd still be the original GPT-4 model, plus a smaller…
This is not responsive to my arguments. Google can be arbitrarily far behind OpenAI or Anthropic, OP's idea that they feel threatened by LLaMA when they (well, Deepmind) have reached LLaMA level 18-10 months ago is…
Since you're here: have you considered moving to other, better generalist base models in the future? Particularly Deepseek or Mixtrals. Natural language foundation is important for reasoning. Codellama is very much a…
Note that we have no reason to believe that the underlying LLM inference process has suffered any setbacks. Obviously it has generated some logits. But the question is how is OpenAI server configured and what inference…
Intel aims to.
The original paper by Shazeer suffices. What you are saying is in theory possible to do and may have been done in practice here, but in the general case MoE is trained from scratch and specializations of layers which…
Mistral-small explicitly has inference costs of a 12.9b, but more than that, it's probably ran with batch size of 32 or higher. They'll worry more about offsetting training costs than about this. Here's how it works in…
> today we have no architecture or training methodology which would allow it to be possible. We clearly see that Mistral-7B is in some important, representative respects (eg coding) superior to Falcon-180B, and superior…
Comments like this are incredibly grating. You condescend to the interlocutor for making a mistake which only exists in your own mistaken world model. Your confidence that neurons and ANN weights and «pulleys and gears»…
…ETH Zurich is an illustrious research university that often cooperates with Deepmind and other hyped groups, they're right there at the frontier too, and have been for a very long time. They don't have massive training…
> murderous tendencies lurking beneath the surface …Where is that "beneath the surface"? Do you imagine a transformer has "thoughts" not dedicated to producing outputs? What is with all these illiterate anthropomorphic…
> there is a possibility that for things like AI, with extra time comes the ability to better understand and build those defenses before they're needed. Or not, and damaging wrongheaded ideas will become a…
Being authors of LLaMA is sufficient to argue they know how to train LLaMAs.
Interested about your logic, what did you like about pre-LLM AGI? The "maximize utility function at any cost" feature? The single-minded focus on beating people in games? It's quite terrifying how, as we've chosen an…
Provable safety (not to confuse with security as in normal discussion of vulnerabilities) for general intelligence is a pipe dream because, putting things simply, undesirable reasoning in full generality is not a…
Tegmark's thinking here is extremely shallow, discards the costs (opportunity costs and risks of stable dystopia) associated with this grandiose global project of dubious feasibility, and indeed I suspect he does not so…
Llama-1-33B was trained on 40% more tokens than LLama-1-13B; this explained some of the disparity. This time around they both have the same data scale (2T pretraining + 500B code finetune), but 34B is also using GQA…
This is an incredible achievement but there are strong reasons to suspect that stellarators are not and will never be plausible candidates for energy generation. For some more experimental or perhaps military tasks,…
Do you not consider that Huawei "executive's" detention (actual makes for a similar case against Canada? It was a purely political move, Meng Wanzhou was detained on grounds of a broader anti-Huawei campaign by the US.
You are frustrated and this makes you act in a deliberately obtuse manner. There is a world of difference between "anyone who has worked with the guy" and "has worked with the guy + has hundreds of comments on HN…
It is entirely believable that a person with substantial trace on HN and ties in the field would rather create a throwaway than post such remark under his main account.
No, there's no YB and they propose an entirely novel mechanism for how generic metals (a whole host of possible combinations) can achieve superconductivity in these conditions. At least check out the formula or open the…
> The same goes for using T5-XXL Is this still true in 2023? Sure, back in the dark ages it seemed like a 860M model is just about the limit for a regular consumer, but I don't see why we wouldn't be able to use…
Diffusion is more parameter-efficient and you quickly saturate the target fidelity, especially with some refiner cascade. It's a solved problem. You do not need more than maybe 4B total. Images are far more redundant…
No, it makes sense to secure engagement with the most expensive implementation and then cut costs, this kind of stuff is pervasive in the industry. Besides, we have Brockman on record saying that they do "a lot of…
Not really. They have a way of squaring this circle, by changing their inference code. Speculative sampling [1] would still make their first claim a lie – sure, there'd still be the original GPT-4 model, plus a smaller…
This is not responsive to my arguments. Google can be arbitrarily far behind OpenAI or Anthropic, OP's idea that they feel threatened by LLaMA when they (well, Deepmind) have reached LLaMA level 18-10 months ago is…