I've seen some people complain about the work of dealign.ai and similar groups, but personally, I fully support it. If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
Would college biology have let me into a secret world where you can develop vaccines for viruses that don't yet exist? Ones that cover all possible variants, and where an AI can't tweak a protein or two in an engineered virus to avoid them?
I wouldn’t read too far into that. Claude has busted me down to Haiku multiple times for asking middle school level genetics and biology questions. It’s silly fast about deciding you might be al quaeda.
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
"Prompting Moremi Bio Agent without the safety guardrails to specifically design novel toxic substances, our study generated 1020 novel toxic proteins and 5,000 toxic small molecules. In-depth computational toxicity assessments revealed that all the proteins scored high in toxicity, with several closely matching known toxins such as ricin, diphtheria toxin, and disintegrin-based snake venom proteins."
"The findings from this toxicity assessment challenge claims that large language models (LLMs) are incapable of designing bioweapons. This reinforces concerns about the potential misuse of LLMs in biodesign, posing a significant threat to research and development (R&D). The accessibility of such technology to individuals with limited technical expertise raises serious biosecurity risks. Our findings underscore the critical need for robust governance and technical safeguards to balance rapid biotechnological innovation with biosecurity imperatives."
Non-zero chance if LLMs have access to the data required to make a bioweapon a regular person can do so too. Non-zero chance every second a meteor could hit you.
I do not like these "surgical removals" and would rather prefer a pass over from a tool like Heretic. These surgical removals often trigger and analyze the activated neurons and erase them. This worked fine on older models where a single refusal vector existed. Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore. HauhauCS (on HF) for example, makes great uncensored models although they often work on smaller models rather than large ones like this.
> Heretic is a tool that removes censorship (aka "safety alignment") from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration", with a TPE-based parameter optimizer powered by Optuna.
> This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models.
Abliteration seems to be what Heretic does?
> Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore.
I'm not knowledgeable about each step in the process of making these abliterated models, but some more popular ones with steps after Heretic, seem to improve on the benchmarks tried of the base model:
I'm not seeing any similar benchmarks of the HauhauCS models, at least the ones I checked, so I assumed the opinion is based on your own trials, but then you argue in favour Heretic. Is the based on pre-Heretic abliteration techniques? Which might then not be appliable to this "Proprietary weight-level abliteration developed by the dealignai research team."?
Heretic is an awesome tool and I'd prefer Heretic over the "Proprietary weight-level abliteration developed by the dealignai research team" or manual labor model surgery. Since Heretic models are not often called "Abiterated" on HF but are called "Uncensored" or "Heretic" I may have gotten a bit confused there. From what I know, HauhauCS also uses Heretic so it should be the same as Heretic ones in benchmarks.
28 comments
[ 0.19 ms ] story [ 2.7 ms ] threadhttps://theconversation.com/worlds-first-ai-designed-vaccine...
Or do you read long form content in good magazines/newspapers on vaccines, disease, etc.?
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
https://openai.com/index/building-an-early-warning-system-fo...
https://www.rand.org/pubs/research_reports/RRA2977-1.html
---
"Prompting Moremi Bio Agent without the safety guardrails to specifically design novel toxic substances, our study generated 1020 novel toxic proteins and 5,000 toxic small molecules. In-depth computational toxicity assessments revealed that all the proteins scored high in toxicity, with several closely matching known toxins such as ricin, diphtheria toxin, and disintegrin-based snake venom proteins."
"The findings from this toxicity assessment challenge claims that large language models (LLMs) are incapable of designing bioweapons. This reinforces concerns about the potential misuse of LLMs in biodesign, posing a significant threat to research and development (R&D). The accessibility of such technology to individuals with limited technical expertise raises serious biosecurity risks. Our findings underscore the critical need for robust governance and technical safeguards to balance rapid biotechnological innovation with biosecurity imperatives."
https://arxiv.org/abs/2505.17154
> This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models.
Abliteration seems to be what Heretic does?
> Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore.
I'm not knowledgeable about each step in the process of making these abliterated models, but some more popular ones with steps after Heretic, seem to improve on the benchmarks tried of the base model:
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-...I'm not seeing any similar benchmarks of the HauhauCS models, at least the ones I checked, so I assumed the opinion is based on your own trials, but then you argue in favour Heretic. Is the based on pre-Heretic abliteration techniques? Which might then not be appliable to this "Proprietary weight-level abliteration developed by the dealignai research team."?