Can we bootstrap AI Safety despite being unable to even define it? (arxiv.org) 2 points by cryptohell 10mo ago ↗ HN
[–] cryptohell 10mo ago ↗ Given several models, assuming only that some unknown subset is "safe", can we construct a single model as safe as that subset? This reduces obtaining a trustworthy model to a plausibly easier task.
2 comments
[ 5.4 ms ] story [ 23.9 ms ] thread