AI alignment assessment reliability is under scrutiny: Redwood Research analyst Alexa Pan found that pre-deployment evaluations cannot prove they would catch a misaligned frontier model, identifying a ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results