
Hugging Face CEO Clement Delangue says he's launching an Open Alignment Initiative, led by co-founder Thomas Wolf, and has publicly applied to join Anthropic's freshly proposed "embedded evaluator" program. If Anthropic approves, Hugging Face's evaluators could be stationed on-site long-term with near-employee-level access, inspecting model training and safety measures directly.
Delangue argues that AI alignment can't keep getting solved behind closed doors at a handful of frontier labs — more outside researchers need in. Anthropic has already promised to open that kind of access to third parties: desks, badge access, company laptops, and tools close to its internal risk team, along with the right for evaluators to publish their conclusions independently.
https://twitter.com/ClementDelangue/status/2098790988034580852