Over 100 AI experts signed a letter begging Anthropic and OpenAI to let someone independent check whether their models might kill us all. The labs have been evaluating their own safety this whole time. Like letting a drunk driver administer his own breathalyzer.
The experts want transparency. They want independence. They want someone who doesn't collect a paycheck from Sam Altman to verify that the thing trained on the entire internet won't decide humans are inefficient. Anthropic and OpenAI have responded by continuing to do exactly what they were doing before, which is building the models faster and pinky-swearing everything's fine.
Foundation model labs currently operate under a self-certification system that would make Boeing's 737 MAX safety review process look rigorous. Build the model. Test the model with your own people. Announce the model is safe. Ship the model. Collect venture capital. Repeat until something goes wrong or everyone gets rich. Guess which one happens first.
The letter doesn't legally require the labs to do anything. It's a polite suggestion from people who understand how transformers work that maybe someone other than the guy trying to hit his Q3 revenue target should verify the kill switch actually works. But suggestions are just suggestions, and OpenAI has a product roadmap to maintain.
The real innovation here is convincing 100 experts to waste time writing a letter that will be read by people who already decided the answer is no. That's harder than training GPT-5.
Photo by on Unsplash

Leave a Comment