A small group of third-party AI safety evaluators just found out they're important. This happened because someone decided to ask them questions. The questions were about whether AI might kill everyone. The evaluators said maybe. Now they're at the center of a multitrillion-dollar industry, which is the exact place you want people who until last week were checking if chatbots could make bombs.
These are the gatekeepers. They evaluate models before release. They run tests. They write reports. Nobody read the reports. Now everyone reads the reports because the AI safety debate intensified, which is what happens when you let tech CEOs testify to Congress. The evaluators stepped into the spotlight. They didn't ask for this. They were fine in the shadows running hypothetical scenarios about rogue AGI. Now they're supposed to stop Sam Altman from building God.
The industry is worth trillions. The evaluators are worth dozens of people. Maybe a hundred if you count the interns. They're supposed to evaluate models that cost hundreds of millions to train. They get a few weeks and a API key. They write up their findings. The findings say concerning or requires further monitoring or within acceptable parameters. Then the model ships anyway because the product roadmap doesn't care about your p-values.
Retail traders are now pricing in AI safety evaluator sentiment. They're buying stocks based on whether some guy with a PhD said the new model passed the vibe check. They're selling because an evaluator flagged potential misuse risks in a 47-page PDF nobody will read. They're convinced this matters. They're wrong. The evaluators have exactly as much power as the companies give them, which is the amount of power you give the health inspector when you really need to open the restaurant tomorrow.
The gatekeepers are in the spotlight. The gate is already open. The horses left. We're three towns over.
Photo by Jeremy Vejgman on Unsplash

Leave a Comment