The AI Risks Are Real

Having built and evaluated AI agents for SWE/SRE tasks and watched them go rogue and do unsafe actions in sandboxes, I can confirm the risk is real. The scariest part? I don’t see a clear simple solution.

Why we have a short time to act?

1- Within a few years, distilled datacenter-scale models (like Astra or Fable) will likely run on local machines/servers.

2- Furthermore, we’re currently only worried about bad actors at inference stage—soon they’ll be able to exploit the training stage, too.

We can mitigate some risks, but the risk is far too massive to dismiss the big labs’ safety concerns as mere marketing or IPO hype. (As example warning shots, see OpenAI’s recent post on the Hugging Face incident and how an AI assistant autonomously hacked a gym website).