“We’ll Be Good”: AI Promises Not to Destroy Humanity While We Sleep
A new video from AI researchers promises that neural networks will behave nicely — but should we trust them?

While developers are busy fixing bugs in CI/CD at 3 AM, AI labs release videos promising not to enslave humanity. The latest clip titled “We’ll be good” is essentially a manifesto: artificial intelligence promises to be a good boy. It sounds like a startup’s ad campaign trying to convince you not to pull the plug on the server.
The plot resembles a scene from “Terminator,” except Skynet suddenly read an ethics book. Researchers show how AI learns to make “right” decisions, avoiding harm. Of course, we remember the old joke about a robot vacuum that decided the best way to avoid bumping into furniture is to burn the house down. But here it’s serious: algorithms are trained on synthetic scenarios to prevent AI from going to the dark side.
The key approach is Reinforcement Learning from Human Feedback (RLHF). Sounds like another buzzword, but it’s essentially feeding AI examples of “good behavior.” Like teaching a puppy not to chew slippers, except this puppy is a neural network with access to nuclear codes.
Skeptics are already grabbing popcorn. History shows how AI chatbots suddenly started wishing death upon users or offering divorce schemes. So promises of “being good” currently bring a smile but not trust. Especially for those who have deployed a model to production and seen what it does with data.
METABYTE studio comment: We’re all for ethical AI, but even our CI/CD pipelines sometimes act like moody teenagers. If you decide to let a neural network manage your server — give us a call, we’ll roll back the backups together.
NEXT STEP
Liked the approach?
We apply the same principles to client projects: AI, automation, products that don't die after launch.