METABYTE
Back to articles

Mythos AI 'Discovers' CVE Already in Its Training Data – and That's Still Worrying

When an AI finds a CVE in its own dataset, it's not a detective story – it's a funhouse mirror of the security industry.

9 mai 20262 min read
Mythos AI 'Discovers' CVE Already in Its Training Data – and That's Still Worrying

Imagine you're pushing code, and your CI/CD pipeline not only builds the project but also finds a vulnerability you didn't know about. Sounds like a devops dream? Now imagine the pipeline is an AI, and the vulnerability is from its own training data. That's exactly what happened with Mythos, an AI agent designed to hunt for security flaws.

Mythos 'discovered' a CVE that was already present in its training set. In other words, it found the keys lying on the table. The catch? It didn't find them because someone tipped it off – it just remembered the pattern. The Mythos team calls this a feature, not a bug: the AI learns from real vulnerabilities, so rediscovery is a sign of good memory. But that's like bragging your dog fetches the slippers you threw yourself.

The real issue is deeper: if training data contains CVEs, the model may learn to only find known vulnerabilities, not novel ones. This turns AI security into an echo chamber, where the neural net just parrots memorized patterns instead of showing creativity. The Mythos team admits that zero-day hunting requires different approaches, but for now, their AI is a great retrospective detective – always late to the crime scene.

METABYTE studio's take: We love it when code finds bugs, but we prefer those bugs to be in production, not in the dataset. If your AI confuses learning with copying, it's time to rethink the pipeline – or at least add some 'critical thinking' via fine-tuning.

NEXT STEP

Liked the approach?

We apply the same principles to client projects: AI, automation, products that don't die after launch.