Listen to this article
Ever heard of the theory that supports how AI may be eating itself? The internet was supposed to be AI’s own personal buffet. There are billions of pages that exist with information. The internet is, simply put, an enormous archive of humans who think, write, and even communicate.
Then AI arrived and it started to produce content at a speed humans could never match. So? Where is the problem, you ask? Ask yourself, what happens when the machines start eating what they cooked? This is where the whole idea of cannibalism originated.
AI Eats, Creates, Then Eats Again
The theory has no chemistry. It is simple. AI trains on human-generated content. It simply produces new content. That certain content gets published online. After that, future AI models scrape the internet and encounter later on that synthetic content alongside human work, of course.
So here goes the loop for you. It all starts with human data, AI, synthetic data, AI, and then more synthetic data. This may indeed sound harmless enough. After all, if AI-generated content is good, then why shouldn’t another AI learn from it? Got intrigued? Allow me to indulge your curiosity.
The Copy Gets a Little Worse
Because we’re humans who tend to love to know everything behind everything, studies have been made regarding what happens when models get to train on their own generated data. This is when something called a “model collapse” took place in our dictionary of how to navigate AI in life.
Studies found out that indiscriminate recursive training can cause models to progressively lose information from the original data distribution source with less common or “tail” information disappearing first. Think of it as making a photocopy of a photocopy.
The first copy looks perfectly fine. The tenth still looks acceptable at some point. However, a few more copies and you’ll find yourself staring at a blurry version of something nobody remembers seeing nevertheless printing in the first place.
It Gets Interesting
Synthetic data isn’t automatically bad. AI-generated data can be useful, abundant and considerably cheaper than collecting new human examples. Researchers are actively exploring ways to use it without causing a model collapse. These methods include approaches that retain real data and carefully work on curating synthetic examples. This is when you learn that the real problem isn’t AI learning from AI. It’s AI learning from AI without knowing what it is learning from. See the difference?
The Human Data Paradox
The irony is almost too perfect. We built AI to generate more content because we wanted more content. Now, the more synthetic content we create, the more important genuinely human data becomes.
The weirdest outcome may not be an AI apocalypse. Sorry for disappointing you, Ultron, for the 100th time. It could simply be an internet where everything looks increasingly perfect, too similar, and too familiar… while becoming progressively less diverse, surprising and human.
This leaves us with an uncomfortable question: If AI eventually learns mostly from AI, will it still be learning about us? Or mostly learning about what it previously thought we were? Here’s something to think of at 3 AM.