Every living thing faces the same challenge of reproducing itself. The TWO primary strategies are copying and merging, and both choices shape nearly everything that follows.
Copying is relatively simple with a cell dividing, and each daughter cell inherits the parent’s full instruction set. Most of the time, that is exactly what an organism needs. Copying closes wounds, enables growth, and replaces the lining of the gut every few days. But replication also creates opportunities for error. Most errors are repaired or eliminated but when one escapes, its descendants inherit it.
The strength of copying becomes dangerous when a system faithfully repeats what it has become. Cancer, for example, arises when ordinary capacities for growth, survival, and replication are released from the signals that normally tell them when to stop.
Merging, by contrast, combines genetic material from two lineages to produce new gene combinations. Recombination gives natural selection an option that strict copying cannot by breaking apart an inherited package. Harmful alleles may separate from beneficial neighbors, while beneficial alleles from different lineages can come together.
Sexual reproduction carries costs that copying avoids: organisms must find mates, coordinate timing, and transmit only half of each parent’s nuclear genome. Under many conditions, asexual reproduction therefore has an immediate advantage. Yet sex remains widespread across complex life, suggesting that recombination has delivered enduring benefits over billions of years of natural selection.
More differences, however, are not always better. When populations are too distant or adapted to incompatible conditions, useful combinations can break apart, forming what biologists call outbreeding depression. The goal is not maximum distance but compatible difference: enough variation to create novelty, enough similarity to remain functional. Difference that fits.
Artificial intelligence is beginning to reproduce the same choice. The open internet is filling with model-generated output, and future models will inevitably train on some of it. That alone is not the problem. Synthetic material can expand a system’s range when it is tested against something outside the model: verifiable mathematics, executable code, falsifiable experiments, or measurements that can prove an answer wrong. External reality remains in the loop, able to contradict the system.
Satya Nadella described Microsoft’s approach in similar operational terms: “We train models against the actual product harness, interactions, and outcomes they will encounter.” The crucial word is outcomes. A model is not rewarded merely for reproducing its own distribution; it must answer to whether something worked outside itself. The danger appears when that loop closes. Like biological copying, a system begins learning from material already shaped by its own assumptions, leaving too little independent information to correct them. The model begins to inherit itself. Researchers call one form of this model collapse, though the term can be misleading because the decline is often gradual rather than sudden.
In closed recursive training, without enough fresh data or external correction, the first sign of decline is not gibberish. Instead, low-frequency material like rare constructions, unusual events, minority patterns, and edge cases gradually disappears. The distribution contracts toward its center. Each generation becomes a more confident rendering of the previous generation’s average while remaining fluent throughout the decline.
This resembles a cell inheriting itself without limits as errors accumulate, and range disappears before competence seems to. Because the degradation is smooth, the system may still look polished and sound intelligent even as it loses the rare, unfamiliar, and contradictory material that once expanded its understanding. The problem, then, is not synthetic data itself. When curated, verified, and anchored to independent reality, synthetic data can create controlled examples, surface edge cases, and extend coverage. The danger is allowing the machine to become its own environment.
When a system trains on outputs already filtered through its own assumptions, with nothing outside the loop to correct it, copying begins to replace recombination. The center strengthens, the tails vanish, and the model grows more certain of a world that increasingly resembles itself. That is the copying path.
Merging is harder for AI because it requires material with an independent origin: human experience, original reporting, experiments, measurements, physical feedback, dissent, adversarial testing, and facts produced through contact with the world rather than by another model. As recursive copying accelerates, this independent material becomes harder to find. The system increasingly encounters information already shaped by systems like itself.
The key distinction is not simply human versus machine, but independent provenance versus recursive inheritance. A model needs contact with histories it did not generate, because only something outside its existing structure can contradict it, expand it, or create a genuinely new combination. Independent material can expose an assumption, restore a missing edge case, or produce an insight the system could not reach by repeatedly refining its own outputs. Without that contact, learning becomes inheritance: fluency may improve even as range narrows.
Mamoru Oshii dramatized the same argument in the 1995 film Ghost in the Shell. The Puppet Master has become self-aware but recognizes that copying is insufficient. Identical copies share the same point of failure in that one flaw can destroy them all. It therefore seeks offspring with variation but since it cannot create it alone, accepts mortality because variation requires surrendering perfect continuity.
The same pattern applies to in human learning when someone who reads only what confirms existing beliefs. It can feel like learning because fluency remains even as perspective narrows. A person may absorb more facts, cite more sources, and speak more persuasively, yet become less open when every input passes through the same assumptions. An open mind encounters material that did not originate within its existing structure. That material must be familiar enough to be intelligible but different enough to create something new. Too little common ground produces noise; too little difference produces repetition. Merger occurs in the narrow space between them.
A copy inherits. A merger creates what no single history could produce alone. We are building the most consequential learning systems in history, but their path will not be chosen in a single moment. It will emerge cumulatively, one training run at a time, as material with an independent origin quietly becomes harder to find.
The question is not whether machines can learn, but whether we will leave them anything independent to learn from.



