Buried in the paper that made headlines for producing the first AI-written viral genomes is a stranger result than the headline suggests. The models did not simply reshuffle what evolution had already tried. They wrote genetic solutions that appear nowhere in any known natural sequence, and several of them worked.
The study, published in Science by Samuel King, Brian Hie, and colleagues at Stanford University and the Arc Institute, used the genome language models Evo 1 and Evo 2 to design complete bacteriophage genomes. Sixteen of them booted up inside E. coli and killed it.
The number that deserves attention is not 16. It is 13.
That distinction is what separates this result from a very expensive act of copying. A model that reproduces known biology is a compression tool. A model that reaches combinations the databases have never recorded, and whose output then replicates inside a living cell, is doing something closer to design.
Thirteen Genomes Carried Mutations Found Nowhere in the Databases
Each of the 16 working phages diverged from its closest natural relative by between 67 and 392 mutations, according to the Arc Institute's technical account of the work. Thirteen of them contained mutations the researchers could not locate in any known natural sequence.
One design, Evo-Φ2147, accumulated 392 mutations and shares an average nucleotide identity of 93.0 percent with the phage NC51. Under some taxonomic thresholds, the team notes that it would qualify as a new species.
The template for all of this was ΦX174, a small phage 5,386 nucleotides long, encoding 11 genes, and a landmark in its own right. It was the first complete genome ever sequenced, by Frederick Sanger's group, and later the first genome ever chemically synthesized, by Craig Venter's. Its genes overlap, which means a mutation in a shared stretch has to satisfy two proteins at once. That is a punishing constraint for a design system.
The Capsid Result Is the One Biologists Are Talking About
The most technically arresting finding involves a single protein.
One generated phage, designated Evo-Φ36, used a DNA-packaging J protein taken from G4, a considerably more distant relative. That protein is shorter than the one ΦX174 uses, 25 amino acids against 38. Slotting a mismatched component into a viral shell should not work, and earlier attempts to engineer exactly this swap by hand had failed.
Cryo-electron microscopy showed the shorter protein adopting a distinct orientation within the capsid and functioning there, implying that the model had made compensating changes elsewhere in the genome to accommodate it. That is a different order of achievement than editing a gene. It suggests the model was coordinating across the whole genome rather than optimizing one part at a time.
The Success Rate Was Low, and the Host Range Stayed Narrow
The scale of the failure is as informative as the success. From thousands of generated candidates, the team chemically synthesized and tested nearly 300 designs. Sixteen produced working phages, with a hit rate of 5.6 percent.
Screening that many designs required a purpose-built assay. Synthetic genomes were assembled, transformed into competent E. coli C, and monitored for growth collapse in 96-well plates, with infections showing a sharp drop in optical density within two to three hours.
That failure rate is worth holding onto. Genome design is not yet a matter of asking a model for a virus and receiving one. It is a matter of generating enormous numbers of candidates, filtering them computationally, building the survivors, and accepting that most will be inert.
Crucially, host range did not drift. All 16 functional phages infected only E. coli C and the related E. coli W strain, with no growth on six other tested strains. The team had required generated sequences to retain a spike protein similar to ΦX174's, because that protein determines what the virus can infect. Substantial genomic divergence, in other words, was compatible with tightly constrained targeting.
The team also evolved three E. coli strains resistant to ΦX174 through mutations in the waa operon, which governs bacterial surface receptors. Cocktails of the generated phages overcame resistance in all three within one to five passages, while ΦX174 alone failed completely. The phages that broke through were mosaics, recombinants combining elements from two or three separate AI designs.
The Governance Question Sits in the Same Issue
Thomas Inglesby and Moritz Hanke of the Johns Hopkins Center for Health Security published an accompanying commentary in Science arguing that generating functional viral genomes carries urgent biosafety and biosecurity implications. The capability to compose viral genomes with generative AI now exists, they wrote, while "the governance to safely steer it does not."
The Stanford and Arc team excluded sequences from viruses capable of infecting humans, animals or plants from the training data, worked with a phage that infects a non-pathogenic laboratory strain, and ran the experiments in dedicated biosafety cabinets. Their own recommendation is that safety and security specialists be involved from the outset of any whole-genome design effort.
For readers, the practical stakes are still distant. These are bacteria-killing viruses in petri dishes, not therapies. What changed is narrower and stranger: a model has now written functional biology that evolution, as far as the databases can tell, never got around to writing.
Key Questions Answered
What did the AI actually design?
Complete genomes for bacteriophages, viruses that infect bacteria. The models generated full genome sequences using the natural phage ΦX174 as a design template.
Why does the number 13 matter?
Thirteen of the 16 working genomes contained mutations not found in any known natural sequence, meaning the models produced combinations that have not been observed in nature.
What was unusual about Evo-Φ36?
It used a DNA-packaging protein from a more distantly related phage. Cryo-electron microscopy revealed the shorter protein working inside the capsid, a swap that had defeated earlier hand-engineering efforts.
Can these viruses infect people?
No. All 16 infected only two related laboratory E. coli strains and failed to grow on six others. Human, animal, and plant virus sequences were excluded from training.
How efficient was the process?
Thousands of candidate genomes were generated; nearly 300 were chemically synthesized and tested, and 16 proved viable, a success rate of nearly 5.6 percent.
Why are biosecurity experts concerned?
Because the same capability could, in principle, be pointed at pathogens. The accompanying commentary argues the technology has outpaced the oversight built to govern it.