A Memory Of Every Virus

Long before anyone thought of editing a gene, bacteria were already doing the hard part. A bacterium under viral attack cuts a fragment out of the invader and files it in its own chromosome, in an array of captured scraps that its descendants inherit. It is an immune system with a written memory, in an organism with no cells to spare. A Spanish microbiologist noticed the odd repeats in 1993 and worked out what they were for by 2005; a yoghurt company proved it in 2007; and by 2012 two researchers had shown the bacterium's search-and-destroy machine could be handed any address you liked. This is the story of the mechanism, not the medicine. And our own file's sources on it turned out to be, for the first time in this series, entirely correct where they existed, and missing in nine places where they should not have been.
Consider the problem from the bacterium's side. You are a single cell. You have no antibodies, no lymphocytes, no thymus, no capacity to learn in the way an immune system is usually said to learn. And you are under constant attack by viruses that outnumber you, that evolve faster than you, and that kill by injecting their own genetic material into you and turning your machinery against you. You get one shot. If you survive an infection, the useful thing would be to remember it.
Bacteria remember it. When a cell survives a viral attack, it cuts a short fragment out of the invader's DNA and files it into a dedicated stretch of its own chromosome, between short repeated sequences, like a scrap pasted into a ledger. The ledger is heritable. Its descendants are born already carrying the record, and they transcribe those scraps into RNA and use them as search terms, patrolling for the same sequence and destroying anything that matches. That is CRISPR. Everything that has happened since, the Nobel Prize, the patents, the medicine, the arguments about babies, is downstream of a filing system in a microbe.
01The Archive
Each spacer is a captured piece of something that tried to kill an ancestor. The name is a description of the architecture: Clustered Regularly Interspaced Short Palindromic Repeats. In the genome there is an array where short repeated sequences alternate with spacers, and each spacer is a fragment of a virus or plasmid that once attacked the cell line. Beside the array sit the cas genes, which encode the proteins that do the capturing and the cutting. Because the array is part of the chromosome, it is inherited. A bacterium is born carrying its lineage's list of old enemies, and it uses that list to recognise them again. This is an adaptive immune system, with a memory written down, in an organism that has no room for anything it does not need.

A Spanish microbiologist found the repeats in 1993 and worked out what they were for in 2005. Francisco Mojica, at the University of Alicante, first noticed the unusual repeated sequences in Haloferax mediterranei, an archaeon from a salt marsh. Nobody knew what they were. Over the following decade he kept at them, and by 2005 he had recognised the thing that makes the whole story: the spacers between the repeats matched foreign viral and plasmid DNA. If the sequences in the ledger were copies of viruses, the ledger was an immune record. Alexander Bolotin and Gilles Vergnaud reached the same recognition independently the same year. It is worth pausing on the shape of that: the central insight came from a curiosity-driven study of a microbe in a Spanish salt marsh, and it took twelve years to interpret.

The experimental proof came out of an industrial problem with cheese cultures. Rodolphe Barrangou and Philippe Horvath were working at Danisco, a food-ingredients company, on a thoroughly commercial difficulty: bacteriophages destroy the starter cultures used to make yoghurt and cheese, and that costs money. In 2007 they published in Science the demonstration that CRISPR provides acquired resistance against phages in Streptococcus thermophilus. They challenged the bacteria with phage, watched the survivors acquire new spacers matching that phage, and showed the survivors were now resistant. Add a spacer, gain immunity; remove it, lose immunity. The first proof that this was an immune system at all came from a dairy laboratory.
02From A Salt Marsh To A Human Cell In Twenty Years
The steps are short and each one is somebody's whole career. 2010: the Moineau group shows that Cas9 cuts the target DNA within the protospacer, so Cas9 is the blade. 2011: Charpentier's group identifies the tracrRNA, a second RNA needed to mature the guide, which is the missing component that made the system reproducible outside a bacterium. 2012: Jinek and colleagues, in the Doudna and Charpentier collaboration, fuse the two RNAs into a single synthetic guide and show that Cas9 can then be programmed to cut purified DNA at any chosen target. 2013: Feng Zhang, George Church and Doudna's groups independently get it working inside mammalian cells. Nineteen years from an unexplained repeat in a salt-marsh archaeon to editing a human genome.

03How It Finds One Address In A Whole Genome

Cas9 does not read the genome. It looks for a three-letter motif and only then checks the address. SpCas9 is a protein of 1,368 amino acids with two nuclease domains that cut one strand each: RuvC takes the non-target strand and HNH takes the target strand, and between them they make a blunt double-strand break. The search works like this. The enzyme is inert until it binds its guide RNA, at which point it undergoes a large conformational change that opens a channel for scanning DNA. It then hunts for a PAM, a protospacer adjacent motif, which for SpCas9 is the three letters 5'-NGG-3'. Finding one, the PAM-interacting domain melts the helix locally, and the guide RNA tries to base-pair with the exposed strand. Only on full twenty-nucleotide complementarity do the nucleases fire, cutting three base pairs upstream of the PAM.
The twenty letters of the guide are not equally strict, and the slack end is the whole problem. Specificity is dominated by roughly twelve seed nucleotides nearest the PAM. A mismatch in that region strongly reduces cutting. A mismatch at the far end may be tolerated. So the guide is really a short strict password followed by a longer, laxer one, and the laxity at the tail is exactly where off-target cuts come from. This is not a flaw someone introduced. It is a property of a system that evolved to catch viruses which are themselves mutating, where a little tolerance for mismatch is an advantage. What is a virtue in a bacterium under attack is a liability in a tool meant to hit one site and no other.
Everything called gene editing is done by the cell's own repair machinery. Cas9 breaks the DNA and stops. What happens next is the cell's response to a double-strand break, and there are two routes. Non-homologous end joining simply shoves the ends back together; it is error-prone and typically leaves small insertions or deletions, which disrupt the gene. That is how a gene gets switched off. Homology-directed repair copies from a template, so if a donor sequence is supplied the cell will write that in, which is how a sequence gets corrected or inserted. Which route the cell takes depends on the cell type and where it is in its cycle, and that is why the same edit works beautifully in one tissue and poorly in another.
04The Generation That Stopped Cutting
If the break is the dangerous part, remove the break. Base editors do exactly that. Komor and colleagues in David Liu's laboratory fused a catalytically dead Cas9, which still finds and holds the target but cannot cut, to a cytidine deaminase, converting C to T chemically, in place, in 2016. Gaudelli and colleagues in the same laboratory then evolved a modified transfer-RNA adenosine deaminase to convert A to G, in 2017. Neither makes a double-strand break at all. The enzyme becomes a delivery vehicle: it stops being scissors and starts being a way to carry a small chemical modification to a chosen letter.
Prime editing lets the guide RNA specify not just where, but what. Anzalone and colleagues, again in Liu's laboratory, published it in 2019. A Cas9 nickase, which cuts only one strand, is fused to a reverse transcriptase, and the guide is extended into a pegRNA that carries both the address and the replacement text. The reverse transcriptase writes the new sequence directly into the nicked strand. It can make any of the twelve possible point mutations, insertions up to about 44 base pairs and deletions up to about 80, with no double-strand break and no donor template. The name is precise: search, and replace.
The amount of machinery built purely to detect mistakes is itself the answer about precision. Our file names four separate methods developed to map off-target cleavage across a whole genome: GUIDE-seq, CIRCLE-seq, Digenome-seq and DISCOVER-seq. Alongside them sit engineered high-fidelity variants, eSpCas9, SpCas9-HF1 and HiFi Cas9, which reduce off-target activity by altering the protein-DNA interface. For any therapeutic use, edited cells get whole-genome sequenced to check what else happened. Nobody builds four detection platforms for a problem that does not exist.
It is not molecular scissors and it is not perfect. Our file marks the early media framing debunked, and the list of what has actually been documented at cut sites is specific: off-target effects, mosaicism, large-scale chromosomal rearrangements, and chromothripsis, which is the shattering of a chromosome and its faulty reassembly. The file's own verdict is the one to keep: CRISPR is powerful but not infallible. And note where that refusal comes from. Every item on that list is a finding published by the same field that built the tool, which is roughly the opposite of a technology being oversold.
A gene drive is designed to spread through a wild population faster than Mendel allows. Ordinarily an engineered allele passes to half the offspring. A CRISPR gene drive copies itself onto the partner chromosome, so nearly all the offspring inherit it, and it can sweep through a population. Kyrou and colleagues showed in 2018 that a drive targeting the gene doublesex produced complete population suppression in caged Anopheles gambiae, the mosquito that carries malaria. Andrea Crisanti and Austin Burt at Imperial College London led the work. Our file states the position without decoration, and it has not changed in the file since: field release has not occurred, because of ecological and ethical concerns.
The hard part is no longer cutting. It is getting there. Our file names the constraint precisely: adeno-associated virus vectors, the workhorse of gene delivery, have a packaging limit of about 4.7 kilobases, and SpCas9 plus its guide RNA does not fit. Lipid nanoparticles do work, but currently deliver efficiently mainly to the liver. So a tool that can in principle address any sequence in any genome is limited in practice by a cargo-space problem and an organ preference. That constraint, incidentally, is why smaller Cas9 proteins from other bacterial species are of interest at all, and why the enzyme in this article's opening image is not the famous one.
One sentence on that, because it is a different story from this one. Casgevy, approved by the FDA on 8 December 2023 for sickle cell disease and transfusion-dependent beta-thalassemia, was the first approved CRISPR therapy anywhere, and it works by an elegant sidestep: rather than repair the faulty adult haemoglobin gene, it edits a patient's own blood stem cells to disrupt an enhancer of BCL11A and switch the fetal haemoglobin gene back on. The clinical account, and the ethics of editing embryos, belong to our companion articles on gene therapy and on designed humans. This one is about the bacterium.
05A Note On Our Own Sources, And The First Clean File
Twenty-six research files into this project, this is the first with no wrong pointer in it. Five of the fourteen entries carry an identifier. All five resolve exactly, on every field: journal, volume, issue, pages, year, authors. Jinek 2012 in Science, Cong 2013 in Science, Barrangou 2007 in Science, Mojica 2005 in the Journal of Molecular Evolution, Nishimasu 2014 in Cell. And not one of them is a review, which is what the pattern we have been testing across this series requires of a bibliography containing no books. We wrote that prediction down before checking, and also wrote down that it was a weak test, because these are famous papers whose identifiers look right on their face. Nothing failed, so nothing could fail in an informative way.
That is the real finding, and all nine of them should have had one. The other nine entries are papers in Nature, Science, the New England Journal of Medicine, Nature Biotechnology and the Annual Review of Biochemistry. Journals of that kind register an identifier for everything they publish; there is no such thing as one of these papers not having one. We found all nine and verified each against our file's own claimed volume, issue, pages, year and first author. Every one matched. That is a sixty-four per cent omission rate in the easiest case there is, and it is the largest single recovery we have made from one document.
An empty identifier field is not automatically wrong. It is not automatically right either. An earlier article in this series worked through a bibliography of old German monographs where five identifiers were all wrong, and established something counterintuitive: the correct repair there was to leave the field blank, because those books have no registered identifier anywhere in the world and filling the gap could only mean filling it with something false. This file is the mirror image. Here nine blanks are all defects. Put the two side by side and you get the rule: whether a blank is honest depends entirely on whether that kind of work gets registered at all. Anyone repairing a bibliography needs both halves. With only the first, they will leave modern papers bare; with only the second, they will invent identifiers for books that never had any.
Searching for the missing nine kept turning up other records wearing the paper's own title. A Faculty Opinions recommendation of the prime-editing paper. A Commentary on the adenine base-editor paper, in a different journal. A Publisher Correction to that same paper. And for the sickle cell trial, a piece of correspondence in the same journal whose title is nearly identical to the article's. Earlier in this series we described this decoy problem as something that happens to books, on the reasoning that a book has no neighbouring article and the only other record carrying its title is a review of it. That reasoning was too narrow. Articles accrete corrections, correspondence, commentaries and recommendation records. The decoy is a property of being worth citing, not of being a book.
The New England Journal of Medicine writes the answer into the string. Its identifiers carry a stem: nejmoa marks an original article, nejmc marks correspondence. The two records for the sickle cell trial differ by exactly that, and knowing the convention settles in one glance which is the paper and which is a letter about it. Almost no other publisher offers that help, and the reason it matters is the mechanism underneath all of this: when a bibliography is assembled by matching titles, the satellite and the paper look identical, and the identifier is the only thing that can tell them apart. Which is precisely the thing that keeps going missing.
The entry for the man who started all of this has his name broken in half. Our file's author column reads Mojica, Francisco J, with M., et al stranded out in a column of its own. The registry gives his name as Francisco J. M. Mojica. The orphaned M. is part of the man's name, cut loose and parked in the wrong field. We have now found this same mechanism shattering author lists in four different sections of our library, and it always breaks at an initial. It is a small thing next to a wrong citation, and it is worth recording because it is mechanical, repeatable, and therefore fixable everywhere at once.
| Entry | In Our File | Result |
|---|---|---|
| 1. Jinek et al 2012, Science 337(6096):816-821 | 10.1126/science.1225829 | EXACT on every field |
| 2. Cong et al 2013, Science 339(6121):819-823 | 10.1126/science.1231143 | EXACT |
| 3. Barrangou et al 2007, Science 315(5819):1709-1712 | 10.1126/science.1138140 | EXACT |
| 4. Mojica et al 2005, J Mol Evol 60(2):174-182 | 10.1007/s00239-004-0046-3 | EXACT. And the registry gives the author as Francisco J. M. Mojica, where our file has the M. stranded in another column |
| 5. Nishimasu et al 2014, Cell 156(5):935-949 | 10.1016/j.cell.2014.02.001 | EXACT. This paper's own figure appears above |
| 6. Komor et al 2016, Nature 533(7603):420-424 | no identifier | RECOVERED 10.1038/nature17946, verified |
| 7. Anzalone et al 2019, Nature 576(7785):149-157 | no identifier | RECOVERED 10.1038/s41586-019-1711-4. A plain title search returns a Faculty Opinions recommendation of it instead |
| 8. Frangoul et al 2021, NEJM 384(3):252-260 | no identifier | RECOVERED 10.1056/nejmoa2031054. Correspondence with a near-identical title sits at nejmc2103481 |
| 9. Tsai et al, Nature Biotechnology 33(2):187-197 | no identifier | RECOVERED 10.1038/nbt.3117. Dated 2014 online where our file says 2015 in print; both defensible |
| 10. Kyrou et al 2018, Nat Biotechnol 36(11):1062-1066 | no identifier | RECOVERED 10.1038/nbt.4245, verified |
| 11. Doudna and Charpentier 2014, Science 346:1258096 | no identifier | RECOVERED 10.1126/science.1258096. The suffix is the article number our file gives, which confirms it twice over |
| 12. Ledford 2019, Nature 570(7761):293-296 | no identifier | RECOVERED 10.1038/d41586-019-01906-z. News journalism, not a research paper, and the stem says so |
| 13. Gaudelli et al 2017, Nature 551(7681):464-471 | no identifier | RECOVERED 10.1038/nature24644. A Commentary on it and a Publisher Correction to it both carry its title |
| 14. Wang et al 2016, Annu Rev Biochem 85:227-264 | no identifier | RECOVERED 10.1146/annurev-biochem-060815-014607, verified |
Fast Facts
- What CRISPR is
- An array in a bacterial genome where short repeats alternate with spacers, each spacer a captured fragment of a virus or plasmid that attacked the lineage. It is heritable, which makes it an immune memory
- Discovery
- Francisco Mojica noticed the repeats in Haloferax mediterranei in 1993 and recognised in 2005 that the spacers matched foreign DNA. Bolotin and Vergnaud reached it independently the same year
- First proof
- Barrangou and Horvath at Danisco, 2007, in Streptococcus thermophilus, a yoghurt and cheese bacterium. The motivation was phages destroying starter cultures
- The 2012 result
- Jinek et al. fused the two natural RNAs into one synthetic guide and showed Cas9 could be programmed to cut any chosen DNA target. Nobel Prize in Chemistry 2020 to Doudna and Charpentier
- The enzyme
- SpCas9, 1,368 amino acids, two nuclease domains: RuvC cuts the non-target strand and HNH the target strand, together making a blunt double-strand break
- Finding the target
- It hunts for a PAM, three letters (5'-NGG-3' for SpCas9), melts the helix there, then requires full 20-nucleotide guide complementarity before cutting, 3 base pairs upstream of the PAM
- Why off-targets happen
- Specificity is dominated by about 12 seed nucleotides nearest the PAM. Mismatches at the far end can be tolerated
- Who does the editing
- The cell, not Cas9. NHEJ leaves error-prone indels that disrupt a gene; HDR copies from a supplied template to correct or insert
- Without breaking DNA
- Base editors (Komor 2016 for C to T; Gaudelli 2017 for A to G) and prime editing (Anzalone 2019), which can make all 12 point mutations, insertions to about 44 bp and deletions to about 80 bp
- Gene drives
- Kyrou et al. 2018 achieved complete population suppression in caged Anopheles gambiae by targeting doublesex. No field release has occurred
- The bottleneck
- Delivery. AAV vectors cap at about 4.7 kb, too small for SpCas9 plus guide; lipid nanoparticles currently reach mainly the liver
- Refused
- That CRISPR is precise molecular scissors. Off-target cuts, mosaicism, large chromosomal rearrangements and chromothripsis are all documented
- Identifier audit
- 14 entries checked 26 August 2026. 5 had identifiers and all 5 are exact, the first clean file in this series. 9 had none and all 9 were recovered and verified
What We Can Actually Stand Behind
CRISPR is a genuine adaptive immune system in bacteria and archaea, storing captured fragments of past invaders as heritable spacers. Mojica saw the repeats in 1993 and interpreted them in 2005; Barrangou and Horvath proved the immune function experimentally in 2007; tracrRNA was identified in 2011; Jinek and colleagues made it programmable in 2012; three groups had it working in mammalian cells by 2013; Doudna and Charpentier took the 2020 Nobel Prize in Chemistry. SpCas9 has 1,368 amino acids and two strand-specific nuclease domains, requires a PAM and full 20-nucleotide complementarity, cuts blunt 3 base pairs from the PAM, and leaves the repair to the cell. All five identifiers our file carries for this material resolve exactly.
Base editing and prime editing do what is claimed for them and avoid the double-strand break, though their long-term behaviour in tissues is still being characterised. Off-target cutting is real, is detectable by at least four published genome-wide methods, and is reduced but not eliminated by high-fidelity variants. Gene drives have achieved complete suppression in cages and have not been released. Delivery, not cutting, is the binding constraint, and the AAV packaging limit is the specific reason.
Heritable germline editing is genuinely unsettled and is deliberately left to our companion article on designed humans, as the clinical story is left to our article on gene therapy. This is a boundary our own library sets between three articles on the same subject, and it is worth naming rather than quietly observing.
CRISPR is not precise molecular scissors. Off-target effects, mosaicism, large-scale chromosomal rearrangements and chromothripsis are all documented at cut sites, and the documentation comes from the field that built the tool. Powerful, and not infallible.
Fourteen entries checked live on 26 August 2026 against a prediction recorded beforehand. Five carried an identifier and every one is exact, which makes this the first file in this series with no wrong pointer in it. Nine carried no identifier at all, in a bibliography made entirely of modern journal articles, and all nine were recovered and verified field by field. Put beside an earlier file where the correct repair was to leave fields blank, this gives the rule that a blank identifier is neither automatically right nor automatically wrong. The recovery work also showed that the same-title decoy problem, which we had described as a book phenomenon, applies to articles too, through corrections, correspondence, commentaries and recommendation records. Treat five out of five as a weak test, which we said in advance: nothing failed, so nothing could fail informatively.
Sources & further reading
The biology above is drawn from our research library on Theories of Anything. Every identifier below was checked live against Crossref on 26 August 2026. Five were carried by our file and all five proved exact; the other nine our file omitted entirely, and each was recovered and then verified field by field against the journal, volume, issue, pages, year and first author our file claims. Open the full research file to check the sourcing and go deeper.
Image credits
- Cas9 from Staphylococcus aureus with guide RNA and target DNA, based on PDB 5AXW Thomas Splettstoesser (scistyle.com), via Wikimedia Commons (CC BY-SA 4.0). CC BY-SA 4.0 Source.
- Simplified diagram of a CRISPR locus: cas genes, leader and repeat-spacer array AnnaJune, via Wikimedia Commons (CC BY-SA 3.0). CC BY-SA 3.0 Source.
- Bacteriophage, transmission electron micrograph with 100 nm scale bar SnaxMikn, via Wikimedia Commons (CC BY-SA 4.0). CC BY-SA 4.0 Source.
- Crystal structure of Cas9 in complex with guide RNA and target DNA Hiroshi Nishimasu, F. Ann Ran, Patrick D. Hsu and colleagues, via Wikimedia Commons (CC BY-SA 3.0). CC BY-SA 3.0 Source.
- Emmanuelle Charpentier and Jennifer Doudna, a composite of two separate portraits Charpentier portrait by Bianca Fioretti of Hallbauer and Fioretti; composite via Wikimedia Commons (CC BY-SA 4.0). CC BY-SA 4.0 Source.