His name was Frederick Sanger. He is sometimes called the father of genomics and almost nobody outside a lab has heard of him. He went home to his garden most evenings and disliked fuss so thoroughly that he turned down a knighthood, because he did not fancy being called Sir. He won the Nobel Prize in Chemistry twice, in 1958 and again in 1980, and was the only person to have done so until Barry Sharpless matched him in 2022. The second prize was for reading DNA.
Here is what he cracked. We knew DNA was a string of four letters. We had no way to read their order. Order is the whole game. It is the difference between a healthy gene and a lethal one, the same way "united" and "untied" are the same letters behaving very differently.
The honest layman's version, which is the only version I am qualified to give: he made copies of a DNA strand that each stopped at a known letter, sorted them by length, and read the sequence off the ladder. Slow, fiddly, reliable. It ran as the standard for thirty years.
The step at the end
It is not a museum piece either.
I sat in a lab in Hyderabad while a result went back for confirmation. The sequencer had produced it and the pipeline had called it. The request to check it again came from the clinician, who wanted the finding seen once more by the older method before acting on it.
When a modern sequencer flags a variant that will change how a patient is treated, the result is, in plenty of labs, still confirmed with Sanger's own chemistry before it reaches the report. The run that produced the finding is fast, parallel and statistical. The check that lets it leave the building is a single reaction from 1977.
Whether that check is still necessary is a genuine argument, and it is worth noticing who is having it with whom. The case against is technical. Coverage is deep, confidence scores are high, the second read adds days and cost and usually agrees. The case for, in the room I was sitting in, was not technical at all. It came from the person who had to act on the result and wanted it seen once more first.
Those are different problems. One of them gets solved by better data. The other one is about who signs, and it does not go away when the specifications improve.
The argument is its own quiet compliment to the old method. It is the thing people trust enough to debate retiring.
The part I nearly got wrong
There is a second thing in this story worth checking, and I nearly published it wrong myself.
You will read that Sanger sequenced the first genome in 1977. It is a lovely story and it is slightly wrong. A year earlier, in 1976, the Belgian scientist Walter Fiers had already read the complete genome of a virus called MS2. So MS2 was first. The catch is that MS2's code is written in RNA. Sanger's virus, phiX174, was the first genome written in DNA, the molecule your own genes are made of.
First genome, MS2. First DNA genome, phiX174. The distinction costs nothing and marks you as someone who actually checked.
Having made that correction, I then wrote that phiX174 is 5,386 letters long, and cited the 1977 paper for it.
The 1977 paper does not say that. It reports approximately 5,375 nucleotides. The figure of 5,386 comes from the complete sequence published the following year, in 1978. Eleven bases, in a piece whose entire argument was about getting details right.
Two more things that paper says, which most people repeating the story have not read. It describes nine known genes. The modern annotation is eleven. And it was produced with the plus and minus method, not with chain termination. The dideoxy method that everyone now calls Sanger sequencing was published separately, in the same year, in a different journal. The first DNA genome was not sequenced by Sanger sequencing in the sense that phrase is normally used.
None of this reduces the achievement. It changes what you are entitled to say about it.
The 1977 paper also carried a genuine surprise. Two pairs of genes were coded by the same stretch of DNA, read in different frames to mean different things, like a sentence that reads one way forwards and another if you start on the second word. Evolution had invented file compression. The number of known overlaps has gone up since.
Selling the grandchildren
phiX174 is barely alive. Yours is about three billion letters. The virus was never the point. The point was proof that a genome is finite, that it has a first letter and a last letter, and that we could travel from one to the other. Everything after that is a question of scale, and scale is a subject for another day.
When I set up Genique's first lab, I opened every box we bought and taught myself what each machine did. Thermal cyclers, qPCR, the lot. I could tell you what they cost and how they earned their keep. I still could not have told you who Sanger was. I was selling the grandchildren of his method before I knew the man existed. Still am: real-time PCR from Seasun Biomaterials, pharmacogenetics kits from Diatech Pharmacogenetics, a polymerase and a primer and a readout.
I have spent the years since then learning this field by osmosis, which mostly means learning what I am not entitled to claim. A specification I have not seen in current partner documentation. A performance number I cannot source. A date I read in a blog post and repeated.
I sell instruments that read three billion letters. The last word on what they found is often still a single reaction from 1977.
Sources
Sanger F, Air GM, Barrell BG, Brown NL, Coulson AR, Fiddes CA, Hutchison CA, Slocombe PM, Smith M. Nucleotide sequence of bacteriophage phi X174 DNA. Nature 265, 687-695 (1977). Reports approximately 5,375 nucleotides, nine known genes, two pairs of overlapping genes, plus and minus method.
Sanger F, Coulson AR, Friedmann T, et al. The nucleotide sequence of bacteriophage phiX174. J Mol Biol 125, 225-246 (1978). Source of the 5,386 figure.
Sanger F, Nicklen S, Coulson AR. DNA sequencing with chain-terminating inhibitors. PNAS 74, 5463-5467 (1977). The dideoxy method.
Fiers W, Contreras R, Duerinck F, et al. Complete nucleotide sequence of bacteriophage MS2 RNA: primary and secondary structure of the replicase gene. Nature 260, 500-507 (1976). The MS2 genome is 3,569 nucleotides.
Modern phiX174 annotation of eleven protein coding genes: NCBI reference sequence NC_001422.
Siddhartha Mukherjee, The Gene: An Intimate History, for the narrative background.