Last updated: 2026-07-23

02. Gene sequencing — turning life into data

Atoms to Bits

What sequencing is

What does it mean to “sequence” a genome?

Atoms to bits
Atoms to bits

Gene sequencing (more precisely DNA sequencing) is the process of determining the exact order of bases — A, C, G, T — along a DNA molecule. When the target is an organism’s complete DNA, we call it whole-genome sequencing (WGS). When the target is only the protein-coding exons, it is whole-exome sequencing. When the target is a panel of disease genes, it is targeted sequencing.

Sequencing does not by itself “understand” disease or design a therapy. It produces a digital string (and quality scores). Interpretation — mapping reads to a reference, calling variants, linking variants to function — is a separate computational and clinical layer. Still, without cheap, accurate reads, the rest of synthetic biology starves for ground truth.

In OOM EmTech terms, Gene Sequencing is the atoms-to-bits leg of Synthetic Biology: physical polymer → information.

From moonshot to industrial routine

How did we go from a multi-billion-dollar project to a routine lab service?

The Human Genome Project (HGP) produced the first high-quality draft of a human reference genome around the turn of the millennium, with a project cost on the order of $2.7 billion and a multi-year international effort dominated by Sanger sequencing. That price was not a sticker for “one person’s clinical report”; it was the cost of building the map itself.

The field then entered a long cost collapse driven by next-generation sequencing (NGS): massively parallel short reads, better chemistry, better cameras and fluidics, better base-calling software, and brutal competition. The U.S. National Human Genome Research Institute (NHGRI) tracked sequencing costs for years; the public graphs show orders-of-magnitude decline after ~2007 as NGS displaced capillary Sanger for large projects.

Industrial milestones that matter for orientation (not an exhaustive vendor history):

  • ~$1,000 genome era (mid-2010s): whole human genomes became thinkable for large research cohorts and some clinical use cases.
  • Illumina NovaSeq X series (announced/positioned in the early 2020s): vendor claims of list economics around ~$200 per 30× human genome and on the order of 20,000+ genomes per year per high-end instrument, with marketing language of roughly $2/Gb. These are list/reagent-framed figures, not a guarantee of hospital all-in cost including labor, interpretation, compliance, and confirmatory testing.
  • Challenger platforms (long-read leaders such as PacBio and Oxford Nanopore, and low-cost short-read contenders) keep pressure on price, read length, and completeness.

Trend classification (mandatory)

Trend: Exponential decline in cost per human-equivalent genome (and cost per gigabase) for high-throughput short-read sequencing.
Metric: Inflation-adjusted cost per genome / cost per Gb at defined coverage (NHGRI methodology historically used ~3 Gb genome and platform-specific coverage assumptions).
Period: Roughly 2001–2020s for the dramatic NGS-era collapse; HGP-era Sanger baseline in the 1990s–early 2000s.
Rate: Multi-order-of-magnitude drop over ~15–20 years of NGS maturation — far faster than a simple Moore’s-law chip curve alone would suggest for the full stack. Recent years show continued improvement but with a flattening risk as chemistry and optics near local optima and as interpretation, not raw bases, becomes the bottleneck.
Mechanism: Digitization + parallelization + learning-curve manufacturing + software substitution in base calling and secondary analysis + market competition. Positive feedback: cheaper sequencing → more datasets → better algorithms and clinical utility → more instrument demand → more R&D.
Bottlenecks & next paradigm: Sample prep and clinical interpretation costs; rare-error modes that matter for diagnostics; long repetitive regions poorly covered by short reads; privacy and consent for population-scale data. Next paradigms already visible: high-accuracy long reads as first-line rather than specialty; multi-omics (RNA, methylation, proteins) as the default companion; on-device real-time sequencing; AI variant effect prediction reducing expert hours per genome.
Reach: When genomes are cheap, medicine, agriculture, pandemic surveillance, and industrial strain engineering all shift from sampling a few genes to routine whole-system reads.

Classification: Exponential for raw sequencing price-performance across the NGS era; not automatically exponential forever. Treat late-stage claims of “$100” or “$10” genomes as conditional on what is included in the price.

What the machines actually do (intuition, not a manual)

How does a sequencer read letters it cannot see with the naked eye?

Different platforms use different physics, but the shared idea is: break DNA into fragments, generate many parallel signals that depend on base identity, convert signals to letters, then reassemble by computation.

  • Short-read NGS (Illumina-class): fragments are amplified on a surface; each cycle adds a fluorescently labeled base; a camera records colors; software calls bases. Extremely high throughput; limited native read length.
  • Long-read (nanopore / SMRT): a single molecule is read for thousands to millions of bases, improving structural variant detection and telomere-to-telomere completeness, historically with different error profiles that improved rapidly.

Assembly is like reconstructing a book from millions of torn phrases. A reference genome makes “resequencing” easier: you align reads to the map and look for differences (variants). De novo assembly without a good reference is harder and more expensive computationally.

Capabilities unlocked by cheap sequencing

What can we do now that was fantasy at HGP prices?

Capability What it enables
Clinical rare-disease diagnosis Find causal variants in undiagnosed patients
Cancer genomics Tumor-normal comparison; therapy matching; MRD monitoring
Pathogen surveillance Track outbreaks and resistance genes in near real time
Population biobanks Association studies at millions of people
Breeding and ag Genomic selection in crops and livestock
Synbio DBTL loop Verify constructs; evolve strains; debug designs
Ancient DNA Recover genomes from fossils (foundation for de-extinction inputs)

Limits and honest caveats

What does a cheap genome still fail to give you?

  • A sequence is not a phenotype. Most variants are benign or ambiguous (VUS — variants of uncertain significance).
  • List price ≠ clinical bill. Counseling, orthogonal confirmation, and liability dominate in healthcare settings.
  • Short reads miss structure. Large insertions, repetitive regions, and some haplotypes need long reads or specialized assays.
  • Privacy is a one-way door. Genomes re-identify; relatives are implicated without consenting.
  • Equity: instruments and bioinformaticians cluster in rich systems; the data gap can widen health gaps even as per-genome reagent cost falls.

Milestone map (sequencing)

HGP reference map → NGS parallelization → $1,000 research genome → population biobanks → $200-class industrial WGS claims → long-read completeness / multi-omics default → (aspirational) always-on clinical genome + continuous methylome as vital signs.

Bottom line

Gene sequencing is the measurement layer of the living world. Its cost curve is one of the cleanest exponential EmTech stories of the early twenty-first century, with the important footnote that value has shifted from generating letters to using them safely and wisely. Every later chapter assumes this digital mirror of life exists and keeps getting cheaper.

← DNA · genes · proteinsGene editing →