Skip to content
Sunday, August 30, 2026
Engevity NewsScience & health
Research · Learning · Evidence
Science News

What the telomere-to-telomere genome added to 'complete' human DNA

The 2022 T2T consortium sequence filled the last eight percent of the human genome — centromeres, satellite arrays, and other tracts the 2003 draft could not read.

Diagram comparing 2003 draft genome with gapless T2T sequence

The telomere-to-telomere (T2T) genome, published in the journal Science in 2022 by a consortium led by the National Human Genome Research Institute, added roughly 200 million base pairs that the famous 2003 draft had left unread. The additions are concentrated in centromeres — the pinched middles of chromosomes — and in other repetitive tracts that older sequencing technology could not resolve. For the first time, every chromosome of a human genome was sequenced end to end with no gaps.

What was missing from the 'complete' genome of 2003?

The Human Genome Project declared success in 2003 with about 92 percent of the genome in hand. The missing eight percent was not a rounding error; it was a technological blind spot. The project's sequencing method read DNA in short fragments, and assembly software stitched those fragments together by finding overlaps. Repetitive regions — where the same motif repeats thousands of times, sometimes for millions of base pairs — defeat that logic, because fragments from opposite ends of a repeat look identical.

What fell in the blind spot was not junk by default. Centromeres, the attachment points that pull chromosomes apart during cell division, sit inside some of the largest repeats. Short arms of the five acrocentric chromosomes, ribosomal RNA gene arrays, and duplicons — copied segments that fuel evolution but also disease — were all fragmentarily assembled or absent.

How did newer technology close the gaps?

Two changes mattered. First, long-read sequencing, commercialized by Pacific Biosciences and Oxford Nanopore, reads DNA in stretches tens of thousands to hundreds of thousands of bases long — enough to span a repeat and anchor it uniquely. Second, the T2T consortium used a peculiar cell line, CHM13, derived from a hydatidiform mole. This tissue carries two identical copies of each chromosome rather than the usual paternal and maternal versions, which spared the team the hardest assembly problem of all: telling a paternal repeat array from a maternal one.

The result, the T2T-CHM13 reference, was published across six papers in Science in March and April 2022. In 2023, a separate effort completed the human Y chromosome end to end for the first time, reported in Nature, and in 2024 the same community released a draft human pangenome — 94 ancestry-diverse genomes assembled with T2T-era methods, published in Nature.

What did the new 200 million base pairs contain?

The T2T consortium's annotations surprised even close observers. The added DNA contained roughly 2,000 predicted genes, most of them partial or non-coding, though about 100 to 115 may encode proteins — many already implicated in disease because they sit near known variant hotspots. The sequence also clarified the architecture of centromeres, showing that each is built from a repeated 171-base-pair unit, the alpha satellite, arranged in hierarchical blocks whose sizes help explain how centromeres form and occasionally move.

One concrete payoff involves segmental duplications. Because these duplicated blocks previously collapsed into single assemblies, medical sequencing had to mask them; clinically relevant variants inside them went uncalled. The T2T reference restored them to analysis.

How do we know the new sequence is right?

The consortium validated the assembly in several independent ways, and its papers describe the checks plainly. Ultralong nanopore reads spanning entire repeat arrays confirmed continuity. Optical maps — fluorescent labels on long DNA molecules — verified the spacing of landmarks. For centromeres, chromatin profiling located the sites where the cell's division machinery actually binds, and those functional marks landed on the newly assembled alpha-satellite arrays rather than beside them.

Limitations remain stated in the literature. CHM13 is a single haplotype of largely European descent, so it is a structural reference, not a diversity baseline — the reason the pangenome effort exists. Some regions, including parts of the acrocentric short arms and arrays of ribosomal genes, were resolved only within groups rather than positioned with full confidence. And a reference genome is a scaffold for interpretation, not a readout of any individual's DNA.

Does this change medicine yet?

Slowly, at the level of diagnostics. Reference genomes are the coordinate systems against which clinical sequencing is aligned; adding repeat-rich regions means variants there can now, in principle, be discovered. Consortium papers pointed to immediate candidates: expansions in satellite DNA associated with some inherited diseases, and duplication-spanning deletions relevant to developmental disorders.

But clinical adoption requires validated assays, not just a longer reference, and laboratories move cautiously with newly mappable territory because the same repeats that were invisible can also generate false calls when handled by untested pipelines. The honest summary is that the T2T genome widened the lens; reading clearly through it is proceeding study by study.

How does T2T change how new genomes are built?

The reference that genomics laboratories align against is no longer a single linear sequence by necessity. Since 2023, the Human Pangenome Reference Consortium has assembled dozens of genomes with T2T-era long reads, so that each individual's variation appears as its own structure rather than as deviations from one mostly European yardstick. The practical effect, described in the consortium's Nature publications, is that alignment misses near structurally variable regions drop sharply when a genome is compared against many references instead of one.

Sequencing vendors have moved in parallel. Long-read instruments now sit in a growing number of clinical and research laboratories, and consortium methods have been folded into standard pipelines for assembling and quality-checking new genomes. What began as a gap-filling exercise has become the default way a complete genome is produced at all.

Why 'telomere to telomere'?

Telomeres are the protective caps at chromosome ends. Sequencing from one cap to the other — telomere to telomere — is the literal definition of a gapless chromosome. The name is also a reminder of what the 2003 announcement quietly meant: complete as far as the instruments could see. The 2022 sequence did not change the human genome; it changed how much of it science can actually examine.

Frequently Asked Questions

What is the telomere-to-telomere genome?
It is the first gapless sequence of a human genome, published in Science in 2022 by the T2T consortium. Built from the CHM13 cell line with long-read sequencing, it added roughly 200 million base pairs of repeat-rich DNA — about eight percent — that the 2003 Human Genome Project could not resolve.
Wasn't the human genome already completed in 2003?
The 2003 draft covered about 92 percent, leaving centromeres, satellite arrays, and other highly repetitive tracts unassembled. Those regions were invisible to short-read technology, not considered unimportant. The 2022 T2T sequence closed them, and a complete Y chromosome followed in 2023.
Are there new genes in the added sequence?
Annotations of the added DNA predict roughly 2,000 gene features, mostly non-coding, with on the order of one hundred to 115 candidate protein-coding genes. Many sit in duplicon-rich regions near known disease variants, which is part of why clinicians are interested in them.
Does the T2T reference represent everyone?
No. CHM13 is a single haplotype of largely European ancestry. It serves as a structural reference; diversity is addressed by the human pangenome, a draft of 94 ancestry-diverse assemblies published in Nature in 2024 and still being extended.