Tandem repeats are the unmeasured variable in organoid donor lines
Tandem repeats make up nearly 8% of the human genome and mutate at rates orders of magnitude above single-letter variants. A comparative catalog across seven ape genomes shows they are under strong constraint exactly where gene regulation lives, and that they concentrate in nervous-system and synaptic genes. The uncomfortable implication for organoid science: the fastest-moving part of the donor genome is the part nobody measures when a cell line is qualified.
Source: Comparative genomics of Tandem Repeat variation in apes, bioRxiv, 2026. Primary source. Read: full text retrieved from bioRxiv on the run date; every number below was checked against the retrieved text.
What the work claims
This is a primary comparative-genomics result, not a model paper, and it should be weighted as such. de Lima Adam, Rocha, Sudmant and Rohlfs built a population-aware catalog of tandem repeats across the telomere-to-telomere reference genomes of seven ape species, then genotyped long-read sequencing data from 46 humans and 23 chimpanzees to measure within-species diversity and between-species divergence at the same loci. They identified over 3 million tandem repeat loci per ape genome and 1,905,903 homologous loci shared between humans and chimpanzees, of which 872,237 (45.7%) were variable rather than invariant1.
The central claims are three. First, repeat diversity and conservation are strongly structured by genomic context: coding and untranslated regions show reduced polymorphism within species and reduced divergence between species, consistent with stabilizing selection, while intronic and intergenic repeats are far more dynamic. Second, a divergence-diversity ratio framework, which flags loci whose between-species divergence is extreme relative to their within-species diversity, recovers repeats concentrated in genes for nervous system development, synaptic organization and cell signaling, in both the high-divergence (directional selection) and high-diversity (balancing selection) tails. Third, repeat-length divergence tracks gene-expression divergence weakly but significantly across thousands of orthologous genes, and differentially expressed genes are preferentially enriched for repeats with prior evidence of regulatory function.
How it works
A tandem repeat is a short DNA sequence, from a single nucleotide to motifs of dozens of base pairs, repeated head to tail. Replication slippage makes these loci mutate at rates orders of magnitude higher than single-nucleotide variants, which is why they are the basis of forensic DNA fingerprinting and the cause of repeat-expansion disorders. Their length can modulate transcription in a non-linear, length-dependent way, the tuning-knob model: a promoter or 5 prime UTR repeat can act as a quantitative dial on expression of its gene.
The paper's most decision-relevant numbers are about constraint. Of the shared human-chimpanzee repeats, 89.7% of coding-region repeats and 74.2% of 5 prime UTR repeats were invariant, against roughly 55% of 3 prime UTR and 50% of intronic repeats. Where repeats do sit in coding sequence, frame-preserving motifs dominate: trinucleotide repeats are enriched 5.2-fold and hexanucleotide repeats 2.7-fold, while repeat-destroying mono- and dinucleotide repeats are depleted 100-fold and 20-fold respectively. In 5 prime UTRs, trinucleotide repeats are enriched 4.3-fold. In plain terms: selection keeps repeat length stable precisely in the regions that control translation and protein sequence, and tolerates variation elsewhere. The regulatory compartment of the genome is also where the enrichment sits, which tells you these elements are doing something selection cares about.
On expression, the authors analyzed 7,050 orthologous genes across six tissues (brain, cerebellum, heart, kidney, liver, testis). Average divergence of repeat-allele lengths correlated with gene-expression divergence at Pearson r = 0.02 (FDR-adjusted p = 2.2e-16): significant because the dataset is enormous, but tiny in variance explained. The locus-level picture is sharper. Of 4,125 genes differentially expressed between humans and chimpanzees, those carrying expression-associated repeats were significantly more likely to be differentially expressed (odds ratio 1.45, p = 0.012), even though repeat-carrying genes as a class were not overrepresented among differentially expressed genes (odds ratio 0.84, p = 0.987). The signal is not repeats in general; it is repeats already known to regulate expression1.
Where a skeptic should push
The most load-bearing assumption in moving from this paper to any experimental field is that a genome-wide average of r = 0.02 licenses concern about individual loci. A correlation that small explains about 0.04% of variance across all genes, and the authors are honest that the effect is heterogeneous and context-dependent. The defense is that the genome-wide mean is the wrong unit: fine-mapped expression-associated repeats and yeast experiments both show single loci where repeat length changes expression substantially and non-linearly. An average across roughly 1.9 million loci, most of which are inert, will always dilute the loci that matter. But it does mean the paper cannot tell you which repeats matter for which gene in which cell type. It identifies a class of plausible regulatory variants, not a list of confirmed effectors.
Two further limits deserve weight. The expression data are bulk RNA-seq from six organs, originally generated for cross-species comparison, not single-cell data and not developmental time courses. And although the catalog spans seven ape genomes, the population genotyping that powers the diversity and divergence-diversity analyses covers only humans and chimpanzees; the authors themselves flag that everything about selection regimes in the other apes is inference from single reference genomes. For the field that reads this paper, the correct posture is: treat repeat length as a candidate regulatory variable of unknown effect size per locus, not as a proven driver of any given phenotype.
Tandem repeats as an organoid donor-line confound
Here is the implication the paper never states. Every organoid inherits one diploid set of tandem-repeat alleles from its donor, and standard iPSC and organoid quality control does not measure them. A typical qualification panel covers karyotype, pluripotency markers, mycoplasma and a SNP fingerprint for identity. SNP arrays and short-read sequencing handle single-nucleotide variation well and tandem repeats poorly; the repetitive sequence that makes these loci mutate fast also makes them the hardest class of variant to genotype with the tools most labs actually use. So a donor panel of ten lines, presented as biological replication, is ten uncharacterized repeat-length genotypes. The between-donor differences that organoid studies routinely attribute to genetic background are, at least in part, differences in a variant class nobody has typed.
Why this matters more for organoids than for most model systems is where the repeats concentrate. The enrichment the paper reports for nervous-system development, synaptic organization and signaling genes maps almost exactly onto the readouts that neural organoid screens depend on: differentiation trajectories, synaptic marker expression, network maturation, responses to neuroactive compounds. A repeat in the 5 prime UTR of a synaptic gene, differing in length between two donor lines, is precisely the kind of variant that could shift a differentiation protocol's output or a drug-response curve while leaving the karyotype pristine and the SNP panel indistinguishable. Repeat-expansion disorders make the same point from the disease side: a neuronal organoid model of such a disorder assumes the repeat length is the modeled variable, but high slippage mutation rates raise the question, flagged here as an open one, of whether repeat length stays stable through reprogramming and extended culture, or whether it drifts and silently changes the model.
The opportunity is concrete and cheap relative to the cost of a failed screen. Long-read sequencing of donor lines, run once and deposited as metadata, would turn an invisible variable into a recorded one; donors could be matched or stratified by repeat length at loci relevant to the assay. More ambitiously, natural repeat-length variation is a ready-made allele series: instead of engineering expression constructs, a lab could select donor lines that carry long and short alleles at a repeat of interest and read out the phenotypic dose-response. The threat is the mirror image. As organoid drug screens feed decisions with real money attached, an uncharacterized donor-level variable in regulatory regions of neural genes is a reproducibility liability waiting to be discovered by someone else's failed replication, and a ready-made explanation for exactly the kind of lab-to-lab discrepancy the field is trying to shed.
The bottom line
Established by this paper: tandem repeats are pervasive (nearly 8% of the genome), extremely mutable, under measurable stabilizing selection in coding sequence and 5 prime UTRs, enriched in the regulatory regions of nervous-system and synaptic genes, and weakly but significantly associated with gene-expression divergence, with the association concentrated in repeats already known to regulate expression. Hypothesis, not established: that repeat-length differences between organoid donor lines measurably confound differentiation and drug-response readouts. What would confirm it is straightforward to describe and rarely done: long-read repeat genotyping of donor panels and their derived organoids, longitudinal checks on repeat stability through reprogramming and culture, and association of specific repeat alleles with specific organoid phenotypes. What would weaken it is evidence that repeat lengths are stable in culture and phenotypically silent across the donor range used in practice. Until one of those is in hand, repeat length belongs on the list of things a serious organoid study reports, alongside passage number and cell line identity.
Frequently asked questions
What is a tandem repeat?
A short DNA sequence repeated consecutively many times, from single-nucleotide runs to motifs of dozens of base pairs. Roughly 8% of the human genome consists of such repeats, and their copy number varies between individuals far more mutably than single-letter variants.
What did the study measure?
Over 3 million tandem repeat loci in each of seven telomere-to-telomere ape reference genomes, plus long-read genotyping of 46 humans and 23 chimpanzees at 1,905,903 homologous loci, to quantify within-species diversity, between-species divergence, and signatures of selection on repeat length.
Where is repeat length constrained?
In coding sequence and 5 prime UTRs: 89.7% of coding repeats and 74.2% of 5 prime UTR repeats were invariant between the species, versus roughly half of intronic repeats. Frame-preserving repeat motifs (multiples of three base pairs) are strongly enriched in coding regions.
Does the weak correlation with expression mean repeats do not matter?
No. The genome-wide correlation (r = 0.02) averages nearly 2 million loci, most inert. Genes differentially expressed between humans and chimpanzees were significantly more likely to carry repeats with prior evidence of regulatory function (odds ratio 1.45), and known regulatory repeats can have large, non-linear effects at individual loci.
Was any of this measured in organoids?
No, and a correction is worth making explicit: one library summary of this paper described the expression correlation as measured in neurodevelopmental organoids. The paper's methods describe bulk RNA-seq from six tissues in humans and chimpanzees. The organoid connection in this analysis is an expert inference from the genomic results, not a finding of the paper.
What should organoid labs do differently?
As a recommendation rather than a paper finding: add long-read repeat genotyping to donor-line characterization, record repeat length at loci relevant to the assay, and check stability of disease-model repeats through reprogramming and extended culture. None of this is standard practice today.
References
- de Lima Adam C, Rocha JL, Sudmant PH, Rohlfs R. Comparative genomics of Tandem Repeat variation in apes. bioRxiv. 2026. doi:10.64898/2026.01.20.700717. https://www.biorxiv.org/content/10.64898/2026.01.20.700717. Accessed 2026-09-10.