Why Rooibos Has Been So Hard to Sequence, and Why the Ministry Wants It Done
Rooibos carries a genome roughly a billion base pairs long, packed with the same tannin-like compounds that make its tea, and those compounds fight every attempt to read its DNA. Here is what scientists have found so far, and why a finished genome matters for a crop already running short on stress tolerance.
The rooibos in your cup depends on a crop that is already struggling, and the plant's own chemistry is slowing down the one project meant to help it. Commercial rooibos plants are living shorter, more stress-prone lives than growers want, and a university team in Cape Town has spent several years just working out how to read the plant's DNA cleanly enough to breed a hardier one. The catch is almost funny: the same leaf chemistry that gives rooibos its tea also sabotages the lab step every sequencing project depends on. Rooibos does not yet have a finished genome, and that gap is not for want of trying.
The problem starts with what rooibos is made of
Aspalathus linearis is an extraordinarily phenolic plant. Polyphenols, the same broad family of compounds behind the tea's antioxidant reputation, can make up close to 30 percent of the leaf's dry weight. Aspalathin alone accounts for up to 13.5 percent of it, a University of the Western Cape team reported in the journal Plants in 2022. Those compounds are wonderful in a cup. In a DNA extraction, they are a nuisance. Phenolics and polysaccharides bind to DNA as it comes out of the cell, fouling both its purity and the amount you can recover. Every standard plant DNA protocol the team tried needed reworking before it produced DNA clean enough to sequence. The method that finally worked reliably, a silica-membrane cleanup kit built for exactly this kind of contamination, only became part of their pipeline after testing several less successful alternatives first.
That single fact, that rooibos fights its own sequencing, is the reason a plant this economically important still lacks a public reference genome nearly two decades after tools like this became standard for other crops.
How big a genome, and why the number matters first
Before you can sequence a genome you need to know roughly how large a target you are aiming at. So a separate 2020 study measured it two ways. Flow cytometry, which estimates size from how much a cell's DNA absorbs a fluorescent dye, put the rooibos genome at 1.24 billion base pairs. A computational method called k-mer analysis, which infers genome size from patterns in short sequencing reads without ever assembling them, put it lower: 1.03 billion base pairs, the same University of the Western Cape group reported. The two figures do not agree exactly. The paper's own explanation is not a shrug: rooibos's abundant phenolics likely inflate the dye-based estimate, by binding sites the stain also targets. Either number lands rooibos in the same range as pea and narrow-leafed lupin, two other well-studied legumes. This is a solidly mid-sized genome by legume standards, not a small one.
More than half of it, the team estimated, is repetitive DNA: sequence that repeats itself across the genome rather than coding for anything, which is exactly the kind of terrain that trips up older, short-read sequencing methods. A repeat longer than your sequencing read is invisible as a repeat. It just looks like more of the same sequence, and the assembly software has no way to tell how many times it actually occurs.
Why long reads, and why a device the size of a phone charger
This is the reason the 2022 paper's real subject is not the genome itself but the sequencing method. Long-read sequencing, where a single read can span tens of thousands of base pairs instead of a few hundred, is what lets an assembly bridge a repetitive stretch instead of getting lost inside it. Oxford Nanopore's MinION delivers exactly that, and does it cheaply: the device works by pulling a DNA strand through a nanoscale pore under an electrical charge, and each base disrupts the current passing through in a characteristic way that the software reads back as a sequence. A MinION starter kit costs roughly a thousand US dollars, and the sequencer itself is small enough to run on a laptop's USB port, portable enough that Oxford Nanopore has, separately, flown one to the International Space Station.
Getting ultra-pure, high-molecular-weight DNA into that pore was the hard part for the Cape Town team, exactly because of the phenolics described above. Once they had it, they tested nine different assembly programs against the resulting long reads: Platanus, MaSuRCA, Haslr, Wengan, Flye, Canu, Raven, Redbean, and NextDenovo. No single program won on every measure. NextDenovo built the most contiguous assembly of the nine, anchored by the single longest continuous stretch any of them assembled, a contig running 3.4 million base pairs. But its total assembly size only reached 66 to 82 percent of the genome the team expected to find, meaning real sequence was still missing. Redbean ran a close second on contiguity, and did it for a fraction of the computing cost: about 32 hours on 205 CPU-hours and 54 gigabytes of memory, and after a polishing step recovered 99.2 percent of the complete, single-copy genes a benchmarking tool expects any healthy plant genome to have. Flye landed closest to the full predicted size, at 1.1 billion base pairs, with 97.3 percent of the benchmark genes complete. Canu, by contrast, needed 112 processor cores, roughly 3,000 gigabytes of memory, and 35,401 CPU-hours. That is an order of magnitude more computing than Redbean, for a result the smaller-footprint methods matched or beat. For a university lab without a supercomputer on hand, that gap between Canu and Redbean is the difference between a project that runs and one that does not.
None of this is a finished, published reference genome yet. It is the working-out of the method, the necessary first act, and the paper is honest that the assemblies it produced are drafts to be refined, not a settled final sequence.
The genome is only half the story: what the transcriptome adds
A genome tells you what genetic material a plant carries. It does not tell you which of those genes a rooibos plant actually switches on, in which tissue, and when. That second question is what a transcriptome answers, by sequencing RNA rather than DNA, and the same University of the Western Cape group has published two of those studies as well, six years apart.
The first, in 2020, sequenced rooibos across 44 individual wild plants representing five of the plant's eight known growth types (the Red type, the commercial ancestor, plus Black, Nardouwsberg, Grey sprouter, and Nieuwoudtville sprouter forms), and the chemistry difference between them is stark, the team reported. The commercial Red type produces roughly 66.5 milligrams of aspalathin per 100 grams of dry leaf. The wild growth types sampled ranged from 0 to about 9.24 milligrams per 100 grams, and some wild plants make essentially none at all, relying instead on two other compounds, orientin and rutin, as their main phenolics. A commercial crop that concentrates one compound to roughly seven times its wild relatives' upper range, while some of the same species makes almost none of it, is a genetic difference substantial enough to build a targeted breeding program around, once the underlying genes are known.
The second transcriptome study, published in Plants in May 2026, sequenced leaf and root tissue separately from a single six-month-old seedling using Oxford Nanopore's long-read method, Tanweer Beckett and Uljana Hesse reported. It found leaves running photosynthesis and carbon-fixation genes hard, unsurprisingly, and roots running a different program almost entirely: genes tied to hormone production, stress response, and defence, especially in the pathway that builds the plant hormone auxin, plus the biosynthesis routes for a broader set of protective secondary compounds. That leaf-versus-root split is not a curiosity. It is the kind of tissue-specific map a breeding program needs to know which genes to target and where.
What this means for the plant behind your cup
The 2026 paper states the practical stakes plainly: commercial rooibos plant longevity is declining, largely because of low stress tolerance, a shortage of resilient planting material behind it. A genetic improvement program for the crop already exists, aimed at exactly the traits a warming, drying Cederberg is putting under pressure: sturdier drought and pest tolerance, more consistent biomass, and steadier yields of the phytochemicals, aspalathin among them, that the whole industry is built on. Wild ecotype diversity, the kind the 2020 transcriptome study sampled across five growth types, is the raw material a breeding program draws from. A gene known to sit behind a wild plant's drought tolerance is only useful to a breeder once the genome and the transcriptome between them show which gene it is and where it sits.
Nobody buying a box of rooibos will ever read a BUSCO score or a k-mer plot. But the plant behind that box is the one this research is trying to keep alive and productive through a harder climate than it has faced before, and a finished reference genome is the tool a breeder needs to do that on purpose rather than by luck. That is the plain case for finishing the job the 2022 assembly paper began: not a finished genome for its own sake, but the reference map a crop under real pressure needs, drawn from a plant whose own chemistry has made it unusually stubborn to read.
Sources
- Establishing MinION Sequencing and Genome Assembly Procedures for the Analysis of the Rooibos (Aspalathus linearis) Genome, Mgwatyu, Hesse et al., Plants, 2022, on the DNA purification challenge, the nine assembly programs tested, and their comparative performance.
- Rooibos (Aspalathus linearis) Genome Size Estimation Using Flow Cytometry and K-Mer Analyses, Mgwatyu et al., Plants, 2020, on the genome size estimates and the repetitive-DNA fraction.
- Transcriptomics of the Rooibos (Aspalathus linearis) Species Complex, Stander, Hesse et al., BioTech, 2020, on the growth-type sampling and the aspalathin concentration differences between wild and commercial rooibos.
- Transcriptome Profiling of Leaves and Roots from Rooibos (Aspalathus linearis) Using Oxford Nanopore Sequencing, Beckett and Hesse, Plants, 2026, on the leaf/root gene-expression split and the stated breeding motivation.