A broader cannabis genome reveals hidden genetic diversity

Four haplotype-resolved genome assemblies and a reference-free 66-haplotype pangenome graph for Cannabis sativa.

Scientific data • • Highly Relevant
🤖

AI Summary

Researchers created four haplotype-resolved, chromosome-scale genome assemblies for Cannabis sativa and combined them with previously available genetic data. In total, the resulting reference-free pangenome graph represents 66 haplotypes and contains 14.87 million SNPs and 6.40 million indels, capturing genetic variation across cannabis rather than relying on a single reference genome.

The optimized graph-building approach reduced misleading links among repetitive DNA and produced a structure that more clearly reflects the organization of plant chromosomes. K-mer analysis also indicates that the cannabis pangenome is not yet complete: additional genotypes are needed, especially from the plant’s region of origin, which appears to be underrepresented. This work does not test cannabis effects in people or report direct benefits for consumers, but it provides a stronger foundation for future research into cannabis diversity, genetics, and breeding.

💡 Key Findings

1
The study produced four chromosome-scale, haplotype-resolved genome assemblies for Cannabis sativa.
High
90%
2
A reference-free pangenome graph combining 66 haplotypes captures broad cannabis genetic variation, including 14.87 million SNPs and 6.40 million indels.
High
90%
3
Optimizing the graph-building parameters reduced spurious connections among repetitive DNA and better represented the linear structure of plant chromosomes.
High
85%
4
K-mer analysis indicates that the cannabis pangenome remains incomplete and that more genotypes—particularly from the plant’s region of origin—are needed.
High
85%

📄 Original Abstract

We present 4 haplotype-resolved, chromosome scale diploid assemblies of cannabis, assembled from ONT R9.4.1 reads via a novel pipeline based on Hi-C phasing. These assemblies, while low in QV, offer contiguity and genic content comparable to recent HiFi assemblies. Along with a trio-binned assembly previously produced by us and 56 haplotypes selected from the Salk Institute Cannabis Pangenome project, we use these assemblies to create a reference-free pangenome graph. Within a total length of 6.48 Gb, it contains 162.14 M nodes, 228.27 M edges, 14.87 M SNPs, and 6.40 M indels. By optimizing parameters within the Pangenome Graph Builder (PGGB), we avoid many spurious connections among repeat elements, reduce processing time, and arrive at a data structure that visibly recapitulates the linear nature of plant chromosomes. Via k-mer analysis, we corroborate that more genotypes are needed to close the cannabis pangenome, and that, in particular, the region of origin likely remains undersampled.

Explore More Research

Stay informed about the latest cannabis science.

Your stash, decoded.