NobleBlocks

State Key Laboratory of Genetic Resources and Evolution

facilityKunming, China

Research output, citation impact, and the most-cited recent papers from State Key Laboratory of Genetic Resources and Evolution. Aggregated across the NobleBlocks index of 300M+ scholarly works.

Total works
3.2K
Citations
426.8K
h-index
274
i10-index
5.5K
Also known as
Key Laboratory of Genetic Resources and EvolutionState Key Lab of Genetic Resources & EvolutionState Key Laboratory of Genetic Resources and Evolution遗传资源与进化国家重点实验室

Top-cited papers from State Key Laboratory of Genetic Resources and Evolution

Towards complete and error-free genome assemblies of all vertebrate species
Arang Rhie, Shane McCarthy, Olivier D. Fedrigo, Joana Damas +4 more
2021· Nature3.3Kdoi:10.1038/s41586-021-03451-0

Abstract High-quality and complete reference genome assemblies are fundamental for the application of genomics to biology, disease, and biodiversity conservation. However, such assemblies are available for only a few non-microbial species 1–4 . To address this issue, the international Genome 10K (G10K) consortium 5,6 has worked over a five-year period to evaluate and develop cost-effective methods for assembling highly accurate and nearly complete reference genomes. Here we present lessons learned from generating assemblies for 16 species that represent six major vertebrate lineages. We confirm that long-read sequencing technologies are essential for maximizing genome quality, and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Our assemblies correct substantial errors, add missing sequence in some of the best historical reference genomes, and reveal biological discoveries. These include the identification of many false gene duplications, increases in gene sizes, chromosome rearrangements that are specific to lineages, a repeated independent chromosome breakpoint in bat genomes, and a canonical GC-rich pattern in protein-coding genes and their regulatory regions. Adopting these lessons, we have embarked on the Vertebrate Genomes Project (VGP), an international effort to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species and to help to enable a new era of discovery across the life sciences.

The sequence and de novo assembly of the giant panda genome
Ruiqiang Li, Wei Fan, Geng Tian, Hongmei Zhu +4 more
2009· Nature1.2Kdoi:10.1038/nature08696

Using next-generation sequencing technology alone, we have successfully generated and assembled a draft sequence of the giant panda genome. The assembled contigs (2.25 gigabases (Gb)) cover approximately 94% of the whole genome, and the remaining gaps (0.05 Gb) seem to contain carnivore-specific repeats and tandem repeats. Comparisons with the dog and human showed that the panda genome has a lower divergence rate. The assessment of panda genes potentially underlying some of its unique traits indicated that its bamboo diet might be more dependent on its gut microbiome than its own genetic composition. We also identified more than 2.7 million heterozygous single nucleotide polymorphisms in the diploid genome. Our data and analyses provide a foundation for promoting mammalian genetic research, and demonstrate the feasibility for using next-generation sequencing technologies for accurate, cost-effective and rapid de novo assembly of large eukaryotic genomes. The genome of the giant panda — specifically of the female Beijing Olympics mascot Jingjing — has been determined using short-read sequencing technology, a first for such a complex genome. It consists of some 2.4 billion DNA base pairs, compared to 3 billion in humans, and contains around 21,000 protein-encoding genes, similar to the human genome. Genomic diversity reflected in the sequence is high, raising hopes that despite a population of only about 2,500, conservation efforts can keep the species from extinction. Intriguingly, the panda appears to have all the genes needed for a carnivorous digestive system but lacks digestive cellulase genes. It may therefore depend on its gut microbiome to handle its famously limited bamboo diet. Taste may be a diet-limiting factor: loss of function of the T1R1 gene means that pandas may not experience the umami taste associated with high-protein foods. Technical aspects of this work pave the way for the use of next-generation sequencing for rapid de novo assembly of large eukaryotic genomes. Here, a draft sequence of the giant panda genome is assembled using next-generation sequencing technology alone. Genome analysis reveals a low divergence rate in comparison with dog and human genomes and insights into panda-specific traits; for example, the giant panda's bamboo diet may be more dependent on its gut microbiome than its own genetic composition.

Resequencing 302 wild and cultivated accessions identifies genes related to domestication and improvement in soybean
Zhengkui Zhou, Yu Jiang, Zheng Wang, Zhiheng Gou +4 more
2015· Nature Biotechnology1.2Kdoi:10.1038/nbt.3096

Understanding soybean (Glycine max) domestication and improvement at a genetic level is important to inform future efforts to further improve a crop that provides the world's main source of oilseed. We detect 230 selective sweeps and 162 selected copy number variants by analysis of 302 resequenced wild, landrace and improved soybean accessions at >11× depth. A genome-wide association study using these new sequences reveals associations between 10 selected regions and 9 domestication or improvement traits, and identifies 13 previously uncharacterized loci for agronomic traits including oil content, plant height and pubescence form. Combined with previous quantitative trait loci (QTL) information, we find that, of the 230 selected regions, 96 correlate with reported oil QTLs and 21 contain fatty acid biosynthesis genes. Moreover, we observe that some traits and loci are associated with geographical regions, which shows that soybean populations are structured geographically. This study provides resources for genomics-enabled improvements in soybean breeding.

Earth BioGenome Project: Sequencing life for the future of life
Harris A. Lewin, Gene E. Robinson, W. John Kress, William J. Baker +4 more
2018· Proceedings of the National Academy of Sciences1.1Kdoi:10.1073/pnas.1720115115

Increasing our understanding of Earth's biodiversity and responsibly stewarding its resources are among the most crucial scientific and social challenges of the new millennium. These challenges require fundamental new knowledge of the organization, evolution, functions, and interactions among millions of the planet's organisms. Herein, we present a perspective on the Earth BioGenome Project (EBP), a moonshot for biology that aims to sequence, catalog, and characterize the genomes of all of Earth's eukaryotic biodiversity over a period of 10 years. The outcomes of the EBP will inform a broad range of major issues facing humanity, such as the impact of climate change on biodiversity, the conservation of endangered species and ecosystems, and the preservation and enhancement of ecosystem services. We describe hurdles that the project faces, including data-sharing policies that ensure a permanent, freely available resource for future scientific discovery while respecting access and benefit sharing guidelines of the Nagoya Protocol. We also describe scientific and organizational challenges in executing such an ambitious project, and the structure proposed to achieve the project's goals. The far-reaching potential benefits of creating an open digital repository of genomic information for life on Earth can be realized only by a coordinated international effort.

The yak genome and adaptation to life at high altitude
Qiang Qiu, Guojie Zhang, Tao Ma, Wubin Qian +4 more
2012· Nature Genetics1.0Kdoi:10.1038/ng.2343

Domestic yaks (Bos grunniens) provide meat and other necessities for Tibetans living at high altitude on the Qinghai-Tibetan Plateau and in adjacent regions. Comparison between yak and the closely related low-altitude cattle (Bos taurus) is informative in studying animal adaptation to high altitude. Here, we present the draft genome sequence of a female domestic yak generated using Illumina-based technology at 65-fold coverage. Genomic comparisons between yak and cattle identify an expansion in yak of gene families related to sensory perception and energy metabolism, as well as an enrichment of protein domains involved in sensing the extracellular environment and hypoxic stress. Positively selected and rapidly evolving genes in the yak lineage are also found to be significantly enriched in functional categories and pathways related to hypoxia and nutrition metabolism. These findings may have important implications for understanding adaptation to high altitude in other animal species and for hypoxia-related diseases in humans.

PLEK: a tool for predicting long non-coding RNAs and messenger RNAs based on an improved k-mer scheme
Aimin Li, Junying Zhang, Zhongyin Zhou
2014· BMC Bioinformatics879doi:10.1186/1471-2105-15-311

BACKGROUND: High-throughput transcriptome sequencing (RNA-seq) technology promises to discover novel protein-coding and non-coding transcripts, particularly the identification of long non-coding RNAs (lncRNAs) from de novo sequencing data. This requires tools that are not restricted by prior gene annotations, genomic sequences and high-quality sequencing. RESULTS: We present an alignment-free tool called PLEK (predictor of long non-coding RNAs and messenger RNAs based on an improved k-mer scheme), which uses a computational pipeline based on an improved k-mer scheme and a support vector machine (SVM) algorithm to distinguish lncRNAs from messenger RNAs (mRNAs), in the absence of genomic sequences or annotations. The performance of PLEK was evaluated on well-annotated mRNA and lncRNA transcripts. 10-fold cross-validation tests on human RefSeq mRNAs and GENCODE lncRNAs indicated that our tool could achieve accuracy of up to 95.6%. We demonstrated the utility of PLEK on transcripts from other vertebrates using the model built from human datasets. PLEK attained >90% accuracy on most of these datasets. PLEK also performed well using a simulated dataset and two real de novo assembled transcriptome datasets (sequenced by PacBio and 454 platforms) with relatively high indel sequencing errors. In addition, PLEK is approximately eightfold faster than a newly developed alignment-free tool, named Coding-Non-Coding Index (CNCI), and 244 times faster than the most popular alignment-based tool, Coding Potential Calculator (CPC), in a single-threading running manner. CONCLUSIONS: PLEK is an efficient alignment-free computational tool to distinguish lncRNAs from mRNAs in RNA-seq transcriptomes of species lacking reference genomes. PLEK is especially suitable for PacBio or 454 sequencing data and large-scale transcriptome data. Its open-source software can be freely downloaded from https://sourceforge.net/projects/plek/files/.

Whole-genome sequence of a flatfish provides insights into ZW sex chromosome evolution and adaptation to a benthic lifestyle
Songlin Chen, Guojie Zhang, Changwei Shao, Quanfei Huang +4 more
2014· Nature Genetics870doi:10.1038/ng.2890

Songlin Chen and colleagues sequenced the whole genomes of a male (ZZ) and a female (ZW) Chinese half-smooth tongue sole, Cynoglossus semilaevis. Their analysis provides insights into the structure and evolution of the sex chromosomes and adaptation to the benthic lifestyle of this flatfish. Genetic sex determination by W and Z chromosomes has developed independently in different groups of organisms. To better understand the evolution of sex chromosomes and the plasticity of sex-determination mechanisms, we sequenced the whole genomes of a male (ZZ) and a female (ZW) half-smooth tongue sole (Cynoglossus semilaevis). In addition to insights into adaptation to a benthic lifestyle, we find that the sex chromosomes of these fish are derived from the same ancestral vertebrate protochromosome as the avian W and Z chromosomes. Notably, the same gene on the Z chromosome, dmrt1, which is the male-determining gene in birds, showed convergent evolution of features that are compatible with a similar function in tongue sole. Comparison of the relatively young tongue sole sex chromosomes with those of mammals and birds identified events that occurred during the early phase of sex-chromosome evolution. Pertinent to the current debate about heterogametic sex-chromosome decay, we find that massive gene loss occurred in the wake of sex-chromosome 'birth'.

A Global Deal For Nature: Guiding principles, milestones, and targets
Eric Dinerstein, Carly Vynne, Enric Sala, Anup R. Joshi +4 more
2019· Science Advances812doi:10.1126/sciadv.aaw2869

The Global Deal for Nature (GDN) is a time-bound, science-driven plan to save the diversity and abundance of life on Earth. Pairing the GDN and the Paris Climate Agreement would avoid catastrophic climate change, conserve species, and secure essential ecosystem services. New findings give urgency to this union: Less than half of the terrestrial realm is intact, yet conserving all native ecosystems-coupled with energy transition measures-will be required to remain below a 1.5°C rise in average global temperature. The GDN targets 30% of Earth to be formally protected and an additional 20% designated as climate stabilization areas, by 2030, to stay below 1.5°C. We highlight the 67% of terrestrial ecoregions that can meet 30% protection, thereby reducing extinction threats and carbon emissions from natural reservoirs. Freshwater and marine targets included here extend the GDN to all realms and provide a pathway to ensuring a more livable biosphere.

Biodiversity soup: metabarcoding of arthropods for rapid biodiversity assessment and biomonitoring
Douglas W. Yu, Yinqiu Ji, Brent C. Emerson, Xiaoyang Wang +3 more
2012· Methods in Ecology and Evolution789doi:10.1111/j.2041-210x.2012.00198.x

Summary 1. Traditional biodiversity assessment is costly in time, money and taxonomic expertise. Moreover, data are frequently collected in ways (e.g. visual bird lists) that are unsuitable for auditing by neutral parties, which is necessary for dispute resolution. 2. We present protocols for the extraction of ecological, taxonomic and phylogenetic information from bulk samples of arthropods. The protocols combine mass trapping of arthropods, mass‐PCR amplification of the COI barcode gene, pyrosequencing and bioinformatic analysis, which together we call ‘metabarcoding’. 3. We construct seven communities of arthropods (mostly insects) and show that it is possible to recover a substantial proportion of the original taxonomic information. We further demonstrate, for the first time, that metabarcoding allows for the precise estimation of pairwise community dissimilarity (beta diversity) and within‐community phylogenetic diversity (alpha diversity), despite the inevitable loss of taxonomic information inherent to metabarcoding. 4. Alpha and beta diversity metrics are the raw materials of ecology and the environmental sciences, facilitating assessment of the state of the environment with a broad and efficient measure of biodiversity.

Reliable, verifiable and efficient monitoring of biodiversity via metabarcoding
Yinqiu Ji, Louise Amy Ashton, Scott M. Pedley, David P. Edwards +4 more
2013· Ecology Letters698doi:10.1111/ele.12162

To manage and conserve biodiversity, one must know what is being lost, where, and why, as well as which remedies are likely to be most effective. Metabarcoding technology can characterise the species compositions of mass samples of eukaryotes or of environmental DNA. Here, we validate metabarcoding by testing it against three high-quality standard data sets that were collected in Malaysia (tropical), China (subtropical) and the United Kingdom (temperate) and that comprised 55,813 arthropod and bird specimens identified to species level with the expenditure of 2,505 person-hours of taxonomic expertise. The metabarcode and standard data sets exhibit statistically correlated alpha- and beta-diversities, and the two data sets produce similar policy conclusions for two conservation applications: restoration ecology and systematic conservation planning. Compared with standard biodiversity data sets, metabarcoded samples are taxonomically more comprehensive, many times quicker to produce, less reliant on taxonomic expertise and auditable by third parties, which is essential for dispute resolution.

Progressive Cactus is a multiple-genome aligner for the thousand-genome era
Joel C. Armstrong, Glenn Hickey, Mark E. Diekhans, Ian T. Fiddes +4 more
2020· Nature688doi:10.1038/s41586-020-2871-y

Abstract New genome assemblies have been arriving at a rapidly increasing pace, thanks to decreases in sequencing costs and improvements in third-generation sequencing technologies 1–3 . For example, the number of vertebrate genome assemblies currently in the NCBI (National Center for Biotechnology Information) database 4 increased by more than 50% to 1,485 assemblies in the year from July 2018 to July 2019. In addition to this influx of assemblies from different species, new human de novo assemblies 5 are being produced, which enable the analysis of not only small polymorphisms, but also complex, large-scale structural differences between human individuals and haplotypes. This coming era and its unprecedented amount of data offer the opportunity to uncover many insights into genome evolution but also present challenges in how to adapt current analysis methods to meet the increased scale. Cactus 6 , a reference-free multiple genome alignment program, has been shown to be highly accurate, but the existing implementation scales poorly with increasing numbers of genomes, and struggles in regions of highly duplicated sequences. Here we describe progressive extensions to Cactus to create Progressive Cactus, which enables the reference-free alignment of tens to thousands of large vertebrate genomes while maintaining high alignment quality. We describe results from an alignment of more than 600 amniote genomes, which is to our knowledge the largest multiple vertebrate genome alignment created so far.

Revealing the History of Sheep Domestication Using Retrovirus Integrations
Bernardo Chessa, Filipe de Souza Pereira, Frédérick Arnaud, António Amorim +4 more
2009· Science585doi:10.1126/science.1170587

The domestication of livestock represented a crucial step in human history. By using endogenous retroviruses as genetic markers, we found that sheep differentiated on the basis of their "retrotype" and morphological traits dispersed across Eurasia and Africa via separate migratory episodes. Relicts of the first migrations include the Mouflon, as well as breeds previously recognized as "primitive" on the basis of their morphology, such as the Orkney, Soay, and the Nordic short-tailed sheep now confined to the periphery of northwest Europe. A later migratory episode, involving sheep with improved production traits, shaped the great majority of present-day breeds. The ability to differentiate genetically primitive sheep from more modern breeds provides valuable insights into the history of sheep domestication.

The sheep genome illuminates biology of the rumen and lipid metabolism
Yu Jiang, Min Xie, Wenbin Chen, Richard Talbot +4 more
2014· Science564doi:10.1126/science.1252806

Sheep (Ovis aries) are a major source of meat, milk, and fiber in the form of wool and represent a distinct class of animals that have a specialized digestive organ, the rumen, that carries out the initial digestion of plant material. We have developed and analyzed a high-quality reference sheep genome and transcriptomes from 40 different tissues. We identified highly expressed genes encoding keratin cross-linking proteins associated with rumen evolution. We also identified genes involved in lipid metabolism that had been amplified and/or had altered tissue expression patterns. This may be in response to changes in the barrier lipids of the skin, an interaction between lipid metabolism and wool synthesis, and an increased role of volatile fatty acids in ruminants compared with nonruminant animals.

Dense sampling of bird diversity increases power of comparative genomics
Shaohong Feng, Josefin Stiller, Yuan Deng, Joel C. Armstrong +4 more
2020· Nature553doi:10.1038/s41586-020-2873-9

Whole-genome sequencing projects are increasingly populating the tree of life and characterizing biodiversity1–4. Sparse taxon sampling has previously been proposed to confound phylogenetic inference5, and captures only a fraction of the genomic diversity. Here we report a substantial step towards the dense representation of avian phylogenetic and molecular diversity, by analysing 363 genomes from 92.4% of bird families—including 267 newly sequenced genomes produced for phase II of the Bird 10,000 Genomes (B10K) Project. We use this comparative genome dataset in combination with a pipeline that leverages a reference-free whole-genome alignment to identify orthologous regions in greater numbers than has previously been possible and to recognize genomic novelties in particular bird lineages. The densely sampled alignment provides a single-base-pair map of selection, has more than doubled the fraction of bases that are confidently predicted to be under conservation and reveals extensive patterns of weak selection in predominantly non-coding DNA. Our results demonstrate that increasing the diversity of genomes used in comparative studies can reveal more shared and lineage-specific variation, and improve the investigation of genomic characteristics. We anticipate that this genomic resource will offer new perspectives on evolutionary processes in cross-species comparative analyses and assist in efforts to conserve species. A dataset of the genomes of 363 species from the Bird 10,000 Genomes Project shows increased power to detect shared and lineage-specific variation, demonstrating the importance of phylogenetically diverse taxon sampling in whole-genome sequencing.

Sequencing and automated whole-genome optical mapping of the genome of a domestic goat (Capra hircus)
Yang Dong, Min Xie, Yu Jiang, Nianqing Xiao +4 more
2012· Nature Biotechnology552doi:10.1038/nbt.2478

We report the ∼2.66-Gb genome sequence of a female Yunnan black goat. The sequence was obtained by combining short-read sequencing data and optical mapping data from a high-throughput whole-genome mapping instrument. The whole-genome mapping data facilitated the assembly of super-scaffolds >5× longer by the N50 metric than scaffolds augmented by fosmid end sequencing (scaffold N50 = 3.06 Mb, super-scaffold N50 = 16.3 Mb). Super-scaffolds are anchored on chromosomes based on conserved synteny with cattle, and the assembly is well supported by two radiation hybrid maps of chromosome 1. We annotate 22,175 protein-coding genes, most of which were recovered in the RNA-seq data of ten tissues. Comparative transcriptomic analysis of the primary and secondary follicles of a cashmere goat reveal 51 genes that are differentially expressed between the two types of hair follicles. This study, whose results will facilitate goat genomics, shows that whole-genome mapping technology can be used for the de novo assembly of large genomes.

Deep RNA sequencing at single base-pair resolution reveals high complexity of the rice transcriptome
Guojie Zhang, Guangwu Guo, Xueda Hu, Yong Zhang +4 more
2010· Genome Research510doi:10.1101/gr.100677.109

Understanding the dynamics of eukaryotic transcriptome is essential for studying the complexity of transcriptional regulation and its impact on phenotype. However, comprehensive studies of transcriptomes at single base resolution are rare, even for modern organisms, and lacking for rice. Here, we present the first transcriptome atlas for eight organs of cultivated rice. Using high-throughput paired-end RNA-seq, we unambiguously detected transcripts expressing at an extremely low level, as well as a substantial number of novel transcripts, exons, and untranslated regions. An analysis of alternative splicing in the rice transcriptome revealed that alternative cis-splicing occurred in approximately 33% of all rice genes. This is far more than previously reported. In addition, we also identified 234 putative chimeric transcripts that seem to be produced by trans-splicing, indicating that transcript fusion events are more common than expected. In-depth analysis revealed a multitude of fusion transcripts that might be by-products of alternative splicing. Validation and chimeric transcript structural analysis provided evidence that some of these transcripts are likely to be functional in the cell. Taken together, our data provide extensive evidence that transcriptional regulation in rice is vastly more complex than previously believed.

Allele-aware chromosome-level genome assembly and efficient transgene-free genome editing for the autotetraploid cultivated alfalfa
Haitao Chen, Yan Zeng, Yongzhi Peter Yang, Lingli Huang +4 more
2020· Nature Communications508doi:10.1038/s41467-020-16338-x

Artificially improving traits of cultivated alfalfa (Medicago sativa L.), one of the most important forage crops, is challenging due to the lack of a reference genome and an efficient genome editing protocol, which mainly result from its autotetraploidy and self-incompatibility. Here, we generate an allele-aware chromosome-level genome assembly for the cultivated alfalfa consisting of 32 allelic chromosomes by integrating high-fidelity single-molecule sequencing and Hi-C data. We further establish an efficient CRISPR/Cas9-based genome editing protocol on the basis of this genome assembly and precisely introduce tetra-allelic mutations into null mutants that display obvious phenotype changes. The mutated alleles and phenotypes of null mutants can be stably inherited in generations in a transgene-free manner by cross pollination, which may help in bypassing the debate about transgenic plants. The presented genome and CRISPR/Cas9-based transgene-free genome editing protocol provide key foundations for accelerating research and molecular breeding of this important forage crop.

Large-scale ruminant genome sequencing provides insights into their evolution and distinct traits
Lei Chen, Qiang Qiu, Yu Jiang, Kun Wang +4 more
2019· Science494doi:10.1126/science.aav6202

Phylogeny and characteristics of ruminants Ruminants are a diverse group of mammals that includes families containing well-known taxa such as deer, cows, and goats. However, their evolutionary relationships have been contentious, as have the origins of their distinctive digestive systems and headgear, including antlers and horns (see the Perspective by Ker and Yang). To understand the relationships among ruminants, L. Chen et al. sequenced 44 species representing 6 families and performed a phylogenetic analysis. From this analysis, they were able to resolve the phylogeny of many genera and document incomplete lineage sorting among major clades. Interestingly, they found evidence for large population reductions among many taxa starting at approximately 100,000 years ago, coinciding with the migration of humans out of Africa. Examining the bony appendages on the head—the so-called headgear—Wang et al. describe specific evolutionary changes in the ruminants and identify selection on cancer-related genes that may function in antler development in deer. Finally, Lin et al. take a close look at the reindeer genome and identify the genetic basis of adaptations that allow reindeer to survive in the harsh conditions of the Arctic. Science , this issue p. eaav6202 , p. eaav6335 , p. eaav6312 ; see also p. 1130

Genomic Comparison of the Ants Camponotus floridanus and Harpegnathos saltator
Roberto Bonasio, Guojie Zhang, Chaoyang Ye, Navdeep S. Mutti +4 more
2010· Science472doi:10.1126/science.1192428

The organized societies of ants include short-lived worker castes displaying specialized behavior and morphology and long-lived queens dedicated to reproduction. We sequenced and compared the genomes of two socially divergent ant species: Camponotus floridanus and Harpegnathos saltator. Both genomes contained high amounts of CpG, despite the presence of DNA methylation, which in non-Hymenoptera correlates with CpG depletion. Comparison of gene expression in different castes identified up-regulation of telomerase and sirtuin deacetylases in longer-lived H. saltator reproductives, caste-specific expression of microRNAs and SMYD histone methyltransferases, and differential regulation of genes implicated in neuronal function and chemical communication. Our findings provide clues on the molecular differences between castes in these two ants and establish a new experimental model to study epigenetics in aging and behavior.

Whole-genome resequencing reveals world-wide ancestry and adaptive introgression events of domesticated cattle in East Asia
Ningbo Chen, Yudong Cai, Qiuming Chen, Ran Li +4 more
2018· Nature Communications463doi:10.1038/s41467-018-04737-0

Cattle domestication and the complex histories of East Asian cattle breeds warrant further investigation. Through analysing the genomes of 49 modern breeds and eight East Asian ancient samples, worldwide cattle are consistently classified into five continental groups based on Y-chromosome haplotypes and autosomal variants. We find that East Asian cattle populations are mainly composed of three distinct ancestries, including an earlier East Asian taurine ancestry that reached China at least ~3.9 kya, a later introduced Eurasian taurine ancestry, and a novel Chinese indicine ancestry that diverged from Indian indicine approximately 36.6-49.6 kya. We also report historic introgression events that helped domestic cattle from southern China and the Tibetan Plateau achieve rapid adaptation by acquiring ~2.93% and ~1.22% of their genomes from banteng and yak, respectively. Our findings provide new insights into the evolutionary history of cattle and the importance of introgression in adaptation of cattle to new environmental challenges in East Asia.