Beijing Institute of Genomics
facilityBeijing, China
Research output, citation impact, and the most-cited recent papers from Beijing Institute of Genomics (China). Aggregated across the NobleBlocks index of 300M+ scholarly works.
Top-cited papers from Beijing Institute of Genomics
SUMMARY: The Sequence Alignment/Map (SAM) format is a generic alignment format for storing read alignments against reference sequences, supporting short and long reads (up to 128 Mbp) produced by different sequencing platforms. It is flexible in style, compact in size, efficient in random access and is the format in which alignments from the 1000 Genomes Project are released. SAMtools implements various utilities for post-processing alignments in the SAM format, such as indexing, variant caller and alignment viewer, and thus provides universal tools for processing read alignments. AVAILABILITY: http://samtools.sourceforge.net.
The human genome holds an extraordinary trove of information about human development, physiology, medicine and evolution. Here we report the results of an international collaboration to produce and make freely available a draft sequence of the human genome. We also present an initial analysis of the data, describing some of the insights that can be gleaned from the sequence.
The goal of the International HapMap Project is to determine the common patterns of DNA sequence variation in the human genome and to make this information freely available in the public domain. An international consortium is developing a map of these patterns across the genome by determining the genotypes of one million or more sequence variants, their frequencies and the degree of association between them, in DNA samples from populations with ancestry from parts of Africa, Asia and Europe. The HapMap will allow the discovery of sequence variants that affect common disease, will facilitate development of diagnostic tools, and will enhance our ability to choose targets for therapeutic intervention.
The genome of the japonica subspecies of rice, an important cereal and model monocot, was sequenced and assembled by whole-genome shotgun sequencing. The assembled sequence covers 93% of the 420-megabase genome. Gene predictions on the assembled sequence suggest that the genome contains 32,000 to 50,000 genes. Homologs of 98% of the known maize, wheat, and barley proteins are found in rice. Synteny and gene homology between rice and the other cereal genomes are extensive, whereas synteny with Arabidopsis is limited. Assignment of candidate rice orthologs to Arabidopsis genes is possible in many cases. The rice genome sequence provides a foundation for the improvement of cereals, our most important crops.
Genome-wide association studies have identified many noncoding variants associated with common diseases and traits. We show that these variants are concentrated in regulatory DNA marked by deoxyribonuclease I (DNase I) hypersensitive sites (DHSs). Eighty-eight percent of such DHSs are active during fetal development and are enriched in variants associated with gestational exposure-related phenotypes. We identified distant gene targets for hundreds of variant-containing DHSs that may explain phenotype associations. Disease-associated variants systematically perturb transcription factor recognition sequences, frequently alter allelic chromatin states, and form regulatory networks. We also demonstrated tissue-selective enrichment of more weakly disease-associated variants within DHSs and the de novo identification of pathogenic cell types for Crohn's disease, multiple sclerosis, and an electrocardiogram trait, without prior knowledge of physiological mechanisms. Our results suggest pervasive involvement of regulatory DNA variation in common human disease and provide pathogenic insights into diverse disorders.
DNase I hypersensitive sites (DHSs) are markers of regulatory DNA and have underpinned the discovery of all classes of cis-regulatory elements including enhancers, promoters, insulators, silencers and locus control regions. Here we present the first extensive map of human DHSs identified through genome-wide profiling in 125 diverse cell and tissue types. We identify ∼2.9 million DHSs that encompass virtually all known experimentally validated cis-regulatory sequences and expose a vast trove of novel elements, most with highly cell-selective regulation. Annotating these elements using ENCODE data reveals novel relationships between chromatin accessibility, transcription, DNA methylation and regulatory factor occupancy patterns. We connect ∼580,000 distal DHSs with their target promoters, revealing systematic pairing of different classes of distal DHSs and specific promoter types. Patterning of chromatin accessibility at many regulatory regions is organized with dozens to hundreds of co-activated elements, and the transcellular DNase I sensitivity pattern at a given region can predict cell-type-specific functional behaviours. The DHS landscape shows signatures of recent functional evolutionary constraint. However, the DHS compartment in pluripotent and immortalized cells exhibits higher mutation rates than that in highly differentiated cells, exposing an unexpected link between chromatin accessibility, proliferative potential and patterns of human variation. An extensive map of human DNase I hypersensitive sites, markers of regulatory DNA, in 125 diverse cell and tissue types is described; integration of this information with other ENCODE-generated data sets identifies new relationships between chromatin accessibility, transcription, DNA methylation and regulatory factor occupancy patterns. This paper describes the first extensive map of human DNaseI hypersensitive sites — markers of regulatory DNA — in 125 diverse cell and tissue types. Integration of this information with other data sets generated by ENCODE (Encyclopedia of DNA Elements) identified new relationships between chromatin accessibility, transcription, DNA methylation and regulatory-factor occupancy patterns. Evolutionary-conservation analysis revealed signatures of recent functional constraint within DNaseI hypersensitive sites.
New sequencing technologies promise a new era in the use of DNA sequence. However, some of these technologies produce very short reads, typically of a few tens of base pairs, and to use these reads effectively requires new algorithms and software. In particular, there is a major issue in efficiently aligning short reads to a reference genome and handling ambiguity or lack of accuracy in this alignment. Here we introduce the concept of mapping quality, a measure of the confidence that a read actually comes from the position it is aligned to by the mapping algorithm. We describe the software MAQ that can build assemblies by mapping shotgun short reads to a reference genome, using quality scores to derive genotype calls of the consensus sequence of a diploid genome, e.g., from a human sample. MAQ makes full use of mate-pair information and estimates the error probability of each read alignment. Error probabilities are also derived for the final genotype calls, using a Bayesian statistical model that incorporates the mapping qualities, error probabilities from the raw sequence quality scores, sampling of the two haplotypes, and an empirical model for correlated errors at a site. Both read mapping and genotype calling are evaluated on simulated data and real data. MAQ is accurate, efficient, versatile, and user-friendly. It is freely available at http://maq.sourceforge.net.
autophagic responses. Here, we critically discuss current methods of assessing autophagy and the information they can, or cannot, provide. Our ultimate goal is to encourage intellectual and technical innovation in the field.
The methyltransferase like 3 (METTL3)-containing methyltransferase complex catalyzes the N6-methyladenosine (m6A) formation, a novel epitranscriptomic marker; however, the nature of this complex remains largely unknown. Here we report two new components of the human m6A methyltransferase complex, Wilms' tumor 1-associating protein (WTAP) and methyltransferase like 14 (METTL14). WTAP interacts with METTL3 and METTL14, and is required for their localization into nuclear speckles enriched with pre-mRNA processing factors and for catalytic activity of the m6A methyltransferase in vivo. The majority of RNAs bound by WTAP and METTL3 in vivo represent mRNAs containing the consensus m6A motif. In the absence of WTAP, the RNA-binding capability of METTL3 is strongly reduced, suggesting that WTAP may function to regulate recruitment of the m6A methyltransferase complex to mRNA targets. Furthermore, transcriptomic analyses in combination with photoactivatable-ribonucleoside-enhanced crosslinking and immunoprecipitation (PAR-CLIP) illustrate that WTAP and METTL3 regulate expression and alternative splicing of genes involved in transcription and RNA processing. Morpholino-mediated knockdown targeting WTAP and/or METTL3 in zebrafish embryos caused tissue differentiation defects and increased apoptosis. These findings provide strong evidence that WTAP may function as a regulatory subunit in the m6A methyltransferase complex and play a critical role in epitranscriptomic regulation of RNA metabolism.
Exosomes are 40-100 nm nano-sized vesicles that are released from many cell types into the extracellular space. Such vesicles are widely distributed in various body fluids. Recently, mRNAs and microRNAs (miRNAs) have been identified in exosomes, which can be taken up by neighboring or distant cells and subsequently modulate recipient cells. This suggests an active sorting mechanism of exosomal miRNAs, since the miRNA profiles of exosomes may differ from those of the parent cells. Exosomal miRNAs play an important role in disease progression, and can stimulate angiogenesis and facilitate metastasis in cancers. In this review, we will introduce the origin and the trafficking of exosomes between cells, display current research on the sorting mechanism of exosomal miRNAs, and briefly describe how exosomes and their miRNAs function in recipient cells. Finally, we will discuss the potential applications of these miRNA-containing vesicles in clinical settings.
We present an integrated stand-alone software package named KaKs_Calculator 2.0 as an updated version. It incorporates 17 methods for the calculation of nonsynonymous and synonymous substitution rates; among them, we added our modified versions of several widely used methods as the gamma series including gamma-NG, gamma-LWL, gamma-MLWL, gamma-LPB, gamma-MLPB, gamma-YN and gamma-MYN, which have been demonstrated to perform better under certain conditions than their original forms and are not implemented in the previous version. The package is readily used for the identification of positively selected sites based on a sliding window across the sequences of interests in 5' to 3' direction of protein-coding sequences, and have improved the overall performance on sequence analysis for evolution studies. A toolbox, including C++ and Java source code and executable files on both Windows and Linux platforms together with a user instruction, is downloadable from the website for academic purpose at https://sourceforge.net/projects/kakscalculator2/.
The Genome Sequence Archive (GSA) is a data repository for archiving raw sequence data, which provides data storage and sharing services for worldwide scientific communities. Considering explosive data growth with diverse data types, here we present the GSA family by expanding into a set of resources for raw data archive with different purposes, namely, GSA (https://ngdc.cncb.ac.cn/gsa/), GSA for Human (GSA-Human, https://ngdc.cncb.ac.cn/gsa-human/), and Open Archive for Miscellaneous Data (OMIX, https://ngdc.cncb.ac.cn/omix/). Compared with the 2017 version, GSA has been significantly updated in data model, online functionalities, and web interfaces. GSA-Human, as a new partner of GSA, is a data repository specialized in human genetics-related data with controlled access and security. OMIX, as a critical complement to the two resources mentioned above, is an open archive for miscellaneous data. Together, all these resources form a family of resources dedicated to archiving explosive data with diverse types, accepting data submissions from all over the world, and providing free open access to all publicly available data in support of worldwide research activities.
Abstract N 6 -methyladenosine (m 6 A) is a chemical modification present in multiple RNA species, being most abundant in mRNAs. Studies on enzymes or factors that catalyze, recognize, and remove m 6 A have revealed its comprehensive roles in almost every aspect of mRNA metabolism, as well as in a variety of physiological processes. This review describes the current understanding of the m 6 A modification, particularly the functions of its writers, erasers, readers in RNA metabolism, with an emphasis on its role in regulating the isoform dosage of mRNAs.
BACKGROUND: Recently, the potential role of gut microbiome in metabolic diseases has been revealed, especially in cardiovascular diseases. Hypertension is one of the most prevalent cardiovascular diseases worldwide, yet whether gut microbiota dysbiosis participates in the development of hypertension remains largely unknown. To investigate this issue, we carried out comprehensive metagenomic and metabolomic analyses in a cohort of 41 healthy controls, 56 subjects with pre-hypertension, 99 individuals with primary hypertension, and performed fecal microbiota transplantation from patients to germ-free mice. RESULTS: Compared to the healthy controls, we found dramatically decreased microbial richness and diversity, Prevotella-dominated gut enterotype, distinct metagenomic composition with reduced bacteria associated with healthy status and overgrowth of bacteria such as Prevotella and Klebsiella, and disease-linked microbial function in both pre-hypertensive and hypertensive populations. Unexpectedly, the microbiome characteristic in pre-hypertension group was quite similar to that in hypertension. The metabolism changes of host with pre-hypertension or hypertension were identified to be closely linked to gut microbiome dysbiosis. And a disease classifier based on microbiota and metabolites was constructed to discriminate pre-hypertensive and hypertensive individuals from controls accurately. Furthermore, by fecal transplantation from hypertensive human donors to germ-free mice, elevated blood pressure was observed to be transferrable through microbiota, and the direct influence of gut microbiota on blood pressure of the host was demonstrated. CONCLUSIONS: Overall, our results describe a novel causal role of aberrant gut microbiota in contributing to the pathogenesis of hypertension. And the significance of early intervention for pre-hypertension was emphasized.
Crop domestications are long-term selection experiments that have greatly advanced human civilization. The domestication of cultivated rice (Oryza sativa L.) ranks as one of the most important developments in history. However, its origins and domestication processes are controversial and have long been debated. Here we generate genome sequences from 446 geographically diverse accessions of the wild rice species Oryza rufipogon, the immediate ancestral progenitor of cultivated rice, and from 1,083 cultivated indica and japonica varieties to construct a comprehensive map of rice genome variation. In the search for signatures of selection, we identify 55 selective sweeps that have occurred during domestication. In-depth analyses of the domestication sweeps and genome-wide patterns reveal that Oryza sativa japonica rice was first domesticated from a specific population of O. rufipogon around the middle area of the Pearl River in southern China, and that Oryza sativa indica rice was subsequently developed from crosses between japonica rice and local wild rice as the initial cultivars spread into South East and South Asia. The domestication-associated traits are analysed through high-resolution genetic mapping. This study provides an important resource for rice breeding and an effective genomics approach for crop domestication research. Whole-genome sequences of wild rice and cultivated rice varieties are used to produce a map of rice genome variation, and show that rice was probably first domesticated in southern China. Cultivated rice (Oryza sativa) is thought to have been domesticated from wild rice (Oryza rufipogon) thousands of years ago. This Chinese/Japanese collaboration reports whole-genome sequences from 446 wild rice isolates from across Asia and Oceana, and from more than 1,000 indica and japonica subspecies of cultivated rice. The resulting map of genome variation will be an important resource for rice breeding and for crop-domestication research.
Mass accuracy is a key parameter of mass spectrometric performance. TOF instruments can reach low parts per million, and FT-ICR instruments are capable of even greater accuracy provided ion numbers are well controlled. Here we demonstrate sub-ppm mass accuracy on a linear ion trap coupled via a radio frequency-only storage trap (C-trap) to the orbitrap mass spectrometer (LTQ Orbitrap). Prior to acquisition of a spectrum, a background ion originating from ambient air is first transferred to the C-trap. Ions forming the MS or MS(n) spectrum are then added to this species, and all ions are injected into the orbitrap for analysis. Real time recalibration on the "lock mass" by corrections of mass shift removes mass error associated with calibration of the mass scale. The remaining mass error is mainly due to imperfect peaks caused by weak signals and is addressed by averaging the mass measurement over the LC peak, weighted by signal intensity. For peptide database searches in proteomics, we introduce a variable mass tolerance and achieve average absolute mass deviations of 0.48 ppm (standard deviation 0.38 ppm) and maximal deviations of less than 2 ppm. For tandem mass spectra we demonstrate similarly high mass accuracy and discuss its impact on database searching. High and routine mass accuracy in a compact instrument will dramatically improve certainty of peptide and small molecule identification.
Long non-coding RNAs (lncRNAs) have been found to perform various functions in a wide variety of important biological processes. To make easier interpretation of lncRNA functionality and conduct deep mining on these transcribed sequences, it is convenient to classify lncRNAs into different groups. Here, we summarize classification methods of lncRNAs according to their four major features, namely, genomic location and context, effect exerted on DNA sequences, mechanism of functioning and their targeting mechanism. In combination with the presently available function annotations, we explore potential relationships between different classification categories, and generalize and compare biological features of different lncRNAs within each category. Finally, we present our view on potential further studies. We believe that the classifications of lncRNAs as indicated above are of fundamental importance for lncRNA studies, helpful for further investigation of specific lncRNAs, for formulation of new hypothesis based on different features of lncRNA and for exploration of the underlying lncRNA functional mechanisms.
Abstract KaKs_Calculator is a software package that calculates nonsynonymous (Ka) and synonymous (Ks) substitution rates through model selection and model averaging. Since existing methods for this estimation adopt their specific mutation (substitution) models that consider different evolutionary features, leading to diverse estimates, KaKs_Calculator implements a set of candidate models in a maximum likelihood framework and adopts the Akaike information criterion to measure fitness between models and data, aiming to include as many features as needed for accurately capturing evolutionary information in protein-coding sequences. In addition, several existing methods for calculating Ka and Ks are also incorporated into this software. KaKs_Calculator, including source codes, compiled executables, and documentation, is freely available for academic use at http://evolution.genomics.org.cn/software.htm.
The role of Fat Mass and Obesity-associated protein (FTO) and its substrate N6-methyladenosine (m6A) in mRNA processing and adipogenesis remains largely unknown. We show that FTO expression and m6A levels are inversely correlated during adipogenesis. FTO depletion blocks differentiation and only catalytically active FTO restores adipogenesis. Transcriptome analyses in combination with m6A-seq revealed that gene expression and mRNA splicing of grouped genes are regulated by FTO. M6A is enriched in exonic regions flanking 5'- and 3'-splice sites, spatially overlapping with mRNA splicing regulatory serine/arginine-rich (SR) protein exonic splicing enhancer binding regions. Enhanced levels of m6A in response to FTO depletion promotes the RNA binding ability of SRSF2 protein, leading to increased inclusion of target exons. FTO controls exonic splicing of adipogenic regulatory factor RUNX1T1 by regulating m6A levels around splice sites and thereby modulates differentiation. These findings provide compelling evidence that FTO-dependent m6A demethylation functions as a novel regulatory mechanism of RNA processing and plays a critical role in the regulation of adipogenesis.
5-methylcytosine (m 5 C) is a post-transcriptional RNA modification identified in both stable and highly abundant tRNAs and rRNAs, and in mRNAs. However, its regulatory role in mRNA metabolism is still largely unknown. Here, we reveal that m 5 C modification is enriched in CG-rich regions and in regions immediately downstream of translation initiation sites and has conserved, tissue-specific and dynamic features across mammalian transcriptomes. Moreover, m 5 C formation in mRNAs is mainly catalyzed by the RNA methyltransferase NSUN2, and m 5 C is specifically recognized by the mRNA export adaptor ALYREF as shown by in vitro and in vivo studies. NSUN2 modulates ALYREF's nuclear-cytoplasmic shuttling, RNA-binding affinity and associated mRNA export. Dysregulation of ALYREF-mediated mRNA export upon NSUN2 depletion could be restored by reconstitution of wild-type but not methyltransferase-defective NSUN2. Our study provides comprehensive m 5 C profiles of mammalian transcriptomes and suggests an essential role for m 5 C modification in mRNA export and post-transcriptional regulation.