Institute of Bioinformatics and Systems Biology
facilityOberschleißheim, Bavaria, Germany
Research output, citation impact, and the most-cited recent papers from Institute of Bioinformatics and Systems Biology (Germany). Aggregated across the NobleBlocks index of 300M+ scholarly works.
Top-cited papers from Institute of Bioinformatics and Systems Biology
Sorghum, an African grass related to sugar cane and maize, is grown for food, feed, fibre and fuel. We present an initial analysis of the ∼730-megabase Sorghum bicolor (L.) Moench genome, placing ∼98% of genes in their chromosomal context using whole-genome shotgun sequence validated by genetic, physical and syntenic information. Genetic recombination is largely confined to about one-third of the sorghum genome with gene order and density similar to those of rice. Retrotransposon accumulation in recombinationally recalcitrant heterochromatin explains the ∼75% larger genome size of sorghum compared with rice. Although gene and repetitive DNA distributions have been preserved since palaeopolyploidization ∼70 million years ago, most duplicated gene sets lost one member before the sorghum–rice divergence. Concerted evolution makes one duplicated chromosomal segment appear to be only a few million years old. About 24% of genes are grass-specific and 7% are sorghum-specific. Recent gene and microRNA duplications may contribute to sorghum’s drought tolerance. The Sorghum bicolor genome sequence is published this week. Sorghum is a cereal grown widely as food, animal feed, fibre and fuel. Tolerant to hot, dry conditions, it is a staple for large populations in the West African Sahel region. Comparisons of the genome with those of maize and rice shed light on the evolution of grasses and of C4 photosynthesis, which is particularly efficient at assimilating carbon at high temperatures. In addition, protein coding genes and miRNAs that could contribute to sorghum's drought tolerance may also be found. Sorghum yield improvement has lagged behind that of other crops and the availability of the genome sequence could provide a vital boost to work on its improvement. Sorghum is an African grass that is grown for food, animal feed and fuel. The current paper presents an initial analysis of the ∼730 megabase genome of Sorghum bicolor. Genome analysis and its comparison with maize and rice shed light on grass genome evolution and also provide insights into the evolution of C4 photosynthesis, as well as protein coding genes and miRNAs that might contribute to sorghum's drought tolerance.
eggNOG is a public resource that provides Orthologous Groups (OGs) of proteins at different taxonomic levels, each with integrated and summarized functional annotations. Developments since the latest public release include changes to the algorithm for creating OGs across taxonomic levels, making nested groups hierarchically consistent. This allows for a better propagation of functional terms across nested OGs and led to the novel annotation of 95 890 previously uncharacterized OGs, increasing overall annotation coverage from 67% to 72%. The functional annotations of OGs have been expanded to also provide Gene Ontology terms, KEGG pathways and SMART/Pfam domains for each group. Moreover, eggNOG now provides pairwise orthology relationships within OGs based on analysis of phylogenetic trees. We have also incorporated a framework for quickly mapping novel sequences to OGs based on precomputed HMM profiles. Finally, eggNOG version 4.5 incorporates a novel data set spanning 2605 viral OGs, covering 5228 proteins from 352 viral proteomes. All data are accessible for bulk downloading, as a web-service, and through a completely redesigned web interface. The new access points provide faster searches and a number of new browsing and visualization capabilities, facilitating the needs of both experts and less experienced users. eggNOG v4.5 is available at http://eggnog.embl.de.
The Gossypium genus is used to investigate emergent consequences of polyploidy in cotton species; comparative genomic analyses reveal a complex evolutionary history including interactions among subgenomes that result in genetic novelty in elite cottons and provide insight into the evolution of spinnable fibres. A phylogenetic and genomic study of plants of the cotton genus Gossypium provides insights into the role of polyploidy in the angiosperm evolution, and specifically, in the emergence of spinnable fibres in domesticated cottons. The authors show that an abrupt five- to sixfold ploidy increase about 60 million years ago, and allopolyploidy reuniting divergent genomes approximately 1–2 million years ago, conferred a roughly 30-fold duplication of ancestral flowering plant genes in the 'elite' cottons G. hirsutum and G. barbadense compared to their presumed progenitor G. raimondii. Polyploidy often confers emergent properties, such as the higher fibre productivity and quality of tetraploid cottons than diploid cottons bred for the same environments1. Here we show that an abrupt five- to sixfold ploidy increase approximately 60 million years (Myr) ago, and allopolyploidy reuniting divergent Gossypium genomes approximately 1–2 Myr ago2, conferred about 30–36-fold duplication of ancestral angiosperm (flowering plant) genes in elite cottons (Gossypium hirsutum and Gossypium barbadense), genetic complexity equalled only by Brassica3 among sequenced angiosperms. Nascent fibre evolution, before allopolyploidy, is elucidated by comparison of spinnable-fibred Gossypium herbaceum A and non-spinnable Gossypium longicalyx F genomes to one another and the outgroup D genome of non-spinnable Gossypium raimondii. The sequence of a G. hirsutum AtDt (in which ‘t’ indicates tetraploid) cultivar reveals many non-reciprocal DNA exchanges between subgenomes that may have contributed to phenotypic innovation and/or other emergent properties such as ecological adaptation by polyploids. Most DNA-level novelty in G. hirsutum recombines alleles from the D-genome progenitor native to its New World habitat and the Old World A-genome progenitor in which spinnable fibre evolved. Coordinated expression changes in proximal groups of functionally distinct genes, including a nuclear mitochondrial DNA block, may account for clusters of cotton-fibre quantitative trait loci affecting diverse traits. Opportunities abound for dissecting emergent properties of other polyploids, particularly angiosperms, by comparison to diploid progenitors and outgroups.
CORUM is a database that provides a manually curated repository of experimentally characterized protein complexes from mammalian organisms, mainly human (64%), mouse (16%) and rat (12%). Protein complexes are key molecular entities that integrate multiple gene products to perform cellular functions. The new CORUM 2.0 release encompasses 2837 protein complexes offering the largest and most comprehensive publicly available dataset of mammalian protein complexes. The CORUM dataset is built from 3198 different genes, representing approximately 16% of the protein coding genes in humans. Each protein complex is described by a protein complex name, subunit composition, function as well as the literature reference that characterizes the respective protein complex. Recent developments include mapping of functional annotation to Gene Ontology terms as well as cross-references to Entrez Gene identifiers. In addition, a 'Phylogenetic Conservation' analysis tool was implemented that analyses the potential occurrence of orthologous protein complex subunits in mammals and other selected groups of organisms. This allows one to predict the occurrence of protein complexes in different phylogenetic groups. CORUM is freely accessible at (http://mips.helmholtz-muenchen.de/genre/proj/corum/index.html).
Bread wheat (Triticum aestivum) is a globally important crop, accounting for 20 per cent of the calories consumed by humans. Major efforts are underway worldwide to increase wheat production by extending genetic diversity and analysing key traits, and genomic resources can accelerate progress. But so far the very large size and polyploid complexity of the bread wheat genome have been substantial barriers to genome analysis. Here we report the sequencing of its large, 17-gigabase-pair, hexaploid genome using 454 pyrosequencing, and comparison of this with the sequences of diploid ancestral and progenitor genomes. We identified between 94,000 and 96,000 genes, and assigned two-thirds to the three component genomes (A, B and D) of hexaploid wheat. High-resolution synteny maps identified many small disruptions to conserved gene order. We show that the hexaploid genome is highly dynamic, with significant loss of gene family members on polyploidization and domestication, and an abundance of gene fragments. Several classes of genes involved in energy harvesting, metabolism and growth are among expanded gene families that could be associated with crop productivity. Our analyses, coupled with the identification of extensive genetic variation, provide a resource for accelerating gene discovery and improving this major crop. Sequencing of the hexaploid bread wheat genome shows that it is highly dynamic, with significant loss of gene family members on polyploidization and domestication, and an abundance of gene fragments. Two groups in this issue report the compilation and analysis of the genome sequences of major cereal crops — bread wheat and barley — providing important resources for future crop improvement. Bread wheat accounts for one-fifth of the calories consumed by humankind. It has a very large and complex hexaploid genome of 17 Gigabases. Michael Bevan and colleagues have analysed the genome using 454 pyrosequencing and compared it with diploid ancestral and progenitor genomes. The authors discovered significant loss of gene family members upon polyploidization and domestication, and expansion of gene classes that may be associated with crop productivity. Barley is one of the earliest domesticated plant crops. Although diploid, it has a very large genome of 5.1 Gigabases. Nils Stein and colleagues describe a physical map anchored to a high-resolution genetic map, on top of which they have overlaid a deep whole-genome shotgun assembly, cDNA and RNA-seq data to provide the first in-depth genome-wide survey of the barley genome.
Sclerotinia sclerotiorum and Botrytis cinerea are closely related necrotrophic plant pathogenic fungi notable for their wide host ranges and environmental persistence. These attributes have made these species models for understanding the complexity of necrotrophic, broad host-range pathogenicity. Despite their similarities, the two species differ in mating behaviour and the ability to produce asexual spores. We have sequenced the genomes of one strain of S. sclerotiorum and two strains of B. cinerea. The comparative analysis of these genomes relative to one another and to other sequenced fungal genomes is provided here. Their 38-39 Mb genomes include 11,860-14,270 predicted genes, which share 83% amino acid identity on average between the two species. We have mapped the S. sclerotiorum assembly to 16 chromosomes and found large-scale co-linearity with the B. cinerea genomes. Seven percent of the S. sclerotiorum genome comprises transposable elements compared to <1% of B. cinerea. The arsenal of genes associated with necrotrophic processes is similar between the species, including genes involved in plant cell wall degradation and oxalic acid production. Analysis of secondary metabolism gene clusters revealed an expansion in number and diversity of B. cinerea-specific secondary metabolites relative to S. sclerotiorum. The potential diversity in secondary metabolism might be involved in adaptation to specific ecological niches. Comparative genome analysis revealed the basis of differing sexual mating compatibility systems between S. sclerotiorum and B. cinerea. The organization of the mating-type loci differs, and their structures provide evidence for the evolution of heterothallism from homothallism. These data shed light on the evolutionary and mechanistic bases of the genetically complex traits of necrotrophic pathogenicity and sexual mating. This resource should facilitate the functional studies designed to better understand what makes these fungi such successful and persistent pathogens of agronomic crops.
Genome-wide association studies (GWAS) with intermediate phenotypes, like changes in metabolite and protein levels, provide functional evidence to map disease associations and translate them into clinical applications. However, although hundreds of genetic variants have been associated with complex disorders, the underlying molecular pathways often remain elusive. Associations with intermediate traits are key in establishing functional links between GWAS-identified risk-variants and disease end points. Here we describe a GWAS using a highly multiplexed aptamer-based affinity proteomics platform. We quantify 539 associations between protein levels and gene variants (pQTLs) in a German cohort and replicate over half of them in an Arab and Asian cohort. Fifty-five of the replicated pQTLs are located in trans. Our associations overlap with 57 genetic risk loci for 42 unique disease end points. We integrate this information into a genome-proteome network and provide an interactive web-tool for interrogations. Our results provide a basis for novel approaches to pharmaceutical and diagnostic applications.
MicroRNA-122 (miR-122), which accounts for 70% of the liver's total miRNAs, plays a pivotal role in the liver. However, its intrinsic physiological roles remain largely undetermined. We demonstrated that mice lacking the gene encoding miR-122a (Mir122a) are viable but develop temporally controlled steatohepatitis, fibrosis, and hepatocellular carcinoma (HCC). These mice exhibited a striking disparity in HCC incidence based on sex, with a male-to-female ratio of 3.9:1, which recapitulates the disease incidence in humans. Impaired expression of microsomal triglyceride transfer protein (MTTP) contributed to steatosis, which was reversed by in vivo restoration of Mttp expression. We found that hepatic fibrosis onset can be partially attributed to the action of a miR-122a target, the Klf6 transcript. In addition, Mir122a(-/-) livers exhibited disruptions in a range of pathways, many of which closely resemble the disruptions found in human HCC. Importantly, the reexpression of miR-122a reduced disease manifestation and tumor incidence in Mir122a(-/-) mice. This study demonstrates that mice with a targeted deletion of the Mir122a gene possess several key phenotypes of human liver diseases, which provides a rationale for the development of a unique therapy for the treatment of chronic liver disease and HCC.
The rapidly evolving field of metabolomics aims at a comprehensive measurement of ideally all endogenous metabolites in a cell or body fluid. It thereby provides a functional readout of the physiological state of the human body. Genetic variants that associate with changes in the homeostasis of key lipids, carbohydrates, or amino acids are not only expected to display much larger effect sizes due to their direct involvement in metabolite conversion modification, but should also provide access to the biochemical context of such variations, in particular when enzyme coding genes are concerned. To test this hypothesis, we conducted what is, to the best of our knowledge, the first GWA study with metabolomics based on the quantitative measurement of 363 metabolites in serum of 284 male participants of the KORA study. We found associations of frequent single nucleotide polymorphisms (SNPs) with considerable differences in the metabolic homeostasis of the human body, explaining up to 12% of the observed variance. Using ratios of certain metabolite concentrations as a proxy for enzymatic activity, up to 28% of the variance can be explained (p-values 10(-16) to 10(-21)). We identified four genetic variants in genes coding for enzymes (FADS1, LIPC, SCAD, MCAD) where the corresponding metabolic phenotype (metabotype) clearly matches the biochemical pathways in which these enzymes are active. Our results suggest that common genetic polymorphisms induce major differentiations in the metabolic make-up of the human population. This may lead to a novel approach to personalized health care based on a combination of genotyping and metabolic characterization. These genetically determined metabotypes may subscribe the risk for a certain medical phenotype, the response to a given drug treatment, or the reaction to a nutritional intervention or environmental challenge.
Sequencing and analysing the diploid genome and transcriptome of Aegilops tauschii provide new insights into the role of this genome in enabling the adaptation of bread wheat and are a step towards understanding the very large and complicated hexaploid genomes of wheat species. The hexaploid genome of bread wheat Triticum aestivum, designated AABBDD, evolved as a result of hybridization between three ancestral grasses. Two papers published in the issue of Nature present genome sequences and analysis of two of these wheat progenitors. First, the genome sequence of the diploid wild wheat T. urartu (ancestor of the A genome), which resembles cultivated wheat more strongly than either Aegilops speltoides (the B ancestor) or Ae. tauschii (the D donor). And second, the Ae. tauschii genome, together with an analysis of its transcriptome. These genomes and their analyses will be powerful tools for the study of complex, polyploid wheat genomes and a valuable resource for genetic improvement of wheat. About 8,000 years ago in the Fertile Crescent, a spontaneous hybridization of the wild diploid grass Aegilops tauschii (2n = 14; DD) with the cultivated tetraploid wheat Triticum turgidum (2n = 4x = 28; AABB) resulted in hexaploid wheat (T. aestivum; 2n = 6x = 42; AABBDD)1,2. Wheat has since become a primary staple crop worldwide as a result of its enhanced adaptability to a wide range of climates and improved grain quality for the production of baker’s flour2. Here we describe sequencing the Ae. tauschii genome and obtaining a roughly 90-fold depth of short reads from libraries with various insert sizes, to gain a better understanding of this genetically complex plant. The assembled scaffolds represented 83.4% of the genome, of which 65.9% comprised transposable elements. We generated comprehensive RNA-Seq data and used it to identify 43,150 protein-coding genes, of which 30,697 (71.1%) were uniquely anchored to chromosomes with an integrated high-density genetic map. Whole-genome analysis revealed gene family expansion in Ae. tauschii of agronomically relevant gene families that were associated with disease resistance, abiotic stress tolerance and grain quality. This draft genome sequence provides insight into the environmental adaptation of bread wheat and can aid in defining the large and complicated genomes of wheat species.
INTRODUCTION: Increasing evidence suggests a role for the gut microbiome in central nervous system disorders and a specific role for the gut-brain axis in neurodegeneration. Bile acids (BAs), products of cholesterol metabolism and clearance, are produced in the liver and are further metabolized by gut bacteria. They have major regulatory and signaling functions and seem dysregulated in Alzheimer's disease (AD). METHODS: Serum levels of 15 primary and secondary BAs and their conjugated forms were measured in 1464 subjects including 370 cognitively normal older adults, 284 with early mild cognitive impairment, 505 with late mild cognitive impairment, and 305 AD cases enrolled in the AD Neuroimaging Initiative. We assessed associations of BA profiles including selected ratios with diagnosis, cognition, and AD-related genetic variants, adjusting for confounders and multiple testing. RESULTS: In AD compared to cognitively normal older adults, we observed significantly lower serum concentrations of a primary BA (cholic acid [CA]) and increased levels of the bacterially produced, secondary BA, deoxycholic acid, and its glycine and taurine conjugated forms. An increased ratio of deoxycholic acid:CA, which reflects 7α-dehydroxylation of CA by gut bacteria, strongly associated with cognitive decline, a finding replicated in serum and brain samples in the Rush Religious Orders and Memory and Aging Project. Several genetic variants in immune response-related genes implicated in AD showed associations with BA profiles. DISCUSSION: We report for the first time an association between altered BA profile, genetic variants implicated in AD, and cognitive changes in disease using a large multicenter study. These findings warrant further investigation of gut dysbiosis and possible role of gut-liver-brain axis in the pathogenesis of AD.
Type 2 diabetes (T2D) can be prevented in pre-diabetic individuals with impaired glucose tolerance (IGT). Here, we have used a metabolomics approach to identify candidate biomarkers of pre-diabetes. We quantified 140 metabolites for 4297 fasting serum samples in the population-based Cooperative Health Research in the Region of Augsburg (KORA) cohort. Our study revealed significant metabolic variation in pre-diabetic individuals that are distinct from known diabetes risk indicators, such as glycosylated hemoglobin levels, fasting glucose and insulin. We identified three metabolites (glycine, lysophosphatidylcholine (LPC) (18:2) and acetylcarnitine) that had significantly altered levels in IGT individuals as compared to those with normal glucose tolerance, with P-values ranging from 2.4×10(-4) to 2.1×10(-13). Lower levels of glycine and LPC were found to be predictors not only for IGT but also for T2D, and were independently confirmed in the European Prospective Investigation into Cancer and Nutrition (EPIC)-Potsdam cohort. Using metabolite-protein network analysis, we identified seven T2D-related genes that are associated with these three IGT-specific metabolites by multiple interactions with four enzymes. The expression levels of these enzymes correlate with changes in the metabolite concentrations linked to diabetes. Our results may help developing novel strategies to prevent T2D.
CORUM is a database that provides a manually curated repository of experimentally characterized protein complexes from mammalian organisms, mainly human (67%), mouse (15%) and rat (10%). Given the vital functions of these macromolecular machines, their identification and functional characterization is foundational to our understanding of normal and disease biology. The new CORUM 3.0 release encompasses 4274 protein complexes offering the largest and most comprehensive publicly available dataset of mammalian protein complexes. The CORUM dataset is built from 4473 different genes, representing 22% of the protein coding genes in humans. Protein complexes are described by a protein complex name, subunit composition, cellular functions as well as the literature references. Information about stoichiometry of subunits depends on availability of experimental data. Recent developments include a graphical tool displaying known interactions between subunits. This allows the prediction of structural interconnections within protein complexes of unknown structure. In addition, we present a set of 58 protein complexes with alternatively spliced subunits. Those were found to affect cellular functions such as regulation of apoptotic activity, protein complex assembly or define cellular localization. CORUM is freely accessible at http://mips.helmholtz-muenchen.de/corum/.
Picoeukaryotes are a taxonomically diverse group of organisms less than 2 micrometers in diameter. Photosynthetic marine picoeukaryotes in the genus Micromonas thrive in ecosystems ranging from tropical to polar and could serve as sentinel organisms for biogeochemical fluxes of modern oceans during climate change. These broadly distributed primary producers belong to an anciently diverged sister clade to land plants. Although Micromonas isolates have high 18S ribosomal RNA gene identity, we found that genomes from two isolates shared only 90% of their predicted genes. Their independent evolutionary paths were emphasized by distinct riboswitch arrangements as well as the discovery of intronic repeat elements in one isolate, and in metagenomic data, but not in other genomes. Divergence appears to have been facilitated by selection and acquisition processes that actively shape the repertoire of genes that are mutually exclusive between the two isolates differently than the core genes. Analyses of the Micromonas genomes offer valuable insights into ecological differentiation and the dynamic nature of early plant evolution.
The human gut microbiome has been associated with many health factors but variability between studies limits exploration of effects between them. Gut microbiota profiles are available for >2700 members of the deeply phenotyped TwinsUK cohort, providing a uniform platform for such comparisons. Here, we present gut microbiota association analyses for 38 common diseases and 51 medications within the cohort. We describe several novel associations, highlight associations common across multiple diseases, and determine which diseases and medications have the greatest association with the gut microbiota. These results provide a reference for future studies of the gut microbiome and its role in human health.
We produced a reference sequence of the 1-gigabase chromosome 3B of hexaploid bread wheat. By sequencing 8452 bacterial artificial chromosomes in pools, we assembled a sequence of 774 megabases carrying 5326 protein-coding genes, 1938 pseudogenes, and 85% of transposable elements. The distribution of structural and functional features along the chromosome revealed partitioning correlated with meiotic recombination. Comparative analyses indicated high wheat-specific inter- and intrachromosomal gene duplication activities that are potential sources of variability for adaption. In addition to providing a better understanding of the organization, function, and evolution of a large and polyploid genome, the availability of a high-quality sequence anchored to genetic maps will accelerate the identification of genes underlying important agronomic traits.
INTRODUCTION BACKGROUND TO METABOLOMICS: Metabolomics is the comprehensive study of the metabolome, the repertoire of biochemicals (or small molecules) present in cells, tissues, and body fluids. The study of metabolism at the global or "-omics" level is a rapidly growing field that has the potential to have a profound impact upon medical practice. At the center of metabolomics, is the concept that a person's metabolic state provides a close representation of that individual's overall health status. This metabolic state reflects what has been encoded by the genome, and modified by diet, environmental factors, and the gut microbiome. The metabolic profile provides a quantifiable readout of biochemical state from normal physiology to diverse pathophysiologies in a manner that is often not obvious from gene expression analyses. Today, clinicians capture only a very small part of the information contained in the metabolome, as they routinely measure only a narrow set of blood chemistry analytes to assess health and disease states. Examples include measuring glucose to monitor diabetes, measuring cholesterol and high density lipoprotein/low density lipoprotein ratio to assess cardiovascular health, BUN and creatinine for renal disorders, and measuring a panel of metabolites to diagnose potential inborn errors of metabolism in neonates. OBJECTIVES OF WHITE PAPER—EXPECTED TREATMENT OUTCOMES AND METABOLOMICS ENABLING TOOL FOR PRECISION MEDICINE: We anticipate that the narrow range of chemical analyses in current use by the medical community today will be replaced in the future by analyses that reveal a far more comprehensive metabolic signature. This signature is expected to describe global biochemical aberrations that reflect patterns of variance in states of wellness, more accurately describe specific diseases and their progression, and greatly aid in differential diagnosis. Such future metabolic signatures will: (1) provide predictive, prognostic, diagnostic, and surrogate markers of diverse disease states; (2) inform on underlying molecular mechanisms of diseases; (3) allow for sub-classification of diseases, and stratification of patients based on metabolic pathways impacted; (4) reveal biomarkers for drug response phenotypes, providing an effective means to predict variation in a subject's response to treatment (pharmacometabolomics); (5) define a metabotype for each specific genotype, offering a functional read-out for genetic variants: (6) provide a means to monitor response and recurrence of diseases, such as cancers: (7) describe the molecular landscape in human performance applications and extreme environments. Importantly, sophisticated metabolomic analytical platforms and informatics tools have recently been developed that make it possible to measure thousands of metabolites in blood, other body fluids, and tissues. Such tools also enable more robust analysis of response to treatment. New insights have been gained about mechanisms of diseases, including neuropsychiatric disorders, cardiovascular disease, cancers, diabetes and a range of pathologies. A series of ground breaking studies supported by National Institute of Health (NIH) through the Pharmacometabolomics Research Network and its partnership with the Pharmacogenomics Research Network illustrate how a patient's metabotype at baseline, prior to treatment, during treatment, and post-treatment, can inform about treatment outcomes and variations in responsiveness to drugs (e.g., statins, antidepressants, antihypertensives and antiplatelet therapies). These studies along with several others also exemplify how metabolomics data can complement and inform genetic data in defining ethnic, sex, and gender basis for variation in responses to treatment, which illustrates how pharmacometabolomics and pharmacogenomics are complementary and powerful tools for precision medicine. CONCLUSIONS KEY SCIENTIFIC CONCEPTS AND RECOMMENDATIONS FOR PRECISION MEDICINE: Our metabolomics community believes that inclusion of metabolomics data in precision medicine initiatives is timely and will provide an extremely valuable layer of data that compliments and informs other data obtained by these important initiatives. Our Metabolomics Society, through its "Precision Medicine and Pharmacometabolomics Task Group", with input from our metabolomics community at large, has developed this White Paper where we discuss the value and approaches for including metabolomics data in large precision medicine initiatives. This White Paper offers recommendations for the selection of state of-the-art metabolomics platforms and approaches that offer the widest biochemical coverage, considers critical sample collection and preservation, as well as standardization of measurements, among other important topics. We anticipate that our metabolomics community will have representation in large precision medicine initiatives to provide input with regard to sample acquisition/preservation, selection of optimal omics technologies, and key issues regarding data collection, interpretation, and dissemination. We strongly recommend the collection and biobanking of samples for precision medicine initiatives that will take into consideration needs for large-scale metabolic phenotyping studies.
Across a variety of Mendelian disorders, ∼50-75% of patients do not receive a genetic diagnosis by exome sequencing indicating disease-causing variants in non-coding regions. Although genome sequencing in principle reveals all genetic variants, their sizeable number and poorer annotation make prioritization challenging. Here, we demonstrate the power of transcriptome sequencing to molecularly diagnose 10% (5 of 48) of mitochondriopathy patients and identify candidate genes for the remainder. We find a median of one aberrantly expressed gene, five aberrant splicing events and six mono-allelically expressed rare variants in patient-derived fibroblasts and establish disease-causing roles for each kind. Private exons often arise from cryptic splice sites providing an important clue for variant prioritization. One such event is found in the complex I assembly factor TIMMDC1 establishing a novel disease-associated gene. In conclusion, our study expands the diagnostic tools for detecting non-exonic variants and provides examples of intronic loss-of-function variants with pathological relevance.
BACKGROUND: Metabolomics is the rapidly evolving field of the comprehensive measurement of ideally all endogenous metabolites in a biological fluid. However, no single analytic technique covers the entire spectrum of the human metabolome. Here we present results from a multiplatform study, in which we investigate what kind of results can presently be obtained in the field of diabetes research when combining metabolomics data collected on a complementary set of analytical platforms in the framework of an epidemiological study. METHODOLOGY/PRINCIPAL FINDINGS: 40 individuals with self-reported diabetes and 60 controls (male, over 54 years) were randomly selected from the participants of the population-based KORA (Cooperative Health Research in the Region of Augsburg) study, representing an extensively phenotyped sample of the general German population. Concentrations of over 420 unique small molecules were determined in overnight-fasting blood using three different techniques, covering nuclear magnetic resonance and tandem mass spectrometry. Known biomarkers of diabetes could be replicated by this multiple metabolomic platform approach, including sugar metabolites (1,5-anhydroglucoitol), ketone bodies (3-hydroxybutyrate), and branched chain amino acids. In some cases, diabetes-related medication can be detected (pioglitazone, salicylic acid). CONCLUSIONS/SIGNIFICANCE: Our study depicts the promising potential of metabolomics in diabetes research by identification of a series of known and also novel, deregulated metabolites that associate with diabetes. Key observations include perturbations of metabolic pathways linked to kidney dysfunction (3-indoxyl sulfate), lipid metabolism (glycerophospholipids, free fatty acids), and interaction with the gut microflora (bile acids). Our study suggests that metabolic markers hold the potential to detect diabetes-related complications already under sub-clinical conditions in the general population.
Abstract Introduction The Alzheimer's Disease Research Summits of 2012 and 2015 incorporated experts from academia, industry, and nonprofit organizations to develop new research directions to transform our understanding of Alzheimer's disease (AD) and propel the development of critically needed therapies. In response to their recommendations, big data at multiple levels are being generated and integrated to study network failures in disease. We used metabolomics as a global biochemical approach to identify peripheral metabolic changes in AD patients and correlate them to cerebrospinal fluid pathology markers, imaging features, and cognitive performance. Methods Fasting serum samples from the Alzheimer's Disease Neuroimaging Initiative (199 control, 356 mild cognitive impairment, and 175 AD participants) were analyzed using the AbsoluteIDQ‐p180 kit. Performance was validated in blinded replicates, and values were medication adjusted. Results Multivariable‐adjusted analyses showed that sphingomyelins and ether‐containing phosphatidylcholines were altered in preclinical biomarker‐defined AD stages, whereas acylcarnitines and several amines, including the branched‐chain amino acid valine and α‐aminoadipic acid, changed in symptomatic stages. Several of the analytes showed consistent associations in the Rotterdam, Erasmus Rucphen Family, and Indiana Memory and Aging Studies. Partial correlation networks constructed for Aβ1–42, tau, imaging, and cognitive changes provided initial biochemical insights for disease‐related processes. Coexpression networks interconnected key metabolic effectors of disease. Discussion Metabolomics identified key disease‐related metabolic changes and disease‐progression‐related changes. Defining metabolic changes during AD disease trajectory and its relationship to clinical phenotypes provides a powerful roadmap for drug and biomarker discovery.