Center for Systems Biology Dresden
facilityDresden, Saxony, Germany
Research output, citation impact, and the most-cited recent papers from Center for Systems Biology Dresden (Germany). Aggregated across the NobleBlocks index of 300M+ scholarly works.
Top-cited papers from Center for Systems Biology Dresden
Abstract High-quality and complete reference genome assemblies are fundamental for the application of genomics to biology, disease, and biodiversity conservation. However, such assemblies are available for only a few non-microbial species 1–4 . To address this issue, the international Genome 10K (G10K) consortium 5,6 has worked over a five-year period to evaluate and develop cost-effective methods for assembling highly accurate and nearly complete reference genomes. Here we present lessons learned from generating assemblies for 16 species that represent six major vertebrate lineages. We confirm that long-read sequencing technologies are essential for maximizing genome quality, and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Our assemblies correct substantial errors, add missing sequence in some of the best historical reference genomes, and reveal biological discoveries. These include the identification of many false gene duplications, increases in gene sizes, chromosome rearrangements that are specific to lineages, a repeated independent chromosome breakpoint in bat genomes, and a canonical GC-rich pattern in protein-coding genes and their regulatory regions. Adopting these lessons, we have embarked on the Vertebrate Genomes Project (VGP), an international effort to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species and to help to enable a new era of discovery across the life sciences.
The field of image denoising is currently dominated by discriminative deep learning methods that are trained on pairs of noisy input and clean target images. Recently it has been shown that such methods can also be trained without clean targets. Instead, independent pairs of noisy images can be used, in an approach known as Noise2Noise (N2N). Here, we introduce Noise2Void (N2V), a training scheme that takes this idea one step further. It does not require noisy image pairs, nor clean target images. Consequently, N2V allows us to train directly on the body of data to be denoised and can therefore be applied when other methods cannot. Especially interesting is the application to biomedical image data, where the acquisition of training targets, clean or noisy, is frequently not possible. We compare the performance of N2V to approaches that have either clean target images and/or noisy image pairs available. Intuitively, N2V cannot be expected to outperform methods that have more information available during training. Still, we observe that the denoising performance of Noise2Void drops in moderation and compares favorably to training-free denoising methods.
The novel coronavirus severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is the cause of COVID-19. The main receptor of SARS-CoV-2, angiotensin I converting enzyme 2 (ACE2), is now undergoing extensive scrutiny to understand the routes of transmission and sensitivity in different species. Here, we utilized a unique dataset of ACE2 sequences from 410 vertebrate species, including 252 mammals, to study the conservation of ACE2 and its potential to be used as a receptor by SARS-CoV-2. We designed a five-category binding score based on the conservation properties of 25 amino acids important for the binding between ACE2 and the SARS-CoV-2 spike protein. Only mammals fell into the medium to very high categories and only catarrhine primates into the very high category, suggesting that they are at high risk for SARS-CoV-2 infection. We employed a protein structural analysis to qualitatively assess whether amino acid changes at variable residues would be likely to disrupt ACE2/SARS-CoV-2 spike protein binding and found the number of predicted unfavorable changes significantly correlated with the binding score. Extending this analysis to human population data, we found only rare (frequency <0.001) variants in 10/25 binding sites. In addition, we found significant signals of selection and accelerated evolution in the ACE2 coding sequence across all mammals, and specific to the bat lineage. Our results, if confirmed by additional experimental data, may lead to the identification of intermediate host species for SARS-CoV-2, guide the selection of animal models of COVID-19, and assist the conservation of animals both in native habitats and in human care.
Abstract Salamanders serve as important tetrapod models for developmental, regeneration and evolutionary studies. An extensive molecular toolkit makes the Mexican axolotl ( Ambystoma mexicanum ) a key representative salamander for molecular investigations. Here we report the sequencing and assembly of the 32-gigabase-pair axolotl genome using an approach that combined long-read sequencing, optical mapping and development of a new genome assembler (MARVEL). We observed a size expansion of introns and intergenic regions, largely attributable to multiplication of long terminal repeat retroelements. We provide evidence that intron size in developmental genes is under constraint and that species-restricted genes may contribute to limb regeneration. The axolotl genome assembly does not contain the essential developmental gene Pax3 . However, mutation of the axolotl Pax3 paralogue Pax7 resulted in an axolotl phenotype that was similar to those seen in Pax3 −/− and Pax7 −/− mutant mice. The axolotl genome provides a rich biological resource for developmental and evolutionary studies.
Deep Learning (DL) methods are powerful analytical tools for microscopy and can outperform conventional image processing pipelines. Despite the enthusiasm and innovations fuelled by DL technology, the need to access powerful and compatible resources to train DL networks leads to an accessibility barrier that novice users often find difficult to overcome. Here, we present ZeroCostDL4Mic, an entry-level platform simplifying DL access by leveraging the free, cloud-based computational resources of Google Colab. ZeroCostDL4Mic allows researchers with no coding expertise to train and apply key DL networks to perform tasks including segmentation (using U-Net and StarDist), object detection (using YOLOv2), denoising (using CARE and Noise2Void), super-resolution microscopy (using Deep-STORM), and image-to-image translation (using Label-free prediction - fnet, pix2pix and CycleGAN). Importantly, we provide suitable quantitative tools for each network to evaluate model performance, allowing model optimisation. We demonstrate the application of the platform to study multiple biological processes.
Animal genomes are folded into loops and topologically associating domains (TADs) by CTCF and loop-extruding cohesins, but the live dynamics of loop formation and stability remain unknown. Here, we directly visualized chromatin looping at the Fbn2 TAD in mouse embryonic stem cells using super-resolution live-cell imaging and quantified looping dynamics by Bayesian inference. Unexpectedly, the Fbn2 loop was both rare and dynamic, with a looped fraction of approximately 3 to 6.5% and a median loop lifetime of approximately 10 to 30 minutes. Our results establish that the Fbn2 TAD is highly dynamic, and about 92% of the time, cohesin-extruded loops exist within the TAD without bridging both CTCF boundaries. This suggests that single CTCF boundaries, rather than the fully CTCF-CTCF looped state, may be the primary regulators of functional interactions.
The nucleus contains diverse phase-separated condensates that compartmentalize and concentrate biomolecules with distinct physicochemical properties. Here, we investigated whether condensates concentrate small-molecule cancer therapeutics such that their pharmacodynamic properties are altered. We found that antineoplastic drugs become concentrated in specific protein condensates in vitro and that this occurs through physicochemical properties independent of the drug target. This behavior was also observed in tumor cells, where drug partitioning influenced drug activity. Altering the properties of the condensate was found to affect the concentration and activity of drugs. These results suggest that selective partitioning and concentration of small molecules within condensates contributes to drug pharmacodynamics and that further understanding of this phenomenon may facilitate advances in disease therapy.
For decades, biologists have relied on software to visualize and interpret imaging data. As techniques for acquiring images increase in complexity, resulting in larger multidimensional datasets, imaging software must adapt. ImageJ is an open-source image analysis software platform that has aided researchers with a variety of image analysis applications, driven mainly by engaged and collaborative user and developer communities. The close collaboration between programmers and users has resulted in adaptations to accommodate new challenges in image analysis that address the needs of ImageJ's diverse user base. ImageJ consists of many components, some relevant primarily for developers and a vast collection of user-centric plugins. It is available in many forms, including the widely used Fiji distribution. We refer to this entire ImageJ codebase and community as the ImageJ ecosystem. Here we review the core features of this ecosystem and highlight how ImageJ has responded to imaging technology advancements with new plugins and tools in recent years. These plugins and tools have been developed to address user needs in several areas such as visualization, segmentation, and tracking of biological entities in large, complex datasets. Moreover, new capabilities for deep learning are being added to ImageJ, reflecting a shift in the bioimage analysis community towards exploiting artificial intelligence. These new tools have been facilitated by profound architectural changes to the ImageJ core brought about by the ImageJ2 project. Therefore, we also discuss the contributions of ImageJ2 to enhancing multidimensional image processing and interoperability in the ImageJ ecosystem.
Abstract Bats possess extraordinary adaptations, including flight, echolocation, extreme longevity and unique immunity. High-quality genomes are crucial for understanding the molecular basis and evolution of these traits. Here we incorporated long-read sequencing and state-of-the-art scaffolding protocols 1 to generate, to our knowledge, the first reference-quality genomes of six bat species ( Rhinolophus ferrumequinum , Rousettus aegyptiacus , Phyllostomus discolor , Myotis myotis , Pipistrellus kuhlii and Molossus molossus ). We integrated gene projections from our ‘Tool to infer Orthologs from Genome Alignments’ (TOGA) software with de novo and homology gene predictions as well as short- and long-read transcriptomics to generate highly complete gene annotations. To resolve the phylogenetic position of bats within Laurasiatheria, we applied several phylogenetic methods to comprehensive sets of orthologous protein-coding and noncoding regions of the genome, and identified a basal origin for bats within Scrotifera. Our genome-wide screens revealed positive selection on hearing-related genes in the ancestral branch of bats, which is indicative of laryngeal echolocation being an ancestral trait in this clade. We found selection and loss of immunity-related genes (including pro-inflammatory NF-κB regulators) and expansions of anti-viral APOBEC3 genes, which highlights molecular mechanisms that may contribute to the exceptional immunity of bats. Genomic integrations of diverse viruses provide a genomic record of historical tolerance to viral infection in bats. Finally, we found and experimentally validated bat-specific variation in microRNAs, which may regulate bat-specific gene-expression programs. Our reference-quality bat genomes provide the resources required to uncover and validate the genomic basis of adaptations of bats, and stimulate new avenues of research that are directly relevant to human health and disease 1 .
Expression of proteins inside cells is noisy, causing variability in protein concentration among identical cells. A central problem in cellular control is how cells cope with this inherent noise. Compartmentalization of proteins through phase separation has been suggested as a potential mechanism to reduce noise, but systematic studies to support this idea have been missing. In this study, we used a physical model that links noise in protein concentration to theory of phase separation to show that liquid droplets can effectively reduce noise. We provide experimental support for noise reduction by phase separation using engineered proteins that form liquid-like compartments in mammalian cells. Thus, phase separation can play an important role in biological signal processing and control.
We present LABKIT, a user-friendly Fiji plugin for the segmentation of microscopy image data. It offers easy to use manual and automated image segmentation routines that can be rapidly applied to single- and multi-channel images as well as to timelapse movies in 2D or 3D. LABKIT is specifically designed to work efficiently on big image data and enables users of consumer laptops to conveniently work with multiple-terabyte images. This efficiency is achieved by using ImgLib2 and BigDataViewer as well as a memory efficient and fast implementation of the random forest based pixel classification algorithm as the foundation of our software. Optionally we harness the power of graphics processing units (GPU) to gain additional runtime performance. LABKIT is easy to install on virtually all laptops and workstations. Additionally, LABKIT is compatible with high performance computing (HPC) clusters for distributed processing of big image data. The ability to use pixel classifiers trained in LABKIT via the ImageJ macro language enables our users to integrate this functionality as a processing step in automated image processing workflows. Finally, LABKIT comes with rich online resources such as tutorials and examples that will help users to familiarize themselves with available features and how to best use LABKIT in a number of practical real-world use-cases.
Phase separating systems that are maintained away from thermodynamic equilibrium via molecular processes represent a class of active systems, which we call active emulsions. These systems are driven by external energy input, for example provided by an external fuel reservoir. The external energy input gives rise to novel phenomena that are not present in passive systems. For instance, concentration gradients can spatially organise emulsions and cause novel droplet size distributions. Another example are active droplets that are subject to chemical reactions such that their nucleation and size can be controlled, and they can divide spontaneously. In this review, we discuss the physics of phase separation and emulsions and show how the concepts that govern such phenomena can be extended to capture the physics of active emulsions. This physics is relevant to the spatial organisation of the biochemistry in living cells, for the development of novel applications in chemical engineering and models for the origin of life.
Phase separating systems that are maintained away from thermodynamic equilibrium via molecular processes represent a class of active systems, which we call active emulsions. These systems are driven by external energy input, for example provided by an external fuel reservoir. The external energy input gives rise to novel phenomena that are not present in passive systems. For instance, concentration gradients can spatially organise emulsions and cause novel droplet size distributions. Another example are active droplets that are subject to chemical reactions such that their nucleation and size can be controlled, and they can divide spontaneously. In this review, we discuss the physics of phase separation and emulsions and show how the concepts that govern such phenomena can be extended to capture the physics of active emulsions. This physics is relevant to the spatial organisation of the biochemistry in living cells, for the development of novel applications in chemical engineering and models for the origin of life.
Identifying the genomic changes that underlie phenotypic adaptations is a key challenge in evolutionary biology and genomics. Loss of protein-coding genes is one type of genomic change with the potential to affect phenotypic evolution. Here, we develop a genomics approach to accurately detect gene losses and investigate their importance for adaptive evolution in mammals. We discover a number of gene losses that likely contributed to morphological, physiological, and metabolic adaptations in aquatic and flying mammals. These gene losses shed light on possible molecular and cellular mechanisms that underlie these adaptive phenotypes. In addition, we show that gene loss events that occur as a consequence of relaxed selection following adaptation provide novel insights into species' biology. Our results suggest that gene loss is an evolutionary mechanism for adaptation that may be more widespread than previously anticipated. Hence, investigating gene losses has great potential to reveal the genomic basis underlying macroevolutionary changes.
Abstract In situ, reversible coacervate formation within lipid vesicles represents a key step in the development of responsive synthetic cellular models. Herein, we exploit the pH responsiveness of a polycation above and below its pK a , to drive liquid–liquid phase separation, to form single coacervate droplets within lipid vesicles. The process is completely reversible as coacervate droplets can be disassembled by increasing the pH above the pK a . We further show that pH‐triggered coacervation in the presence of low concentrations of enzymes activates dormant enzyme reactions by increasing the local concentration within the coacervate droplets and changing the local environment around the enzyme. In conclusion, this work establishes a tunable, pH responsive, enzymatically active multi‐compartment synthetic cell. The system is readily transferred into microfluidics, making it a robust model for addressing general questions in biology, such as the role of phase separation and its effect on enzymatic reactions using a bottom‐up synthetic biology approach.
Loop extrusion by structural maintenance of chromosomes (SMC) complexes has been proposed as a mechanism to organize chromatin in interphase and metaphase. However, the requirements for chromatin organization in these cell cycle phases are different, and it is unknown whether loop extrusion dynamics and the complexes that extrude DNA also differ. Here, we used Xenopus egg extracts to reconstitute and image loop extrusion of single DNA molecules during the cell cycle. We show that loops form in both metaphase and interphase, but with distinct dynamic properties. Condensin extrudes DNA loops non-symmetrically in metaphase, whereas cohesin extrudes loops symmetrically in interphase. Our data show that loop extrusion is a general mechanism underlying DNA organization, with dynamic and structural properties that are biochemically regulated during the cell cycle.
Background: Acute myocarditis (AM) is thought to be a rare cardiovascular complication of COVID-19, although minimal data are available beyond case reports. We aim to report the prevalence, baseline characteristics, in-hospital management, and outcomes for patients with COVID-19–associated AM on the basis of a retrospective cohort from 23 hospitals in the United States and Europe. Methods: A total of 112 patients with suspected AM from 56 963 hospitalized patients with COVID-19 were evaluated between February 1, 2020, and April 30, 2021. Inclusion criteria were hospitalization for COVID-19 and a diagnosis of AM on the basis of endomyocardial biopsy or increased troponin level plus typical signs of AM on cardiac magnetic resonance imaging. We identified 97 patients with possible AM, and among them, 54 patients with definite/probable AM supported by endomyocardial biopsy in 17 (31.5%) patients or magnetic resonance imaging in 50 (92.6%). We analyzed patient characteristics, treatments, and outcomes among all COVID-19–associated AM. Results: AM prevalence among hospitalized patients with COVID-19 was 2.4 per 1000 hospitalizations considering definite/probable and 4.1 per 1000 considering also possible AM. The median age of definite/probable cases was 38 years, and 38.9% were female. On admission, chest pain and dyspnea were the most frequent symptoms (55.5% and 53.7%, respectively). Thirty-one cases (57.4%) occurred in the absence of COVID-19–associated pneumonia. Twenty-one (38.9%) had a fulminant presentation requiring inotropic support or temporary mechanical circulatory support. The composite of in-hospital mortality or temporary mechanical circulatory support occurred in 20.4%. At 120 days, estimated mortality was 6.6%, 15.1% in patients with associated pneumonia versus 0% in patients without pneumonia ( P =0.044). During hospitalization, left ventricular ejection fraction, assessed by echocardiography, improved from a median of 40% on admission to 55% at discharge (n=47; P <0.0001) similarly in patients with or without pneumonia. Corticosteroids were frequently administered (55.5%). Conclusions: AM occurrence is estimated between 2.4 and 4.1 out of 1000 patients hospitalized for COVID-19. The majority of AM occurs in the absence of pneumonia and is often complicated by hemodynamic instability. AM is a rare complication in patients hospitalized for COVID-19, with an outcome that differs on the basis of the presence of concomitant pneumonia.
Abstract The transition from ‘well-marked varieties’ of a single species into ‘well-defined species’—especially in the absence of geographic barriers to gene flow (sympatric speciation)—has puzzled evolutionary biologists ever since Darwin 1,2 . Gene flow counteracts the buildup of genome-wide differentiation, which is a hallmark of speciation and increases the likelihood of the evolution of irreversible reproductive barriers (incompatibilities) that complete the speciation process 3 . Theory predicts that the genetic architecture of divergently selected traits can influence whether sympatric speciation occurs 4 , but empirical tests of this theory are scant because comprehensive data are difficult to collect and synthesize across species, owing to their unique biologies and evolutionary histories 5 . Here, within a young species complex of neotropical cichlid fishes ( Amphilophus spp.), we analysed genomic divergence among populations and species. By generating a new genome assembly and re-sequencing 453 genomes, we uncovered the genetic architecture of traits that have been suggested to be important for divergence. Species that differ in monogenic or oligogenic traits that affect ecological performance and/or mate choice show remarkably localized genomic differentiation. By contrast, differentiation among species that have diverged in polygenic traits is genomically widespread and much higher overall, consistent with the evolution of effective and stable genome-wide barriers to gene flow. Thus, we conclude that simple trait architectures are not always as conducive to speciation with gene flow as previously suggested, whereas polygenic architectures can promote rapid and stable speciation in sympatry.
Annotating coding genes and inferring orthologs are two classical challenges in genomics and evolutionary biology that have traditionally been approached separately, limiting scalability. We present TOGA (Tool to infer Orthologs from Genome Alignments), a method that integrates structural gene annotation and orthology inference. TOGA implements a different paradigm to infer orthologous loci, improves ortholog detection and annotation of conserved genes compared with state-of-the-art methods, and handles even highly fragmented assemblies. TOGA scales to hundreds of genomes, which we demonstrate by applying it to 488 placental mammal and 501 bird assemblies, creating the largest comparative gene resources so far. Additionally, TOGA detects gene losses, enables selection screens, and automatically provides a superior measure of mammalian genome quality. TOGA is a powerful and scalable method to annotate and compare genes in the genomic era.
Long noncoding RNAs (lncRNAs) constitute the majority of transcripts in the mammalian genomes, and yet, their functions remain largely unknown. As part of the FANTOM6 project, we systematically knocked down the expression of 285 lncRNAs in human dermal fibroblasts and quantified cellular growth, morphological changes, and transcriptomic responses using Capped Analysis of Gene Expression (CAGE). Antisense oligonucleotides targeting the same lncRNAs exhibited global concordance, and the molecular phenotype, measured by CAGE, recapitulated the observed cellular phenotypes while providing additional insights on the affected genes and pathways. Here, we disseminate the largest-to-date lncRNA knockdown data set with molecular phenotyping (over 1000 CAGE deep-sequencing libraries) for further exploration and highlight functional roles for ZNF213-AS1 and lnc-KHDC3L-2 .