
Statistical and Applied Mathematical Sciences Institute
facilityDurham, United States
Research output, citation impact, and the most-cited recent papers from Statistical and Applied Mathematical Sciences Institute (United States). Aggregated across the NobleBlocks index of 300M+ scholarly works.
Top-cited papers from Statistical and Applied Mathematical Sciences Institute
Fiber tract trajectories in coherently organized brain white matter pathways were computed from in vivo diffusion tensor magnetic resonance imaging (DT-MRI) data. First, a continuous diffusion tensor field is constructed from this discrete, noisy, measured DT-MRI data. Then a Frenet equation, describing the evolution of a fiber tract, was solved. This approach was validated using synthesized, noisy DT-MRI data. Corpus callosum and pyramidal tract trajectories were constructed and found to be consistent with known anatomy. The method's reliability, however, degrades where the distribution of fiber tract directions is nonuniform. Moreover, background noise in diffusion-weighted MRIs can cause a computed trajectory to hop from tract to tract. Still, this method can provide quantitative information with which to visualize and study connectivity and continuity of neural pathways in the central and peripheral nervous systems in vivo, and holds promise for elucidating architectural features in other fibrous tissues and ordered media.
We propose a new data-augmentation strategy for fully Bayesian inference in models with binomial likelihoods. The approach appeals to a new class of Pólya–Gamma distributions, which are constructed in detail. A variety of examples are presented to show the versatility of the method, including logistic regression, negative binomial regression, nonlinear mixed-effect models, and spatial models for count data. In each case, our data-augmentation strategy leads to simple, effective methods for posterior inference that (1) circumvent the need for analytic approximations, numerical integration, or Metropolis–Hastings; and (2) outperform other known data-augmentation strategies, both in ease of use and in computational efficiency. All methods, including an efficient sampler for the Pólya–Gamma distribution, are implemented in the R package BayesLogit. Supplementary materials for this article are available online.
Bayesian statistical practice makes extensive use of versions of objective Bayesian analysis. We discuss why this is so, and address some of the criticisms that have been raised concerning objective Bayesian analysis. The dangers of treating the issue too casually are also considered. In particular, we suggest that the statistical community should accept formal objective Bayesian techniques with confidence, but should be more cautious about casual objective Bayesian techniques.
Abstract Tree species are expected to track warming climate by shifting their ranges to higher latitudes or elevations, but current evidence of latitudinal range shifts for suites of species is largely indirect. In response to global warming, offspring of trees are predicted to have ranges extend beyond adults at leading edges and the opposite relationship at trailing edges. Large‐scale forest inventory data provide an opportunity to compare present latitudes of seedlings and adult trees at their range limits. Using the USDA Forest Service's Forest Inventory and Analysis data, we directly compared seedling and tree 5th and 95th percentile latitudes for 92 species in 30 longitudinal bands for 43 334 plots across the eastern United States. We further compared these latitudes with 20th century temperature and precipitation change and functional traits, including seed size and seed spread rate. Results suggest that 58.7% of the tree species examined show the pattern expected for a population undergoing range contraction, rather than expansion, at both northern and southern boundaries. Fewer species show a pattern consistent with a northward shift (20.7%) and fewer still with a southward shift (16.3%). Only 4.3% are consistent with expansion at both range limits. When compared with the 20th century climate changes that have occurred at the range boundaries themselves, there is no consistent evidence that population spread is greatest in areas where climate has changed most; nor are patterns related to seed size or dispersal characteristics. The fact that the majority of seedling extreme latitudes are less than those for adult trees may emphasize the lack of evidence for climate‐mediated migration, and should increase concerns for the risks posed by climate change.
Abstract A large array of species distribution model ( SDM ) approaches has been developed for explaining and predicting the occurrences of individual species or species assemblages. Given the wealth of existing models, it is unclear which models perform best for interpolation or extrapolation of existing data sets, particularly when one is concerned with species assemblages. We compared the predictive performance of 33 variants of 15 widely applied and recently emerged SDM s in the context of multispecies data, including both joint SDM s that model multiple species together, and stacked SDM s that model each species individually combining the predictions afterward. We offer a comprehensive evaluation of these SDM approaches by examining their performance in predicting withheld empirical validation data of different sizes representing five different taxonomic groups, and for prediction tasks related to both interpolation and extrapolation. We measure predictive performance by 12 measures of accuracy, discrimination power, calibration, and precision of predictions, for the biological levels of species occurrence, species richness, and community composition. Our results show large variation among the models in their predictive performance, especially for communities comprising many species that are rare. The results do not reveal any major trade‐offs among measures of model performance; the same models performed generally well in terms of accuracy, discrimination, and calibration, and for the biological levels of individual species, species richness, and community composition. In contrast, the models that gave the most precise predictions were not well calibrated, suggesting that poorly performing models can make overconfident predictions. However, none of the models performed well for all prediction tasks. As a general strategy, we therefore propose that researchers fit a small set of models showing complementary performance, and then apply a cross‐validation procedure involving separate data to establish which of these models performs best for the goal of the study.
Diffuse large B-cell lymphoma (DLBCL) is the most common form of lymphoma in adults. The disease exhibits a striking heterogeneity in gene expression profiles and clinical outcomes, but its genetic causes remain to be fully defined. Through whole genome and exome sequencing, we characterized the genetic diversity of DLBCL. In all, we sequenced 73 DLBCL primary tumors (34 with matched normal DNA). Separately, we sequenced the exomes of 21 DLBCL cell lines. We identified 322 DLBCL cancer genes that were recurrently mutated in primary DLBCLs. We identified recurrent mutations implicating a number of known and not previously identified genes and pathways in DLBCL including those related to chromatin modification (ARID1A and MEF2B), NF-κB (CARD11 and TNFAIP3), PI3 kinase (PIK3CD, PIK3R1, and MTOR), B-cell lineage (IRF8, POU2F2, and GNA13), and WNT signaling (WIF1). We also experimentally validated a mutation in PIK3CD, a gene not previously implicated in lymphomas. The patterns of mutation demonstrated a classic long tail distribution with substantial variation of mutated genes from patient to patient and also between published studies. Thus, our study reveals the tremendous genetic heterogeneity that underlies lymphomas and highlights the need for personalized medicine approaches to treating these patients.
A new approach to factor analysis and related latent variable methods is proposed which is based on data reduction using the idea of Bayesian sufficiency. Considerations of symmetry, invariance and independence are used to determine an appropriate family of models. The results are expressed in terms of linear functions of the manifest variables after the manner of principal components analysis. The approach justifies some of the practices based on the normal theory factor model and lays a foundation for the treatment of nonnormal, including categorical, variables.
Imputing genotypes from reference panels created by whole-genome sequencing (WGS) provides a cost-effective strategy for augmenting the single-nucleotide polymorphism (SNP) content of genome-wide arrays. The UK10K Cohorts project has generated a data set of 3,781 whole genomes sequenced at low depth (average 7x), aiming to exhaustively characterize genetic variation down to 0.1% minor allele frequency in the British population. Here we demonstrate the value of this resource for improving imputation accuracy at rare and low-frequency variants in both a UK and an Italian population. We show that large increases in imputation accuracy can be achieved by re-phasing WGS reference panels after initial genotype calling. We also present a method for combining WGS panels to improve variant coverage and downstream imputation accuracy, which we illustrate by integrating 7,562 WGS haplotypes from the UK10K project with 2,184 haplotypes from the 1000 Genomes Project. Finally, we introduce a novel approximation that maintains speed without sacrificing imputation accuracy for rare variants.
Current state-of-the-art diagnostic measures of Alzheimer's disease (AD) are invasive (cerebrospinal fluid analysis), expensive (neuroimaging) and time-consuming (neuropsychological assessment) and thus have limited accessibility as frontline screening and diagnostic tools for AD. Thus, there is an increasing need for additional noninvasive and/or cost-effective tools, allowing identification of subjects in the preclinical or early clinical stages of AD who could be suitable for further cognitive evaluation and dementia diagnostics. Implementation of such tests may facilitate early and potentially more effective therapeutic and preventative strategies for AD. Before applying them in clinical practice, these tools should be examined in ongoing large clinical trials. This review will summarize and highlight the most promising screening tools including neuropsychometric, clinical, blood, and neurophysiological tests.
BACKGROUND/GOALS: Endoscopic ultrasound (EUS)-guided celiac plexus block (CPB) and celiac plexus neurolysis (CPN) have become important interventions in the management of pain due to chronic pancreatitis and pancreatic cancer. However, only a few well-structured studies have been performed to evaluate their efficacy. Given limited data, their use remains controversial. Herein, we evaluate the efficacy of EUS-guided CPB and CPN in alleviating chronic abdominal pain due to chronic pancreatitis and pancreatic cancer respectively. STUDY METHODS: Using Medline, Pubmed, and Embase databases from January 1966 through December 2007, a thorough search of the English literature for studies evaluating the efficacy of EUS-guided CPB and CPN for the management of chronic abdominal pain due to chronic pancreatitis and pancreatic cancer was conducted, along with a hand search of reference lists. Studies that involved less than 10 patients were excluded. Data on pain relief was extracted, pooled, and analyzed. RESULTS: A total of 9 studies were included in the final analysis. For chronic pancreatitis, 6 relevant studies were identified, comprising a total of 221 patients. EUS-guided CPB was effective in alleviating abdominal pain in 51.46% of patients. For pancreatic cancer, 5 relevant studies were identified with a total of 119 patients. EUS-guided CPN was effective in alleviating abdominal pain in 72.54% of patients. CONCLUSIONS: EUS-guided CPB was 51.46% effective in managing chronic abdominal pain in patients with chronic pancreatitis, but warrants improvement in patient selection and refinement of technique, whereas EUS-guided CPN was 72.54% effective in managing pain due to pancreatic cancer and is a reasonable option for patients with tolerance to narcotic analgesics.
Risk sharing arrangements between hospitals and payers together with penalties imposed by the Centers for Medicare and Medicaid (CMS) are driving an interest in decreasing early readmissions. There are a number of published risk models predicting 30day readmissions for particular patient populations, however they often exhibit poor predictive performance and would be unsuitable for use in a clinical setting. In this work we describe and compare several predictive models, some of which have never been applied to this task and which outperform the regression methods that are typically applied in the healthcare literature. In addition, we apply methods from deep learning to the five conditions CMS is using to penalize hospitals, and offer a simple framework for determining which conditions are most cost effective to target.
Abstract Probabilistic forecasts of species distribution and abundance require models that accommodate the range of ecological data, including a joint distribution of multiple species based on combinations of continuous and discrete observations, mostly zeros. We develop a generalized joint attribute model ( GJAM ), a probabilistic framework that readily applies to data that are combinations of presence‐absence, ordinal, continuous, discrete, composition, zero‐inflated, and censored. It does so as a joint distribution over all species providing inference on sensitivity to input variables, correlations between species on the data scale, prediction, sensitivity analysis, definition of community structure, and missing data imputation. GJAM applications illustrate flexibility to the range of species‐abundance data. Applications to forest inventories demonstrate species relationships responding as a community to environmental variables. It shows that the environment can be inverse predicted from the joint distribution of species. Application to microbiome data demonstrates how inverse prediction in the GJAM framework accelerates variable selection, by isolating effects of each input variable's influence across all species.
Propensity score methods are being increasingly used as a less parametric alternative to traditional regression to balance observed differences across groups in both descriptive and causal comparisons. Data collected in many disciplines often have analytically relevant multilevel or clustered structure. The propensity score, however, was developed and has been used primarily with unstructured data. We present and compare several propensity-score-weighted estimators for clustered data, including marginal, cluster-weighted, and doubly robust estimators. Using both analytical derivations and Monte Carlo simulations, we illustrate bias arising when the usual assumptions of propensity score analysis do not hold for multilevel data. We show that exploiting the multilevel structure, either parametrically or nonparametrically, in at least one stage of the propensity score analysis can greatly reduce these biases. We applied these methods to a study of racial disparities in breast cancer screening among beneficiaries of Medicare health plans.
This paper shows that the space of persistence diagrams has properties that allow for the definition of probability measures which support expectations, variances, percentiles and conditional probabilities. This provides a theoretical basis for a statistical treatment of persistence diagrams, for example computing sample averages and sample variances of persistence diagrams. We first prove that the space of persistence diagrams with the Wasserstein metric is complete and separable. We then prove a simple criterion for compactness in this space. These facts allow us to show the existence of the standard statistical objects needed to extend the theory of topological persistence to a much larger set of applications.
The HPN-DREAM community challenge assessed the ability of computational methods to infer causal molecular networks, focusing specifically on the task of inferring causal protein signaling networks in cancer cell lines. It remains unclear whether causal, rather than merely correlational, relationships in molecular networks can be inferred in complex biological settings. Here we describe the HPN-DREAM network inference challenge, which focused on learning causal influences in signaling networks. We used phosphoprotein data from cancer cell lines as well as in silico data from a nonlinear dynamical model. Using the phosphoprotein data, we scored more than 2,000 networks submitted by challenge participants. The networks spanned 32 biological contexts and were scored in terms of causal validity with respect to unseen interventional data. A number of approaches were effective, and incorporating known biology was generally advantageous. Additional sub-challenges considered time-course prediction and visualization. Our results suggest that learning causal relationships may be feasible in complex settings such as disease states. Furthermore, our scoring approach provides a practical way to empirically assess inferred molecular networks in a causal sense.
-like tail behavior. Bayesian computation is straightforward via a simple Gibbs sampling algorithm. We investigate the properties of the maximum a posteriori estimator, as sparse estimation plays an important role in many problems, reveal connections with some well-established regularization procedures, and show some asymptotic results. The performance of the prior is tested through simulations and an application.
OBJECTIVE: To test the effects of maternal periodontal disease treatment on the incidence of preterm birth (delivery before 37 weeks of gestation). METHODS: The Maternal Oral Therapy to Reduce Obstetric Risk Study was a randomized, treatment-masked, controlled clinical trial of pregnant women with periodontal disease who were receiving standard obstetric care. Participants were assigned to either a periodontal treatment arm, consisting of scaling and root planing early in the second trimester, or a delayed treatment arm that provided periodontal care after delivery. Pregnancy and maternal periodontal status were followed to delivery and neonatal outcomes until discharge. The primary outcome (gestational age less than 37 weeks) and the secondary outcome (gestational age less than 35 weeks) were analyzed using a chi test of equality of two proportions. RESULTS: The study randomized 1,806 patients at three performance sites and completed 1,760 evaluable patients. At baseline, there were no differences comparing the treatment and control arms for any of the periodontal or obstetric measures. The rate of preterm delivery for the treatment group was 13.1% and 11.5% for the control group (P=.316). There were no significant differences when comparing women in the treatment group with those in the control group with regard to the adverse event rate or the major obstetric and neonatal outcomes. CONCLUSION: Periodontal therapy did not reduce the incidence of preterm delivery. LEVEL OF EVIDENCE: I.
This paper reports an analysis and comparison of the use of 51 different similarity coefficients for computing the similarities between binary fingerprints for both simulated and real chemical data sets. Five pairs and a triplet of coefficients were found to yield identical similarity values, leading to the elimination of seven of the coefficients. The remaining 44 coefficients were then compared in two ways: by their theoretical characteristics using simple descriptive statistics, correlation analysis, multidimensional scaling, Hasse diagrams, and the recently described atemporal target diffusion model; and by their effectiveness for similarity-based virtual screening using MDDR, WOMBAT, and MUV data. The comparisons demonstrate the general utility of the well-known Tanimoto method but also suggest other coefficients that may be worthy of further attention.
Summary Joint species distribution models ( JSDM ) are increasingly used to analyse community ecology data. Recent progress with JSDM s has provided ecologists with new tools for estimating species associations (residual co‐occurrence patterns after accounting for environmental niches) from large data sets, as well as for increasing the predictive power of species distribution models ( SDM s) by accounting for such associations. Yet, one critical limitation of JSDM s developed thus far is that they assume constant species associations. However, in real ecological communities, the direction and strength of interspecific interactions are likely to be different under different environmental conditions. In this paper, we overcome the shortcoming of present JSDM s by allowing species associations covary with measured environmental covariates. To estimate environmental‐dependent species associations, we utilize a latent variable structure, where the factor loadings are modelled as a linear regression to environmental covariates. We illustrate the performance of the statistical framework with both simulated and real data. Our results show that JSDM s perform substantially better in inferring environmental‐dependent species associations than single SDM s, especially with sparse data. Furthermore, JSDM s consistently overperform SDM s in terms of predictive power for generating predictions that account for environment‐dependent biotic associations. We implemented the statistical framework as a MATLAB package, which includes tools both for model parameterization as well as for post‐processing of results, particularly for addressing whether and how species associations depend on the environmental conditions. Our statistical framework provides a new tool for ecologists who wish to investigate from non‐manipulative observational community data the dependency of interspecific interactions on environmental context. Our method can be applied to answer the fundamental questions in community ecology about how species’ interactions shift in changing environmental conditions, as well as to predict future changes of species’ interactions in response to global change.
In 2005, Holy and Guo advanced the idea that male mice produce ultrasonic vocalizations (USV) with some features similar to courtship songs of songbirds. Since then, studies showed that male mice emit USV songs in different contexts (sexual and other) and possess a multisyllabic repertoire. Debate still exists for and against plasticity in their vocalizations. But the use of a multisyllabic repertoire can increase potential flexibility and information, in how elements are organized and recombined, namely syntax. In many bird species, modulating song syntax has ethological relevance for sexual behavior and mate preferences. In this study we exposed adult male mice to different social contexts and developed a new approach of analyzing their USVs based on songbird syntax analysis. We found that male mice modify their syntax, including specific sequences, length of sequence, repertoire composition, and spectral features, according to stimulus and social context. Males emit longer and simpler syllables and sequences when singing to females, but more complex syllables and sequences in response to fresh female urine. Playback experiments show that the females prefer the complex songs over the simpler ones. We propose the complex songs are to lure females in, whereas the directed simpler sequences are used for direct courtship. These results suggest that although mice have a much more limited ability of song modification, they could still be used as animal models for understanding some vocal communication features that songbirds are used for.