NobleBlocks

Kellogg Biological Station Long Term Ecological Research

facilityKellogg Biological Station, United States

Research output, citation impact, and the most-cited recent papers from Kellogg Biological Station Long Term Ecological Research. Aggregated across the NobleBlocks index of 300M+ scholarly works.

Total works
95
Citations
8.3K
h-index
42
i10-index
92
Also known as
Kellogg Biological Station LTERKellogg Biological Station Long Term Ecological Research

Top-cited papers from Kellogg Biological Station Long Term Ecological Research

Invertebrates, ecosystem services and climate change
Chelse M. Prather, Shannon L. Pelini, Angela Laws, Emily B. Rivest +4 more
2012· Biological reviews/Biological reviews of the Cambridge Philosophical Society344doi:10.1111/brv.12002

The sustainability of ecosystem services depends on a firm understanding of both how organisms provide these services to humans and how these organisms will be altered with a changing climate. Unquestionably a dominant feature of most ecosystems, invertebrates affect many ecosystem services and are also highly responsive to climate change. However, there is still a basic lack of understanding of the direct and indirect paths by which invertebrates influence ecosystem services, as well as how climate change will affect those ecosystem services by altering invertebrate populations. This indicates a lack of communication and collaboration among scientists researching ecosystem services and climate change effects on invertebrates, and land managers and researchers from other disciplines, which becomes obvious when systematically reviewing the literature relevant to invertebrates, ecosystem services, and climate change. To address this issue, we review how invertebrates respond to climate change. We then review how invertebrates both positively and negatively influence ecosystem services. Lastly, we provide some critical future directions for research needs, and suggest ways in which managers, scientists and other researchers may collaborate to tackle the complex issue of sustaining invertebrate-mediated services under a changing climate.

Precipitation control over inorganic nitrogen import–export budgets across watersheds: a synthesis of long‐term ecological research
Evan S. Kane, E. F. Betts, Amy J. Burgin, Hannah M. Clilverd +4 more
2008· Ecohydrology33doi:10.1002/eco.10

Abstract We investigated long‐term and seasonal patterns of N imports and exports, as well as patterns following climate perturbations, across biomes using data from 15 watersheds from nine Long‐Term Ecological Research (LTER) sites in North America. Mean dissolved inorganic nitrogen (DIN) import–export budgets (N import via precipitation–N export via stream flow) for common years across all watersheds was highly variable, ranging from a net loss of − 0·17 ± 0·09 kg N ha−1mo−1 to net retention of 0·68 ± 0·08 kg N ha−1mo−1. The net retention of DIN decreased (smaller import–export budget) with increasing precipitation, as well as with increasing variation in precipitation during the winter, spring, and fall. Averaged across all seasons, net DIN retention decreased as the coefficient of variation (CV) in precipitation increased across all sites (r2 = 0·48, p = 0·005). This trend was made stronger when the disturbed watersheds were withheld from the analysis (r2 = 0·80, p < 0·001, n = 11). Thus, DIN exports were either similar to or exceeded imports in the tropical, boreal, and wet coniferous watersheds, whereas imports exceeded exports in temperate deciduous watersheds. In general, forest harvesting, hurricanes, or floods corresponded with periods of increased DIN exports relative to imports. Periods when water throughput within a watershed was likely to be lower (i.e. low snow pack or El Niño years) corresponded with decreased DIN exports relative to imports. These data provide a basis for ranking diverse sites in terms of their ability to retain DIN in the context of changing precipitation regimes likely to occur in the future. Copyright © 2008 John Wiley & Sons, Ltd.

Using Stakeholder Needs Assessments and Deliberative Dialogue to Inform Climate Change Outreach Efforts
Claire Layman, Julie E. Doll, Cheryl Peters
2013· Journal of Extension15doi:10.34068/joe.51.03.21

Farmers represent a large group of Extension stakeholders who stand to be affected by increased climate variability and change. Yet climate change can be a polarizing topic. In order to be sensitive to this reality, meet stakeholder education needs, and carry out the land-grant mission, we used a participatory decision model known as "deliberation with analysis" to inform climate change programming around agriculture. We designed evaluation tools for each phase of the project. This method strengthened relationships with stakeholders and enabled Michigan State University to move forward with climate change programming.

Surprising effects of cascading higher order interactions
Hsun‐Yi Hsieh, John Vandermeer, Ivette Perfecto
2022· Scientific Reports5doi:10.1038/s41598-022-23763-z

Most species are embedded in multi-interaction networks. Consequently, theories focusing on simple pair-wise interactions cannot predict ecological and/or evolutionary outcomes. This study explores how cascading higher-order interactions (HOIs) would affect the population dynamics of a focal species. Employing a system that involves a myrmecophylic beetle, a parasitic wasp that attacks the beetle, an ant, and a parasitic fly that attacks the ant, the study explores how none, one, and two HOIs affect the parasitism and the sex ratio of the beetle. We conducted mesocosm experiments to examine these HOIs on beetle survival and sex ratio and found that the 1st degree HOI does not change the beetle's survival rate or sex ratio. However, the 2nd degree HOI significantly reduces the beetle's survival rate and changes its sex ratio from even to strongly female-biased. We applied Bayes' theorem to analyze the per capita survival probability of female vs. male beetles and suggested that the unexpected results might arise from complex eco-evolutionary dynamics involved with the 1st and 2nd degree HOIs. Field data suggested the HOIs significantly regulate the sex ratio of the beetle. As the same structure of HOIs appears in other systems, we believe the complexity associated with the 2nd degree HOI would be more common than known and deserve more scientific attention.

The Broken Window: An algorithm for quantifying and characterizing misleading trajectories in ecological processes
Christie A. Bahlai, Easton R. White, Julia D. Perrone, Sarah Cusser +1 more
2020· bioRxiv (Cold Spring Harbor Laboratory)2doi:10.1101/2020.07.07.192211

Abstract A core issue in temporal ecology is the concept of trajectory—that is, when can ecologists have reasonable assurance that they know where a system is going? In this paper, we describe a non-random resampling method to directly address the temporal aspects of scaling ecological observations by leveraging existing data. Findings from long-term research sites have been hugely influential in ecology because of their unprecedented longitudinal perspective, yet short-term studies more consistent with typical grant cycles and graduate programs are still the norm. We use long-term insights to create ‘broken windows,’ that is, reanalyze long-term studies from short-term observational perspectives to examine discontinuities in trends at differing temporal scales. The broken window algorithm connects our observations between the short-term and the long-term with an automated, systematic resampling approach: in short, we repeatedly ‘sample’ moving windows of data from existing long-term time series, and analyze these sampled data as if they represented the entire dataset. We then compile typical statistics used to describe the relationship in the sampled data, through repeated samplings, and then use these derived data to gain insights to the questions: 1) how often are the trends observed in short-term data misleading, and 2) can characteristics of these trends be used to predict our likelihood of being misled? We develop a systematic resampling approach, the ‘broken_window algorithm, and illustrate its utility with a case study of firefly observations produced at the Kellogg Biological Station Long-Term Ecological Research Site (KBS LTER). Through a variety of visualizations, summary statistics, and downstream analyses, we provide a standardized approach to evaluating the trajectory of a system, the amount of observation required to find a meaningful trajectory in similar systems, and a means of evaluating our confidence in our conclusions. Highlights Trends identified in short-term ecology studies can be misleading. Non-random resampling can show how prone different systems are to misleading trends The Broken Window algorithm is a new tool to help synthesize temporal data This tool helps to understand how much data is needed for forecasting to be reliable It can also be used to quantify how likely it is that an observed trend is spurious.

Big Data, Big Changes? A Survey of K-12 Science Teachers in the United States on Which Data Sources and Tools They Use in the Classroom
Joshua M. Rosenberg, Elizabeth H. Schultheis, Melissa K. Kjelvik, Aaron M. Reedy +1 more
20211doi:10.35542/osf.io/tv4zg

The tools that scientists and engineers analyze data are changing—and at the same time, science education standards have shifted to focus on science practices that articulate multiple ways for teachers to support students to make sense of data in science classrooms. Moreover, the types of data and technologies available to teachers and students to support their work with data have advanced. While these changes and features point to the importance of data, practices that relate to data, and the roles of technology, little research has offered a portrait of what teachers presently use. We report on findings from a survey conducted in the United States of 330 science teachers on the data sources, practices, and technologies common to their classroom. We found that teachers predominantly involve their students in analyzing relatively small data sets that they collect. In support of this work, teachers tend to use the technologies that are available to them—namely, calculators and spreadsheets. We discuss what these findings suggest for practice, research, and policy, with an emphasis on supporting teachers based on their needs.

Large-Scale Analysis of Thematic and Geographic Biases in Biodiversity Research Using a Vertebrate Case Study
Aída P. Giozza, Ricardo Santos Magalhães, Marcelle O. Heliópolis, Otávio Augusto Vuolo Marques +1 more
2026· Zenodo (CERN European Organization for Nuclear Research)doi:10.5281/zenodo.20836198

# README — Code and Data Archive ## Overview This repository contains all code and data necessary to reproduce the results reported in the associated manuscript. The analysis pipeline performs a systematic literature screening and classification workflow using machine learning, followed by topic modeling of the accepted articles. --- ## Manuscript Information > **Title:** [Large-Scale Analysis of Thematic and Geographic Biases in Biodiversity Research Using a Vertebrate Case Study ]> **Authors:** [Aída P. Giozza ORCID ID: https://orcid.org/0000-0001-5335-2067, Ricardo Magalhães ORCID ID: https://orcid.org/0000-0001-7477-2191, Marcelle Heliópolis ORCID ID: https://orcid.org/0000-0003-3709-1260, Otavio Marques ORCID ID: https://orcid.org/0000-0002-2830-9558, Luisa Maria Diele-Viegas ORCID ID: https://orcid.org/0000-0002-9225-4678]> **Journal:** The American Naturalist> **Corresponding author:** [Aída P. Giozza, apgiozza@gmail.com]> **Date of deposit:** [24/June/2026]--- ## Repository Structure ```/├── README.md # This file├── data/│ ├── original_data.csv # Full original dataset (~65,000 articles)│ ├── articles_with_scores.csv # Dataset with relevance scores per article│ ├── manually_validated_data.xlsx # Articles manually labeled (accept/reject)│ └── accepted_articles_for_topic_modeling.xlsx # Articles accepted for topic modeling├── models/│ └── optimized_classifier_pipeline.joblib # Trained classification model├── results/│ ├── classified_data.xlsx # Full dataset with ML classification labels│ ├── final_topic_summary.csv # Topic modeling results mapped to categories│ └── classification_report.txt # Model performance report (precision, recall, F1)└── code/ ├── S1_manual_validation_sample.py ├── S2_training_classification_model.py ├── S3_applying_classification_model.py └── S4_bertopic_analysing_themes.py``` --- ## Scripts Description The four scripts must be run **sequentially** in the order listed below. Each script produces an output that serves as the input for the next step. ### S1 — `S1_manual_validation_sample.py` **Purpose:** Creates a random sample of articles from the full dataset for manual validation by human reviewers. **Inputs:**- `original_data.csv` — full article dataset (columns expected: `ID`, `PubYear`, `Title`, `Abstract`) **Outputs:**- An Excel file (`.xlsx`) containing a stratified random sample with a blank `status_manual` column to be filled manually with `accept` or `reject` **Key parameters (edit in the configuration block at the top of the script):**| Parameter | Default | Description ||---|---|---|| `INPUT_PATH` | `"path/to/your/original_data.csv"` | Path to the full dataset || `OUTPUT_PATH` | `"path/to/your/sample_for_manual_validation.xlsx"` | Path for the output sample file || `SAMPLE_SIZE` | `1200` | Number of articles to sample for manual review | **Notes:**- `random_state=42` is set for full reproducibility.- The script also includes a second function (`create_exclusive_sample_v8`) that generates a **second, non-overlapping** validation sample, prioritizing high-scoring articles (score ≥ 8). This is used to expand the training set without re-reviewing already validated articles.- After running S1, **manually fill** the `status_manual` column in the output Excel file with `accept` or `reject` for each article before proceeding to S2. --- ### S2 — `S2_training_classification_model.py` **Purpose:** Trains a text classification model (LinearSVC) using TF-IDF features and optimizes hyperparameters via GridSearchCV. Saves the best model to disk. **Inputs:**- Manually validated Excel file (output of S1, with `status_manual` column completed) **Outputs:**- `optimized_classifier_pipeline.joblib` — the serialized trained model- `classification_report.txt` — precision, recall, F1-score per class- A confusion matrix plot (displayed interactively; save manually if needed) **Key parameters:**| Parameter | Default | Description ||---|---|---|| `VALIDATED_FILE_PATH` | `"path/to/your/validated_data.xlsx"` | Path to manually validated data || `MODEL_OUTPUT_DIR` | `"path/to/your/optimized_model_directory/"` | Directory to save the model || `CV_FOLDS` | `5` | Number of cross-validation folds | **Model details:**- Algorithm: `LinearSVC` with `class_weight='balanced'`- Text features: `TfidfVectorizer` (English stop words removed)- Hyperparameter grid searched: - `ngram_range`: (1,1) or (1,2) - `max_features`: 3,000 or 5,000 - `sublinear_tf`: True or False - `C` (regularization): 0.1, 1, or 10- Optimization metric: `f1_weighted`- Train/test split: 75% / 25% (`random_state=42`) --- ### S3 — `S3_applying_classification_model.py` **Purpose:** Loads the trained model (from S2) and applies it to classify the **full dataset**, producing a predicted label (`accept`/`reject`) for every article. **Inputs:**- `optimized_classifier_pipeline.joblib` — trained model from S2- Full article dataset (`.xlsx`; columns expected: `title`, `abstract`) **Outputs:**- `classified_data.xlsx` — full dataset with an added `ml_classification` column **Key parameters:**| Parameter | Default | Description ||---|---|---|| `MODEL_PATH` | `"path/to/your/optimized_classifier_pipeline.joblib"` | Path to trained model || `INPUT_DATA_PATH` | `"path/to/your/input_data.xlsx"` | Path to full dataset || `OUTPUT_DATA_PATH` | `"path/to/your/classified_data.xlsx"` | Path for classified output | --- ### S4 — `S4_bertopic_analysing_themes.py` **Purpose:** Performs unsupervised topic modeling (BERTopic) on articles classified as `accept` by the ML model. Maps the resulting topics to user-defined thematic categories using cosine similarity of sentence embeddings. **Inputs:**- `accepted_articles_for_topic_modeling.xlsx` — articles accepted by the classifier (from S3), filtered for `classificacao_ML == 'accept'`; columns expected: `title`, `abstract` **Outputs:**- `final_topic_summary.csv` — each BERTopic topic with its ID, top keywords, document count, and the best-fit user-defined category and similarity score **Key parameters:**| Parameter | Default | Description ||---|---|---|| `ACCEPTED_ARTICLES_PATH` | `"path/to/your/accepted_articles_for_topic_modeling.xlsx"` | Path to accepted articles || `RESULTS_DIR` | `"path/to/your/topic_analysis_results/"` | Directory for output files || `DESIRED_NUM_TOPICS` | `40` | Target number of topics after reduction || `SIMILARITY_THRESHOLD` | `0.35` | Minimum cosine similarity to assign a category; topics below this are labeled "Other" | **User-defined categories:** The `user_defined_categories` dictionary in the script must be populated with domain-specific category names and associated keywords before running. See the commented example in the script for guidance. **Models used:**- Sentence embeddings: `all-MiniLM-L6-v2` (via `sentence-transformers`)- Topic modeling: `BERTopic` --- ## Analysis Pipeline Summary ```Raw literature database (~65,000 articles) │ ▼[S1] Random sample → Manual review (accept/reject) │ ▼[S2] Train LinearSVC classifier → Optimized model (.joblib) │ ▼[S3] Apply model to full dataset → Classified dataset (.xlsx) │ ▼[S4] BERTopic on accepted articles → Topic summary (.csv)``` --- ## Software Requirements | Package | Version tested | Purpose ||---|---|---|| Python | ≥ 3.9 | Runtime || pandas | ≥ 1.5 | Data manipulation || scikit-learn | ≥ 1.2 | ML model, GridSearchCV, TF-IDF || joblib | ≥ 1.2 | Model serialization || openpyxl | ≥ 3.0 | Excel file I/O || bertopic | ≥ 0.15 | Topic modeling || sentence-transformers | ≥ 2.2 | Sentence embeddings || seaborn | ≥ 0.12 | Confusion matrix visualization || matplotlib | ≥ 3.6 | Plotting || numpy | ≥ 1.23 | Numerical operations | ### Installation ```bashpip install pandas scikit-learn joblib openpyxl bertopic sentence-transformers seaborn matplotlib numpy``` Or using the provided environment file (if included): ```bashconda env create -f environment.ymlconda activate [env-name]``` --- ## Reproducibility Notes - All random operations use `random_state=42` to ensure reproducibility.- The `GridSearchCV` in S2 uses `n_jobs=-1` (all available CPU cores); results are deterministic regardless of the number of cores used, given the fixed random state.- BERTopic results (S4) may vary slightly across hardware and library versions due to the stochastic nature of the underlying UMAP dimensionality reduction. We recommend using the exact package versions listed above.- The manually validated data (`manually_validated_data.xlsx`) is provided in full so that the model training step (S2) can be reproduced without repeating manual annotation. --- ## Data Availability The original literature dataset was compiled from [Web of Science and Scopus]. --- ## License Code is released under the [MIT License / CC BY 4.0 / insert your license]. See `LICENSE` file for details. --- ## Contact For questions regarding the code or data, please contact the corresponding author.

Large-Scale Analysis of Thematic and Geographic Biases in Biodiversity Research Using a Vertebrate Case Study
Aída P. Giozza, Ricardo Santos Magalhães, Marcelle O. Heliópolis, Otávio Augusto Vuolo Marques +1 more
2026· Zenodo (CERN European Organization for Nuclear Research)doi:10.5281/zenodo.20836199

# README — Code and Data Archive ## Overview This repository contains all code and data necessary to reproduce the results reported in the associated manuscript. The analysis pipeline performs a systematic literature screening and classification workflow using machine learning, followed by topic modeling of the accepted articles. --- ## Manuscript Information > **Title:** [Large-Scale Analysis of Thematic and Geographic Biases in Biodiversity Research Using a Vertebrate Case Study ]> **Authors:** [Aída P. Giozza ORCID ID: https://orcid.org/0000-0001-5335-2067, Ricardo Magalhães ORCID ID: https://orcid.org/0000-0001-7477-2191, Marcelle Heliópolis ORCID ID: https://orcid.org/0000-0003-3709-1260, Otavio Marques ORCID ID: https://orcid.org/0000-0002-2830-9558, Luisa Maria Diele-Viegas ORCID ID: https://orcid.org/0000-0002-9225-4678]> **Journal:** The American Naturalist> **Corresponding author:** [Aída P. Giozza, apgiozza@gmail.com]> **Date of deposit:** [24/June/2026]--- ## Repository Structure ```/├── README.md # This file├── data/│ ├── original_data.csv # Full original dataset (~65,000 articles)│ ├── articles_with_scores.csv # Dataset with relevance scores per article│ ├── manually_validated_data.xlsx # Articles manually labeled (accept/reject)│ └── accepted_articles_for_topic_modeling.xlsx # Articles accepted for topic modeling├── models/│ └── optimized_classifier_pipeline.joblib # Trained classification model├── results/│ ├── classified_data.xlsx # Full dataset with ML classification labels│ ├── final_topic_summary.csv # Topic modeling results mapped to categories│ └── classification_report.txt # Model performance report (precision, recall, F1)└── code/ ├── S1_manual_validation_sample.py ├── S2_training_classification_model.py ├── S3_applying_classification_model.py └── S4_bertopic_analysing_themes.py``` --- ## Scripts Description The four scripts must be run **sequentially** in the order listed below. Each script produces an output that serves as the input for the next step. ### S1 — `S1_manual_validation_sample.py` **Purpose:** Creates a random sample of articles from the full dataset for manual validation by human reviewers. **Inputs:**- `original_data.csv` — full article dataset (columns expected: `ID`, `PubYear`, `Title`, `Abstract`) **Outputs:**- An Excel file (`.xlsx`) containing a stratified random sample with a blank `status_manual` column to be filled manually with `accept` or `reject` **Key parameters (edit in the configuration block at the top of the script):**| Parameter | Default | Description ||---|---|---|| `INPUT_PATH` | `"path/to/your/original_data.csv"` | Path to the full dataset || `OUTPUT_PATH` | `"path/to/your/sample_for_manual_validation.xlsx"` | Path for the output sample file || `SAMPLE_SIZE` | `1200` | Number of articles to sample for manual review | **Notes:**- `random_state=42` is set for full reproducibility.- The script also includes a second function (`create_exclusive_sample_v8`) that generates a **second, non-overlapping** validation sample, prioritizing high-scoring articles (score ≥ 8). This is used to expand the training set without re-reviewing already validated articles.- After running S1, **manually fill** the `status_manual` column in the output Excel file with `accept` or `reject` for each article before proceeding to S2. --- ### S2 — `S2_training_classification_model.py` **Purpose:** Trains a text classification model (LinearSVC) using TF-IDF features and optimizes hyperparameters via GridSearchCV. Saves the best model to disk. **Inputs:**- Manually validated Excel file (output of S1, with `status_manual` column completed) **Outputs:**- `optimized_classifier_pipeline.joblib` — the serialized trained model- `classification_report.txt` — precision, recall, F1-score per class- A confusion matrix plot (displayed interactively; save manually if needed) **Key parameters:**| Parameter | Default | Description ||---|---|---|| `VALIDATED_FILE_PATH` | `"path/to/your/validated_data.xlsx"` | Path to manually validated data || `MODEL_OUTPUT_DIR` | `"path/to/your/optimized_model_directory/"` | Directory to save the model || `CV_FOLDS` | `5` | Number of cross-validation folds | **Model details:**- Algorithm: `LinearSVC` with `class_weight='balanced'`- Text features: `TfidfVectorizer` (English stop words removed)- Hyperparameter grid searched: - `ngram_range`: (1,1) or (1,2) - `max_features`: 3,000 or 5,000 - `sublinear_tf`: True or False - `C` (regularization): 0.1, 1, or 10- Optimization metric: `f1_weighted`- Train/test split: 75% / 25% (`random_state=42`) --- ### S3 — `S3_applying_classification_model.py` **Purpose:** Loads the trained model (from S2) and applies it to classify the **full dataset**, producing a predicted label (`accept`/`reject`) for every article. **Inputs:**- `optimized_classifier_pipeline.joblib` — trained model from S2- Full article dataset (`.xlsx`; columns expected: `title`, `abstract`) **Outputs:**- `classified_data.xlsx` — full dataset with an added `ml_classification` column **Key parameters:**| Parameter | Default | Description ||---|---|---|| `MODEL_PATH` | `"path/to/your/optimized_classifier_pipeline.joblib"` | Path to trained model || `INPUT_DATA_PATH` | `"path/to/your/input_data.xlsx"` | Path to full dataset || `OUTPUT_DATA_PATH` | `"path/to/your/classified_data.xlsx"` | Path for classified output | --- ### S4 — `S4_bertopic_analysing_themes.py` **Purpose:** Performs unsupervised topic modeling (BERTopic) on articles classified as `accept` by the ML model. Maps the resulting topics to user-defined thematic categories using cosine similarity of sentence embeddings. **Inputs:**- `accepted_articles_for_topic_modeling.xlsx` — articles accepted by the classifier (from S3), filtered for `classificacao_ML == 'accept'`; columns expected: `title`, `abstract` **Outputs:**- `final_topic_summary.csv` — each BERTopic topic with its ID, top keywords, document count, and the best-fit user-defined category and similarity score **Key parameters:**| Parameter | Default | Description ||---|---|---|| `ACCEPTED_ARTICLES_PATH` | `"path/to/your/accepted_articles_for_topic_modeling.xlsx"` | Path to accepted articles || `RESULTS_DIR` | `"path/to/your/topic_analysis_results/"` | Directory for output files || `DESIRED_NUM_TOPICS` | `40` | Target number of topics after reduction || `SIMILARITY_THRESHOLD` | `0.35` | Minimum cosine similarity to assign a category; topics below this are labeled "Other" | **User-defined categories:** The `user_defined_categories` dictionary in the script must be populated with domain-specific category names and associated keywords before running. See the commented example in the script for guidance. **Models used:**- Sentence embeddings: `all-MiniLM-L6-v2` (via `sentence-transformers`)- Topic modeling: `BERTopic` --- ## Analysis Pipeline Summary ```Raw literature database (~65,000 articles) │ ▼[S1] Random sample → Manual review (accept/reject) │ ▼[S2] Train LinearSVC classifier → Optimized model (.joblib) │ ▼[S3] Apply model to full dataset → Classified dataset (.xlsx) │ ▼[S4] BERTopic on accepted articles → Topic summary (.csv)``` --- ## Software Requirements | Package | Version tested | Purpose ||---|---|---|| Python | ≥ 3.9 | Runtime || pandas | ≥ 1.5 | Data manipulation || scikit-learn | ≥ 1.2 | ML model, GridSearchCV, TF-IDF || joblib | ≥ 1.2 | Model serialization || openpyxl | ≥ 3.0 | Excel file I/O || bertopic | ≥ 0.15 | Topic modeling || sentence-transformers | ≥ 2.2 | Sentence embeddings || seaborn | ≥ 0.12 | Confusion matrix visualization || matplotlib | ≥ 3.6 | Plotting || numpy | ≥ 1.23 | Numerical operations | ### Installation ```bashpip install pandas scikit-learn joblib openpyxl bertopic sentence-transformers seaborn matplotlib numpy``` Or using the provided environment file (if included): ```bashconda env create -f environment.ymlconda activate [env-name]``` --- ## Reproducibility Notes - All random operations use `random_state=42` to ensure reproducibility.- The `GridSearchCV` in S2 uses `n_jobs=-1` (all available CPU cores); results are deterministic regardless of the number of cores used, given the fixed random state.- BERTopic results (S4) may vary slightly across hardware and library versions due to the stochastic nature of the underlying UMAP dimensionality reduction. We recommend using the exact package versions listed above.- The manually validated data (`manually_validated_data.xlsx`) is provided in full so that the model training step (S2) can be reproduced without repeating manual annotation. --- ## Data Availability The original literature dataset was compiled from [Web of Science and Scopus]. --- ## License Code is released under the [MIT License / CC BY 4.0 / insert your license]. See `LICENSE` file for details. --- ## Contact For questions regarding the code or data, please contact the corresponding author.

Climate change across the air-water interface affects giant salmonfly (Pteronarcys californica) emergence timing and adult lifespan
Lindsey Albertson, Alzada Roche, Alisha Shah, Christine Verhille
2026· DRYADdoi:10.5061/dryad.pg4f4qs4z

Aquatic insects experience complex temperature regimes, including during the vulnerable transition from aquatic to terrestrial environments as they emerge as adults. However, rising temperatures in montane environments across the globe are causing a novel thermal regime. Earlier snow-melt has not yet changed the narrow range of cold spring water temperatures, but both water and air temperatures have been rising in the summer. In southwestern Montana, USA, spring water temperature cues large, synchronous emergence of giant salmonflies (Pteronarcys californica) in early summer, but it is unknown how variable and warmer temperatures that occur after the springtime cue will affect life-history traits. We experimentally tested how changing temperatures during the 6 weeks before and after emergence influenced emergence timing, emergence success, and adult lifespan.

Surprising effects of cascading higher order interactions
Hsun‐Yi Hsieh, John Vandermeer, Ivette Perfecto
2022· Research Squaredoi:10.21203/rs.3.rs-1926117/v2

Abstract Most species are embedded in multi-interaction networks. Consequently, theories focusing on simple pair-wise interactions cannot predict ecological and/or evolutionary outcomes. This study explores how cascading higher-order interactions (HOIs) would affect the population dynamics of a focal species. Employing a system that involves a myrmecophylic beetle, a parasitic wasp that attacks the beetle, an ant, and a parasitic fly that attacks the ant, the study explores how none, one, and two HOIs affect the parasitism and the sex ratio of the beetle. We conducted mesocosm experiments to examine these HOIs on beetle survival and sex ratio and found that the 1st degree HOI does not change the beetle’s survival rate or sex ratio. However, the 2nd degree HOI significantly reduces the beetle’s survival rate and changes its sex ratio from even to strongly female-biased. We applied Bayes’ theorem to analyze the per capita survival probability of female vs. male beetles and suggested that the unexpected results might arise from complex eco-evolutionary dynamics involved with the 1st and 2nd degree HOIs. Field data suggested the HOIs significantly regulate the sex ratio of the beetle. As the same structure of HOIs appears in other systems, we believe the complexity associated with the 2nd degree HOI would be more common than known and deserve more scientific attention.