This dataset contains normalized RNA-seq gene expression counts from Synechococcus sp. WH8102 grown under controlled laboratory phosphorus treatments (High-P, Mid-P, and Low-P) in 2025. Gene expression was measured across three biological replicates per treatment and includes 2,718 genes with locus_tag identifiers, along with gene names and functional product annotations where available. Normalized counts were generated using DESeq2 (median-of-ratios size factor method), enabling comparison acro...
Views
Downloads
These numbers come from web analytics and reflect real user activity on the site. Download counts include both page-based interactions and direct file downloads. They reliably show dataset usage and are mostly free of bot traffic.
RNA samples were collected from Synechococcus sp. WH8102 cultures grown under continuous laboratory phosphorus treatments (High-P, Mid-P, and Low-P), with three biological replicates per treatment. On the final day of the experiment, ~700 mL of each culture was filtered through 47 mm, 0.4 μm GFF filters (previously combusted at 450°C) using a peristaltic pump and stored at -80°C until analysis.
Total RNA was extracted from cells collected on filters using a modified TRIzol-based extraction protocol. Filters were transferred aseptically into screw-cap bead tubes containing 0.1 mm disruptor beads and kept on ice after removal from −80 °C storage. Cells were enzymatically lysed in lysis buffer (30 mM Tris, 10 mM EDTA, and 10 mg mL⁻¹ lysozyme) and incubated at 37 °C for 30 min with slow rotation. Samples were subsequently bead-beaten by vortexing at maximum speed for 5 min to ensure mechanical disruption. Following lysis, RNA was extracted using TRIzol Reagent (Thermo Fisher Scientific, Waltham, MA, USA). TRIzol and chloroform were added to each sample, followed by vigorous mixing and incubation at room temperature prior to centrifugation to separate phases. The aqueous phase containing RNA was transferred to a new tube and RNA was precipitated with saline solution and ice-cold isopropanol. Samples were incubated at −20 °C and centrifuged to pellet RNA. The RNA pellet was washed with 70% ethanol, briefly air-dried, and resuspended in nuclease-free water. To remove contaminating DNA, RNA samples were treated with TURBO DNase (Thermo Fisher Scientific) according to the manufacturer’s instructions. A subsequent phenol–chloroform–isoamyl alcohol purification step was performed, followed by chloroform extraction and ethanol precipitation to further purify RNA. The final RNA pellet was washed with 70% ethanol, air-dried, and resuspended in nuclease-free water. RNA concentration and purity were assessed using NanoDrop Spectrophotometer (Thermo Fisher Scientific) and Qubit Fluorometer (Thermo Fisher Scientific).
Total RNA samples were submitted to SeqCoast for library preparation and sequencing. Ribosomal RNA (rRNA) was depleted and sequencing libraries were prepared using the Illumina Stranded Total RNA Prep Ligation Kit with Ribo-Zero Plus Microbiome (Illumina, #20072063) and Illumina Unique Dual Indexes, following the manufacturer’s protocol. Sequencing was performed on an Illumina NextSeq 2000 platform using a 300-cycle XLEAP-SBS flow cell to generate 2 × 150 bp paired-end reads. A 1–2% PhiX spike-in control was included to support optimal base calling. Base calling, demultiplexing, adapter trimming, and initial run quality control were performed using DRAGEN (v4.2.7; Illumina) onboard the NextSeq 2000 system. Sequencing quality was assessed at both the run level and per-sample level, including inspection of FastQC reports. FastQ files corresponding to paired-end reads were generated for downstream analysis.
The reads were mapped to the Synechococcus sp. WH8102 genome obtained from the JGI Genome Portal using Bowtie2 v.2.4.1 (Langmead et al., 2012) with default parameters in paired-end mode. Counts were generated with the featureCounts function in the Rsubread v.2.12.3 R Bioconductor package (Liao et al., 2019).
Differential gene expression analysis was performed using DESeq2 (v1.42.1). Raw read counts were imported into a DESeqDataSet object with experimental conditions (Low-P, Mid-P, High-P) specified as the design factor. Size-factor normalization was applied using the median-of-ratios method to account for differences in sequencing depth, and gene-wise dispersion estimates were fitted under a negative binomial generalized linear model framework. Two complementary DESeq2 approaches were used to assess differential expression. First, a likelihood ratio test (LRT) was applied to compare a full model including the condition effect against a reduced model excluding it, in order to identify genes exhibiting any significant expression variation across treatments. Second, pairwise contrasts between conditions (Low-P vs High-P, Mid-P vs High-P, and Low-P vs Mid-P) were performed using Wald tests to estimate log₂ fold changes (log2FC). These effect sizes represent shrinkage-adjusted estimates from the DESeq2 model, where positive values indicate higher expression in the first condition of each comparison. P-values from Wald tests were adjusted for multiple testing using the Benjamini–Hochberg false discovery rate (FDR) procedure, and genes with adjusted p-values (padj) < 0.05 were considered statistically significant. For visualization and exploratory analyses, variance-stabilizing transformation (VST) was applied to normalized counts using DESeq2. Mean expression values per condition were computed from VST-transformed data across biological replicates. These values were used exclusively for visualization and interpretation of expression patterns and were not used for statistical inference.
Notes on experimental design:
Raw sequencing data have been deposited in the Gene Expression Omnibus (GEO) at National Center for Biotechnology Information (NCBI) under accession number GSE327000.
Filella, A. (2026). RNA-seq gene expression data for Synechococcus sp. WH8102 under continuous high, mid, and low phosphorus treatments in laboratory EFB experiments in 2025. Biological and Chemical Oceanography Data Management Office (BCO-DMO). (Version 1) Version Date 2026-05-08 [if applicable, indicate subset used]. http://lod.bco-dmo.org/id/dataset/998248 [access date]
Terms of Use
This dataset is licensed under Creative Commons Attribution 4.0.
If you wish to use this dataset, it is highly recommended that you contact the original principal investigators (PI). Should the relevant PI be unavailable, please contact BCO-DMO (info@bco-dmo.org) for additional guidance. For general guidance please see the BCO-DMO Terms of Use document.