Automated Author ProfileE Virginia Armbrust
University of Washington0000-0001-7865-5101
E Virginia Armbrust
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets for this author
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the author's datasets
Total Mentions
Total mentions of the author's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 48.1 (sum of 75 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
DescriptionThe datasets supporting the conclusions of this article, including field measurements of Prochlorococcus division rates, are available in this repository. The R code performs the following tasks:Loads data from various sources, including lab experiments, dilution experiments, and in-situ measurements.Calculates thermal norm predictions using different models (Eppley, Hinshelwood, Eppley-Norberg) to predict division rates based on temperature.Generates figures to visualize the results, including latitudinal and temperature effects on division rates, model predictions compared with observed data, and changes in primary production under different emission scenarios.Fits the Hinshelwood model to culture data and extracts best-fit parameters.Performs bootstrapping to estimate uncertainty in the Hinshelwood model parameters.Calculates confidence intervals for the bootstrapped parameters.R ScriptsRibalet_main.R: This script contains the main analysis code, including data loading, model fitting, figure generation, and bootstrapping.Ribalet_fitting.R: This script defines functions for fitting different growth models to the data and estimating model parameters.RequirementsR version 4.4.2 (2024-10-31)Platform: aarch64-apple-darwin20Running under: macOS Sequoia 15.1.1Matrix products: defaultBLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib LAPACK: /Library/Frameworks/R.framework/Versions/4.4-arm64/Resources/lib/libRlapack.dylib; LAPACK version 3.12.0attached base packages:[1] parallel stats graphics grDevices utils datasets [7] methods base other attached packages: [1] DEoptim_2.2-8 arrow_15.0.1 ggpubr_0.6.0 lubridate_1.9.3 [5] forcats_1.0.0 stringr_1.5.1 dplyr_1.1.4 purrr_1.0.2 [9] readr_2.1.5 tidyr_1.3.1 tibble_3.2.1 ggplot2_3.5.1 [13] tidyverse_2.0.0InstallationInstall the required R packages:Code snippetinstall.packages(c("tidyverse", "ggpubr", "arrow", "DEoptim"))UsageThe scripts will generate figures and output files in the same directory.Input DataThe code requires the following input data files:culture.csvdilution.csvabundance.csvmpm.csvmodel_results.parquetbootstrap_projections.csvmodeled-thermal-traits.tsvsst.parquetculture_syn.csvPlease ensure that these files are present in the same directory as the R script files.Output DataThe code generates the following output files:Figures: Figure1.png, Figure2.png, Figure3.png, FigureS1.png, FigureS2.png, FigureS3.png, FigureS4.png, FigureS5.png, FigureS6.png, FigureS9.png, FigureS11.png, FigureS12.png, FigureS13.png, FigureS14.png, FigureS15.pngCSV files: bootstrap_parameters.csv, cultures_thermal_reactions.csvLicenseThis code is licensed under the MIT License.
Authors
- Ribalet, François ;
- Dutkiewicz, Stephanie ;
- Monier, Erwan ;
- Armbrust, Virginia
DescriptionThe datasets supporting the conclusions of this article, including field measurements of Prochlorococcus division rates, are available in this repository. The R code performs the following tasks:Loads data from various sources, including lab experiments, dilution experiments, and in-situ measurements.Calculates thermal norm predictions using different models (Eppley, Hinshelwood, Eppley-Norberg) to predict division rates based on temperature.Generates figures to visualize the results, including latitudinal and temperature effects on division rates, model predictions compared with observed data, and changes in primary production under different emission scenarios.Fits the Hinshelwood model to culture data and extracts best-fit parameters.Performs bootstrapping to estimate uncertainty in the Hinshelwood model parameters.Calculates confidence intervals for the bootstrapped parameters.R ScriptsRibalet_main.R: This script contains the main analysis code, including data loading, model fitting, figure generation, and bootstrapping.Ribalet_fitting.R: This script defines functions for fitting different growth models to the data and estimating model parameters.RequirementsR version 4.4.2 (2024-10-31)Platform: aarch64-apple-darwin20Running under: macOS Sequoia 15.1.1Matrix products: defaultBLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib LAPACK: /Library/Frameworks/R.framework/Versions/4.4-arm64/Resources/lib/libRlapack.dylib; LAPACK version 3.12.0attached base packages:[1] parallel stats graphics grDevices utils datasets [7] methods base other attached packages: [1] DEoptim_2.2-8 arrow_15.0.1 ggpubr_0.6.0 lubridate_1.9.3 [5] forcats_1.0.0 stringr_1.5.1 dplyr_1.1.4 purrr_1.0.2 [9] readr_2.1.5 tidyr_1.3.1 tibble_3.2.1 ggplot2_3.5.1 [13] tidyverse_2.0.0InstallationInstall the required R packages:Code snippetinstall.packages(c("tidyverse", "ggpubr", "arrow", "DEoptim"))UsageThe scripts will generate figures and output files in the same directory.Input DataThe code requires the following input data files:culture.csvdilution.csvabundance.csvmpm.csvmodel_results.parquetbootstrap_projections.csvmodeled-thermal-traits.tsvsst.parquetculture_syn.csvPlease ensure that these files are present in the same directory as the R script files.Output DataThe code generates the following output files:Figures: Figure1.png, Figure2.png, Figure3.png, FigureS1.png, FigureS2.png, FigureS3.png, FigureS4.png, FigureS5.png, FigureS6.png, FigureS9.png, FigureS11.png, FigureS12.png, FigureS13.png, FigureS14.png, FigureS15.pngCSV files: bootstrap_parameters.csv, cultures_thermal_reactions.csvLicenseThis code is licensed under the MIT License.
Authors
- Ribalet, François ;
- Dutkiewicz, Stephanie ;
- Monier, Erwan ;
- Armbrust, Virginia
DescriptionThe datasets supporting the conclusions of this article, including field measurements of Prochlorococcus division rates, are available in this repository. The R code performs the following tasks:Loads data from various sources, including lab experiments, dilution experiments, and in-situ measurements.Calculates thermal norm predictions using different models (Eppley, Hinshelwood, Eppley-Norberg) to predict division rates based on temperature.Generates figures to visualize the results, including latitudinal and temperature effects on division rates, model predictions compared with observed data, and changes in primary production under different emission scenarios.Fits the Hinshelwood model to culture data and extracts best-fit parameters.Performs bootstrapping to estimate uncertainty in the Hinshelwood model parameters.Calculates confidence intervals for the bootstrapped parameters.R ScriptsRibalet_main.R: This script contains the main analysis code, including data loading, model fitting, figure generation, and bootstrapping.Ribalet_fitting.R: This script defines functions for fitting different growth models to the data and estimating model parameters.RequirementsR version 4.4.2 (2024-10-31)Platform: aarch64-apple-darwin20Running under: macOS Sequoia 15.1.1Matrix products: defaultBLAS: /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib LAPACK: /Library/Frameworks/R.framework/Versions/4.4-arm64/Resources/lib/libRlapack.dylib; LAPACK version 3.12.0attached base packages:[1] parallel stats graphics grDevices utils datasets [7] methods base other attached packages: [1] DEoptim_2.2-8 arrow_15.0.1 ggpubr_0.6.0 lubridate_1.9.3 [5] forcats_1.0.0 stringr_1.5.1 dplyr_1.1.4 purrr_1.0.2 [9] readr_2.1.5 tidyr_1.3.1 tibble_3.2.1 ggplot2_3.5.1 [13] tidyverse_2.0.0InstallationInstall the required R packages:Code snippetinstall.packages(c("tidyverse", "ggpubr", "arrow", "DEoptim"))UsageThe scripts will generate figures and output files in the same directory.Input DataThe code requires the following input data files:culture.csvdilution.csvabundance.parquetmpm.csvmodel_results.parquetbootstrap_projections.csvmodeled-thermal-traits.tsvsst.parquetculture_syn.csvPlease ensure that these files are present in the same directory as the R script files.Output DataThe code generates the following output files:Figures: Figure1.png, Figure2.png, Figure3.png, FigureS1.png, FigureS2.png, FigureS3.png, FigureS4.png, FigureS5.png, FigureS6.png, FigureS9.png, FigureS11.png, FigureS12.png, FigureS13.png, FigureS14.png, FigureS15.pngCSV files: bootstrap_parameters.csv, cultures_thermal_reactions.csvLicenseThis code is licensed under the MIT License.
Authors
- Ribalet, François ;
- Dutkiewicz, Stephanie ;
- Monier, Erwan ;
- Armbrust, Virginia
Data used to develop and test MarPRISM, a model to predict the in situ trophic mode of marine protists. To examine the in situ activity of protists, Lambert et al., 2022 developed a machine learning model to predict the trophic mode of marine protist species based on gene expression from metatranscriptomes.Recent studies (Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021) identified that a number of the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) transcriptomes used for training the Lambert model have a low number of sequences and/or high contamination.trainingData_withContam_withLowSeqs.csv.gz: training data used by Lambert et al. (2022), includes contaminated and low-sequence entries. trainingDataMarPRISM.csv.gz: training data with contaminated and low-sequence entries removed, training data used for MarPRISM.Transcriptomes were removed from the training data and used for testing that had less than 1200 total sequences, less than 500 total assigned Pfam domains, and/or greater than 50% contamination from non-target organisms.trainingDataMarPRISM_binary.csv.gz: trainingDataMarPRISM.csv.gz but TPM values greater than 0 were converted to 1 in order to determine whether a model could be built based on the binary expression of Pfams rathern than continuous expression values. trainingData_micromonasMixToPhot.csv.gz: trainingDataMarPRISM.csv.gz but with any transcriptomes from Micromonas strains that were originally labeled mixotrophic switched to phototrophic. This dataset was created to determine the effect of Micromonas in the training data, as well as the effect of permuting trophic mode labels in the training data as there may be errors in some of the trophic mode labels.The training datasets are unbalanced with more phototrophic transcriptomes than heterotrophic and mixotrophic transcriptomes.So feature selection and hyperparameter search was conducted on versions of the training datasets with phototrophic transcriptomes randomly undersampled. trainingData_withContam_withLowSeqs_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_140phot.zip: 140 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_contamLowSeqsRemoved_50phot.zip: 50 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_binary_50phot.zip: same as trainingData_contamLowSeqsRemoved_50phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_80phot.zip: same as trainingData_contamLowSeqsRemoved_80phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_100phot.zip: same as trainingData_contamLowSeqsRemoved_100phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_120phot.zip: same as trainingData_contamLowSeqsRemoved_120phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_micromonasMixToPhot_50phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 50 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_80phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 80 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_100phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 100 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_120phot.zip: after converting Micromonas mixotrophy labels in trainingDataMarPRISM.csv.gz to phototrophy, 120 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomesFeature selection was run on the training datasets after undersampling phototrophic transcriptomes. Feature selection was run for both XGBoost and Random Forest models. MarPRISM_featurePfams.csv.gz: MarPRISM feature Pfams, XGBoost model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_rfModel_rfFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgRFFeatures.csv.gz: Union of XGBoost and Random Forest feature Pfams, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: XGBoost model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_binary.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, TPM values > 0 were converted to 1Extracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_micromonasMixToPhot.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, mixotrophy labels for any Micromonas strains were converted to phototrophyTranscriptomes not included in the training data and not from the MMETSP were used to test MarPRISM. testTranscriptomes.csv.gz: Transcript per million counts by Pfam ID for these test transcriptomesSome of these transcriptomes were processed and used for testing by Lambert et al., 2022, while other transcriptomes, from Pterosperma cristatum, Amphora coffeaeformis, Chaetoceros sp., and Cylindrotheca closterium, were newly added for testing. Transcripts per million for the latter transcriptomes derived from publicly salmon mappings.Pfam annotations for the newly added transcriptomes were generated through the following: the publicly available assembled transcriptomes for each species was six-frame translated with transeq , the longest reading frame (minimum 100 amino acid length) was selected for each contig, the longest contig was compared to the Pfam database version 34 with hmmsearch, and the Pfam annotation with the best bitscore for each contig was retained (e-value < 1e-05).Transcripts per million were summed by Pfam. testTranscriptomes.xlsx.gz: Accession IDs and references for these test transcriptomes; transcriptome culture conditions. Transcriptomes that were excluded from the training data for MarPRISM due to having high contamination or low-sequence abundance were also used to test MarPRISM. testTranscriptomes_MMETSP.csv.gz: Transcript per million counts by Pfam ID for these excluded transcriptomestestTranscriptomes_MMETSP.xlsx.gz: MMETSP IDs and references for these excluded transcriptomes. Contamination and low-sequence entries were identified by Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021 and curated by Groussman et al., 2023. Identity of ribosomal sequences was analyzed by Groussman et al., 2023.qc_flag_Groussman: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.num_sequences_Groussman: Number of sequences in original sequence file.num_pfams_Groussman: Number of Pfam domains identified in protein sequences.flag_Lasek: Flag notes from Lasek-Nesselquist and Johnson, 2019; CONTAM NOTED; ciliate samples reported as contaminated in this study.flag_VanVlierberghe: Flag for a high level of estimated contamination from 'flag_VanVlierberghe';CONTAM_50PCT; contamination percentages over 50%:flag_ribosomalContamination_Groussman: Flag for a high level of estimated contamination, from ‘ribosomal_contam_pct_Groussman'; CONTAM_50PCT; contamination percentages over 50%. ribosomal_contam_pct_Groussman: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity. Ribosomal taxonomy of most abundant contaminant: For entries with greater than 50% ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, the taxonomic identity of the most abundant ribosomal protein sequences not identified as the recorded identity of the transcriptome in the MMETSP.Expected taxonomy of transcriptome: recorded identity in MMETSP.
Authors
- Thomas, Elaina ;
- Groussman, Mora Jove ;
- Coesel, Sacha ;
- Armbrust, Virginia
G3_depth.assembledReads.id99.fasta.gz G3 depth profiles amino acid assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the cluster representativesThe code for processing and assemblying reads can be found hereAmino acid formatNPac.G3PA_depth.bf100.id99.nt.fasta.gzG3 depth profiles nucleotide assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the nucleotide-encoded contigs that the amino acid-encoded cluster representatives are from The code for processing and assemblying reads can be found hereNucleotide formatNPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was taxonomically annotated using diamond last common ancestor and MarFERReTBash script used: diamond blastp --no-unlink -t ~/ -b 100 -c 1 -p 32 -d /mnt/nfs/projects/marferret/v1/data/marmicrodb/dmnd/MarFERReT.v1.1.MMDB.combined.dmnd -e 1e-5 --top 10 -f 102 -q G3_depth_assembledReads.id99.fasta -o NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tabG3PA_depth.Pfam34.domtblout.tab.gzPfam annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was functionally annotated against the Pfam database (version 34) using hmmsearchBash script used: hmmsearch --cut_tc --domtblout G3PA_depth.Pfam34.domtblout.tab Pfam_34.0/Pfam-A.hmm G3_depth_assembledReads.id99.fastaG3PA_depth.raw.est_counts.csv.gzEstimated counts of transcripts from G3 depth profile samples mapped to G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)The code for processing, assemblying, and mapping reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded but the G3 depth profiles were assembled unstranded so the reads were mapped against the assembly unstranded The G3 depth profile reads were mapped against the G3 depth profile assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant -i NPac.G3PA_depth.id99.nt.idx -o ${SAMPLE} --threads=${N_THREADS} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> G3PA_depth.${SAMPLE}.kallisto.logG2PA_incubations.est_counts.csv.gzEstimated counts of transcripts from G2 incubation samples mapped to G2 surface assemblyTrimmed, quality-controlled reads from the G2 incubations were mapped to the G2 assembly (Gradients2.MGL1704.PA.assemblies.tar.gz)The code for processing reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded; the G2 surface samples were assembled reverse stranded; the G2 incubations were mapped reverse stranded against the G2 surface samplesThe G2 incubation reads were mapped against the G2 surface assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant --rf-stranded -i NPac.G2PA.bf100.id99.nt.idx -o ${SAMPLE} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> ${SAMPLE}.kallisto.logAt 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedchlorophyllA_g2Incubations.csvChlorophyll a measurements from the G2 incubations Samples for metatranscriptomes were collected after zero and 96 hours from the nutrient amendment experiments, along with onboard fluorometer measurements of chlorophyll a.Depth refers to the depth at which water was collected for the incubations. At 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedThe following datasets contain the transcripts per million per Pfam per taxonomic bin per sample:G1PA.tpm_counts.csv.gz G1 surfaceG2PA.tpm_counts.csv.gz G2 surfaceG3PA.tpm_counts.csv.gz G3 surfaceD1PA.tpm_counts.csv.gz Aloha dielG3PA_diel.tpm_counts.csv.gz G3 dielG2PA_incubations.tpm_counts.csv.gz G2 nutrient amendment incubationsG3PA_depth.tpm_counts.csv.gz G3 depth profilesTranscripts per million were calculated with the following stepsDivide the estimated number of reads mapped to each contig by its nucleotide length (in kilobases) to generate reads per kilobase (RPK).Sum the RPK by species and sample, then divide by one million to generate a conversion factor.Divide the RPK by the conversion factor to calculate transcripts per million per contig.Sum transcripts per million by Pfam for each species and sample.Transcript per million counts are only provided for annotated Pfams Transcripts per million were calculated using estimated counts outputted from kallistoEstimated count files for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereEstimated count files for the G2 nutrient amendment incubation and G3 depth profile samples can be found in this repository: G2PA_incubations.est_counts.csv.gz and G3PA_depth.raw.est_counts.csv.gz respectivelyTranscripts per million were calculated per taxonomic binTaxonomic annotations were generated using diamond last common ancestor and MarFERReTTaxonomic annotations for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the taxonomic annotations for the G2 surface assembly can also found in the previous repository: NPac.G2PA.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations the G3 depth profile samples can be found in this repository: NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTranscripts per million were calculated per PfamPfam annotations were generated using hmmsearchPfam annotations for the G1-G3 surface, Aloha diel, and G3 diel samples were generated using version 35 of the Pfam database and can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the Pfam annotations for the G2 surface assembly were generated using version 35 of the Pfam database and can also found in the previous repository: G2PA.Pfam35.domtblout.tab.gzPfam annotations the G3 depth profile samples were generated using version 34 of the Pfam database can be found in this repository: G3PA_depth.Pfam34.domtblout.tab.gz
Authors
- Thomas, Elaina ;
- Groussman, Mora Jove ;
- Coesel, Sacha ;
- Armbrust, Virginia
Data used to develop and test MarPRISM, a model to predict the in situ trophic mode of marine protists. To examine the in situ activity of protists, Lambert et al., 2022 developed a machine learning model to predict the trophic mode of marine protist species based on gene expression from metatranscriptomes.Recent studies (Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021) identified that a number of the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) transcriptomes used for training the Lambert model have a low number of sequences and/or high contamination.trainingData_withContam_withLowSeqs.csv.gz: training data used by Lambert et al. (2022), includes contaminated and low-sequence entries. trainingDataMarPRISM.csv.gz: training data with contaminated and low-sequence entries removed, training data used for MarPRISM.Transcriptomes were removed from the training data and used for testing that had less than 1200 total sequences, less than 500 total assigned Pfam domains, and/or greater than 50% contamination from non-target organisms.trainingDataMarPRISM_binary.csv.gz: trainingDataMarPRISM.csv.gz but TPM values greater than 0 were converted to 1 in order to determine whether a model could be built based on the binary expression of Pfams rathern than continuous expression values. trainingData_micromonasMixToPhot.csv.gz: trainingDataMarPRISM.csv.gz but with any transcriptomes from Micromonas strains that were originally labeled mixotrophic switched to phototrophic. This dataset was created to determine the effect of Micromonas in the training data, as well as the effect of permuting trophic mode labels in the training data as there may be errors in some of the trophic mode labels.The training datasets are unbalanced with more phototrophic transcriptomes than heterotrophic and mixotrophic transcriptomes.So feature selection and hyperparameter search was conducted on versions of the training datasets with phototrophic transcriptomes randomly undersampled. trainingData_withContam_withLowSeqs_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_140phot.zip: 140 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_contamLowSeqsRemoved_50phot.zip: 50 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_binary_50phot.zip: same as trainingData_contamLowSeqsRemoved_50phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_80phot.zip: same as trainingData_contamLowSeqsRemoved_80phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_100phot.zip: same as trainingData_contamLowSeqsRemoved_100phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_120phot.zip: same as trainingData_contamLowSeqsRemoved_120phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_micromonasMixToPhot_50phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 50 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_80phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 80 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_100phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 100 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_120phot.zip: after converting Micromonas mixotrophy labels in trainingDataMarPRISM.csv.gz to phototrophy, 120 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomesFeature selection was run on the training datasets after undersampling phototrophic transcriptomes. Feature selection was run for both XGBoost and Random Forest models. MarPRISM_featurePfams.csv.gz: MarPRISM feature Pfams, XGBoost model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_rfModel_rfFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgRFFeatures.csv.gz: Union of XGBoost and Random Forest feature Pfams, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: XGBoost model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_binary.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, TPM values > 0 were converted to 1Extracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_micromonasMixToPhot.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, mixotrophy labels for any Micromonas strains were converted to phototrophyTranscriptomes not included in the training data and not from the MMETSP were used to test MarPRISM. testTranscriptomes.csv.gz: Transcript per million counts by Pfam ID for these test transcriptomesSome of these transcriptomes were processed and used for testing by Lambert et al., 2022, while other transcriptomes, from Pterosperma cristatum, Amphora coffeaeformis, Chaetoceros sp., and Cylindrotheca closterium, were newly added for testing. Transcripts per million for the latter transcriptomes derived from publicly salmon mappings.Pfam annotations for the newly added transcriptomes were generated through the following: the publicly available assembled transcriptomes for each species was six-frame translated with transeq , the longest reading frame (minimum 100 amino acid length) was selected for each contig, the longest contig was compared to the Pfam database version 34 with hmmsearch, and the Pfam annotation with the best bitscore for each contig was retained (e-value < 1e-05).Transcripts per million were summed by Pfam. testTranscriptomes.xlsx.gz: Accession IDs and references for these test transcriptomes; transcriptome culture conditions. Transcriptomes that were excluded from the training data for MarPRISM due to having high contamination or low-sequence abundance were also used to test MarPRISM. testTranscriptomes_MMETSP.csv.gz: Transcript per million counts by Pfam ID for these excluded transcriptomestestTranscriptomes_MMETSP.xlsx.gz: MMETSP IDs and references for these excluded transcriptomes. Contamination and low-sequence entries were identified by Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021 and curated by Groussman et al., 2023. Identity of ribosomal sequences was analyzed by Groussman et al., 2023.qc_flag_Groussman: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.num_sequences_Groussman: Number of sequences in original sequence file.num_pfams_Groussman: Number of Pfam domains identified in protein sequences.flag_Lasek: Flag notes from Lasek-Nesselquist and Johnson, 2019; CONTAM NOTED; ciliate samples reported as contaminated in this study.flag_VanVlierberghe: Flag for a high level of estimated contamination from 'flag_VanVlierberghe';CONTAM_50PCT; contamination percentages over 50%:flag_ribosomalContamination_Groussman: Flag for a high level of estimated contamination, from ‘ribosomal_contam_pct_Groussman'; CONTAM_50PCT; contamination percentages over 50%. ribosomal_contam_pct_Groussman: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity. Ribosomal taxonomy of most abundant contaminant: For entries with greater than 50% ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, the taxonomic identity of the most abundant ribosomal protein sequences not identified as the recorded identity of the transcriptome in the MMETSP.Expected taxonomy of transcriptome: recorded identity in MMETSP.
Authors
- Thomas, Elaina ;
- Groussman, Mora Jove ;
- Coesel, Sacha ;
- Armbrust, Virginia
G3_depth.assembledReads.id99.fasta.gz G3 depth profiles amino acid assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the cluster representativesThe code for processing and assemblying reads can be found hereAmino acid formatNPac.G3PA_depth.bf100.id99.nt.fasta.gzG3 depth profiles nucleotide assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the nucleotide-encoded contigs that the amino acid-encoded cluster representatives are from The code for processing and assemblying reads can be found hereNucleotide formatNPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was taxonomically annotated using diamond last common ancestor and MarFERReTBash script used: diamond blastp --no-unlink -t ~/ -b 100 -c 1 -p 32 -d /mnt/nfs/projects/marferret/v1/data/marmicrodb/dmnd/MarFERReT.v1.1.MMDB.combined.dmnd -e 1e-5 --top 10 -f 102 -q G3_depth_assembledReads.id99.fasta -o NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tabG3PA_depth.Pfam34.domtblout.tab.gzPfam annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was functionally annotated against the Pfam database (version 34) using hmmsearchBash script used: hmmsearch --cut_tc --domtblout G3PA_depth.Pfam34.domtblout.tab Pfam_34.0/Pfam-A.hmm G3_depth_assembledReads.id99.fastaG3PA_depth.raw.est_counts.csv.gzEstimated counts of transcripts from G3 depth profile samples mapped to G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)The code for processing, assemblying, and mapping reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded but the G3 depth profiles were assembled unstranded so the reads were mapped against the assembly unstranded The G3 depth profile reads were mapped against the G3 depth profile assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant -i NPac.G3PA_depth.id99.nt.idx -o ${SAMPLE} --threads=${N_THREADS} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> G3PA_depth.${SAMPLE}.kallisto.logG2PA_incubations.est_counts.csv.gzEstimated counts of transcripts from G2 incubation samples mapped to G2 surface assemblyTrimmed, quality-controlled reads from the G2 incubations were mapped to the G2 assembly (Gradients2.MGL1704.PA.assemblies.tar.gz)The code for processing reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded; the G2 surface samples were assembled reverse stranded; the G2 incubations were mapped reverse stranded against the G2 surface samplesThe G2 incubation reads were mapped against the G2 surface assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant --rf-stranded -i NPac.G2PA.bf100.id99.nt.idx -o ${SAMPLE} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> ${SAMPLE}.kallisto.logAt 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedchlorophyllA_g2Incubations.csvChlorophyll a measurements from the G2 incubations Samples for metatranscriptomes were collected after zero and 96 hours from the nutrient amendment experiments, along with onboard fluorometer measurements of chlorophyll a.Depth refers to the depth at which water was collected for the incubations. At 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedThe following datasets contain the transcripts per million per Pfam per taxonomic bin per sample:G1PA.tpm_counts.csv.gz G1 surfaceG2PA.tpm_counts.csv.gz G2 surfaceG3PA.tpm_counts.csv.gz G3 surfaceD1PA.tpm_counts.csv.gz Aloha dielG3PA_diel.tpm_counts.csv.gz G3 dielG2PA_incubations.tpm_counts.csv.gz G2 nutrient amendment incubationsG3PA_depth.tpm_counts.csv.gz G3 depth profilesTranscripts per million were calculated with the following stepsDivide the estimated number of reads mapped to each contig by its nucleotide length (in kilobases) to generate reads per kilobase (RPK).Sum the RPK by species and sample, then divide by one million to generate a conversion factor.Divide the RPK by the conversion factor to calculate transcripts per million per contig.Sum transcripts per million by Pfam for each species and sample.Transcript per million counts are only provided for annotated Pfams Transcripts per million were calculated using estimated counts outputted from kallistoEstimated count files for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereEstimated count files for the G2 nutrient amendment incubation and G3 depth profile samples can be found in this repository: G2PA_incubations.est_counts.csv.gz and G3PA_depth.raw.est_counts.csv.gz respectivelyTranscripts per million were calculated per taxonomic binTaxonomic annotations were generated using diamond last common ancestor and MarFERReTTaxonomic annotations for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the taxonomic annotations for the G2 surface assembly can also found in the previous repository: NPac.G2PA.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations the G3 depth profile samples can be found in this repository: NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTranscripts per million were calculated per PfamPfam annotations were generated using hmmsearchPfam annotations for the G1-G3 surface, Aloha diel, and G3 diel samples were generated using version 35 of the Pfam database and can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the Pfam annotations for the G2 surface assembly were generated using version 35 of the Pfam database and can also found in the previous repository: G2PA.Pfam35.domtblout.tab.gzPfam annotations the G3 depth profile samples were generated using version 34 of the Pfam database can be found in this repository: G3PA_depth.Pfam34.domtblout.tab.gz
Authors
- Thomas, Elaina ;
- Groussman, Mora Jove ;
- Coesel, Sacha ;
- Armbrust, Virginia
This data continues with the development of the NPEGC Trinity de novo metatranscriptome assemblies from the protein data repository of The North Pacific Eukaryotic Gene Catalog. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:NPac.G1PA.bf100.id99.nt.fasta.gzNPac.G2PA.bf100.id99.nt.fasta.gzNPac.G3PA.bf100.id99.nt.fasta.gzNPac.G3PA_diel.bf100.id99.nt.fasta.gzNPac.D1PA.bf100.id99.nt.fasta.gzA full description of this data is published in Scientific Data, available here: The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Please cite this publication if your research uses this data:Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Scientific Data, 11(1), 1161.These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3Key processing steps are sampled below with links to the detailed code on the main github code repository: https://github.com/armbrustlab/NPac_euk_gene_catalogCode used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: NPEGC.nt_kallisto_counts.shThere are two main steps:1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts2. Map the short reads from environmental samples back to the assembly indexAs generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.The code in this template script was used for each project: aggregate_kallisto_counts.RThe output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: G1PA.raw.est_counts.csv.gzG2PA.raw.est_counts.csv.gzG3PA.raw.est_counts.csv.gzG3PA_diel.raw.est_counts.csv.gzD1PA.raw.est_counts.csv.gz
Authors
- Groussman, Ryan ;
- Coesel, Sacha ;
- Armbrust, E. Virginia
This data continues with the development of the NPEGC Trinity de novo metatranscriptome assemblies from the protein data repository of The North Pacific Eukaryotic Gene Catalog. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:NPac.G1PA.bf100.id99.nt.fasta.gzNPac.G2PA.bf100.id99.nt.fasta.gzNPac.G3PA.bf100.id99.nt.fasta.gzNPac.G3PA_diel.bf100.id99.nt.fasta.gzNPac.D1PA.bf100.id99.nt.fasta.gzA full description of this data is published in Scientific Data, available here: The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Please cite this publication if your research uses this data:Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Scientific Data, 11(1), 1161.These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3Key processing steps are sampled below with links to the detailed code on the main github code repository: https://github.com/armbrustlab/NPac_euk_gene_catalogCode used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: NPEGC.nt_kallisto_counts.shThere are two main steps:1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts2. Map the short reads from environmental samples back to the assembly indexAs generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.The code in this template script was used for each project: aggregate_kallisto_counts.RThe output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: G1PA.raw.est_counts.csv.gzG2PA.raw.est_counts.csv.gzG3PA.raw.est_counts.csv.gzG3PA_diel.raw.est_counts.csv.gzD1PA.raw.est_counts.csv.gz
Authors
- Groussman, Ryan ;
- Coesel, Sacha ;
- Armbrust, E. Virginia
The dataset consists of BD Influx-based analysis of phytoplankton populations from discrete flow cytometry data collected underway during the University of Washington School of Oceanography 2023 undergraduate senior thesis cruise (TN413) oceanographic research cruise from Hawaii to Fiji. The data consists of cell abundance, cell size (equivalent spherical diameter), carbon quota, and carbon biomass for heterotrophic bacteria, picophytoplankton populations, namely the cyanobacteria Prochlorococcus and Synechococcus, and small eukaryotic phytoplankton (<5 μm ESD). Time is in UTC format, latitude and longitude are in decimal degrees, and depth is in meters. Further information can be found here: https://github.com/fribalet/FCSplankton
Authors
- Cain, Kelsy ;
- Ribalet, François ;
- Armbrust, Virginia