Automated Author Profile

E Virginia Armbrust

University of Washington
0000-0001-7865-5101

Current S-Index

48.1

Sum of Dataset Indices for all datasets

Average Dataset Index per Dataset

0.6

Average Dataset Index per dataset

Total Datasets

75

Total datasets for this author

Average FAIR Score

76.9%

Average FAIR Score per dataset

Total Citations

25

Total citations to the author's datasets

Total Mentions

1

Total mentions of the author's datasets

S-Index Interpretation

S-Index Over Time

Cumulative Citations Over Time

Cumulative Mentions Over Time

Datasets

Future Ocean Warming May Threaten Key Photosynthetic Microbes (Version: v1.3)

DescriptionThe datasets supporting the conclusions of this article, including field measurements of Prochlorococcus division rates, are available in this repository. The R code performs the following tasks:Loads data from various sources, including lab experiments, dilution experiments, and in-situ measurements.Calculates thermal norm predictions using different models (Eppley, Hinshelwood, Eppley-Norberg) to predict division rates based on temperature.Generates figures to visualize the results, including latitudinal and temperature effects on division rates, model predictions compared with observed data, and changes in primary production under different emission scenarios.Fits the Hinshelwood model to culture data and extracts best-fit parameters.Performs bootstrapping to estimate uncertainty in the Hinshelwood model parameters.Calculates confidence intervals for the bootstrapped parameters.R ScriptsRibalet_main.R: This script contains the main analysis code, including data loading, model fitting, figure generation, and bootstrapping.Ribalet_fitting.R: This script defines functions for fitting different growth models to the data and estimating model parameters.RequirementsR version 4.4.2 (2024-10-31)Platform: aarch64-apple-darwin20Running under: macOS Sequoia 15.1.1Matrix products: defaultBLAS:   /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib LAPACK: /Library/Frameworks/R.framework/Versions/4.4-arm64/Resources/lib/libRlapack.dylib;  LAPACK version 3.12.0attached base packages:[1] parallel  stats     graphics  grDevices utils     datasets [7] methods   base     other attached packages: [1] DEoptim_2.2-8   arrow_15.0.1    ggpubr_0.6.0    lubridate_1.9.3 [5] forcats_1.0.0   stringr_1.5.1   dplyr_1.1.4     purrr_1.0.2     [9] readr_2.1.5     tidyr_1.3.1     tibble_3.2.1    ggplot2_3.5.1  [13] tidyverse_2.0.0InstallationInstall the required R packages:Code snippetinstall.packages(c("tidyverse", "ggpubr", "arrow", "DEoptim"))UsageThe scripts will generate figures and output files in the same directory.Input DataThe code requires the following input data files:culture.csvdilution.csvabundance.csvmpm.csvmodel_results.parquetbootstrap_projections.csvmodeled-thermal-traits.tsvsst.parquetculture_syn.csvPlease ensure that these files are present in the same directory as the R script files.Output DataThe code generates the following output files:Figures: Figure1.png, Figure2.png, Figure3.png, FigureS1.png, FigureS2.png, FigureS3.png, FigureS4.png, FigureS5.png, FigureS6.png, FigureS9.png, FigureS11.png, FigureS12.png, FigureS13.png, FigureS14.png, FigureS15.pngCSV files: bootstrap_parameters.csv, cultures_thermal_reactions.csvLicenseThis code is licensed under the MIT License.

Authors

  • Ribalet, François ;
  • Dutkiewicz, Stephanie ;
  • Monier, Erwan ;
  • Armbrust, Virginia
1 Citation0 Mentions73% FAIR0.9 Dataset Index
10.5281/zenodo.110433862025

Future Ocean Warming May Threaten Key Photosynthetic Microbes (Version: v1.3)

DescriptionThe datasets supporting the conclusions of this article, including field measurements of Prochlorococcus division rates, are available in this repository. The R code performs the following tasks:Loads data from various sources, including lab experiments, dilution experiments, and in-situ measurements.Calculates thermal norm predictions using different models (Eppley, Hinshelwood, Eppley-Norberg) to predict division rates based on temperature.Generates figures to visualize the results, including latitudinal and temperature effects on division rates, model predictions compared with observed data, and changes in primary production under different emission scenarios.Fits the Hinshelwood model to culture data and extracts best-fit parameters.Performs bootstrapping to estimate uncertainty in the Hinshelwood model parameters.Calculates confidence intervals for the bootstrapped parameters.R ScriptsRibalet_main.R: This script contains the main analysis code, including data loading, model fitting, figure generation, and bootstrapping.Ribalet_fitting.R: This script defines functions for fitting different growth models to the data and estimating model parameters.RequirementsR version 4.4.2 (2024-10-31)Platform: aarch64-apple-darwin20Running under: macOS Sequoia 15.1.1Matrix products: defaultBLAS:   /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib LAPACK: /Library/Frameworks/R.framework/Versions/4.4-arm64/Resources/lib/libRlapack.dylib;  LAPACK version 3.12.0attached base packages:[1] parallel  stats     graphics  grDevices utils     datasets [7] methods   base     other attached packages: [1] DEoptim_2.2-8   arrow_15.0.1    ggpubr_0.6.0    lubridate_1.9.3 [5] forcats_1.0.0   stringr_1.5.1   dplyr_1.1.4     purrr_1.0.2     [9] readr_2.1.5     tidyr_1.3.1     tibble_3.2.1    ggplot2_3.5.1  [13] tidyverse_2.0.0InstallationInstall the required R packages:Code snippetinstall.packages(c("tidyverse", "ggpubr", "arrow", "DEoptim"))UsageThe scripts will generate figures and output files in the same directory.Input DataThe code requires the following input data files:culture.csvdilution.csvabundance.csvmpm.csvmodel_results.parquetbootstrap_projections.csvmodeled-thermal-traits.tsvsst.parquetculture_syn.csvPlease ensure that these files are present in the same directory as the R script files.Output DataThe code generates the following output files:Figures: Figure1.png, Figure2.png, Figure3.png, FigureS1.png, FigureS2.png, FigureS3.png, FigureS4.png, FigureS5.png, FigureS6.png, FigureS9.png, FigureS11.png, FigureS12.png, FigureS13.png, FigureS14.png, FigureS15.pngCSV files: bootstrap_parameters.csv, cultures_thermal_reactions.csvLicenseThis code is licensed under the MIT License.

Authors

  • Ribalet, François ;
  • Dutkiewicz, Stephanie ;
  • Monier, Erwan ;
  • Armbrust, Virginia
0 Citations0 Mentions73% FAIR0.6 Dataset Index
10.5281/zenodo.154856212025

Future Ocean Warming May Threaten Key Photosynthetic Microbes (Version: v1.2)

DescriptionThe datasets supporting the conclusions of this article, including field measurements of Prochlorococcus division rates, are available in this repository. The R code performs the following tasks:Loads data from various sources, including lab experiments, dilution experiments, and in-situ measurements.Calculates thermal norm predictions using different models (Eppley, Hinshelwood, Eppley-Norberg) to predict division rates based on temperature.Generates figures to visualize the results, including latitudinal and temperature effects on division rates, model predictions compared with observed data, and changes in primary production under different emission scenarios.Fits the Hinshelwood model to culture data and extracts best-fit parameters.Performs bootstrapping to estimate uncertainty in the Hinshelwood model parameters.Calculates confidence intervals for the bootstrapped parameters.R ScriptsRibalet_main.R: This script contains the main analysis code, including data loading, model fitting, figure generation, and bootstrapping.Ribalet_fitting.R: This script defines functions for fitting different growth models to the data and estimating model parameters.RequirementsR version 4.4.2 (2024-10-31)Platform: aarch64-apple-darwin20Running under: macOS Sequoia 15.1.1Matrix products: defaultBLAS:   /System/Library/Frameworks/Accelerate.framework/Versions/A/Frameworks/vecLib.framework/Versions/A/libBLAS.dylib LAPACK: /Library/Frameworks/R.framework/Versions/4.4-arm64/Resources/lib/libRlapack.dylib;  LAPACK version 3.12.0attached base packages:[1] parallel  stats     graphics  grDevices utils     datasets [7] methods   base     other attached packages: [1] DEoptim_2.2-8   arrow_15.0.1    ggpubr_0.6.0    lubridate_1.9.3 [5] forcats_1.0.0   stringr_1.5.1   dplyr_1.1.4     purrr_1.0.2     [9] readr_2.1.5     tidyr_1.3.1     tibble_3.2.1    ggplot2_3.5.1  [13] tidyverse_2.0.0InstallationInstall the required R packages:Code snippetinstall.packages(c("tidyverse", "ggpubr", "arrow", "DEoptim"))UsageThe scripts will generate figures and output files in the same directory.Input DataThe code requires the following input data files:culture.csvdilution.csvabundance.parquetmpm.csvmodel_results.parquetbootstrap_projections.csvmodeled-thermal-traits.tsvsst.parquetculture_syn.csvPlease ensure that these files are present in the same directory as the R script files.Output DataThe code generates the following output files:Figures: Figure1.png, Figure2.png, Figure3.png, FigureS1.png, FigureS2.png, FigureS3.png, FigureS4.png, FigureS5.png, FigureS6.png, FigureS9.png, FigureS11.png, FigureS12.png, FigureS13.png, FigureS14.png, FigureS15.pngCSV files: bootstrap_parameters.csv, cultures_thermal_reactions.csvLicenseThis code is licensed under the MIT License.

Authors

  • Ribalet, François ;
  • Dutkiewicz, Stephanie ;
  • Monier, Erwan ;
  • Armbrust, Virginia
0 Citations0 Mentions73% FAIR0.6 Dataset Index
10.5281/zenodo.153674022025

MarPRISM

Data used to develop and test MarPRISM, a model to predict the in situ trophic mode of marine protists. To examine the in situ activity of protists, Lambert et al., 2022 developed a machine learning model to predict the trophic mode of marine protist species based on gene expression from metatranscriptomes.Recent studies (Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021) identified that a number of the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) transcriptomes used for training the Lambert model have a low number of sequences and/or high contamination.trainingData_withContam_withLowSeqs.csv.gz: training data used by Lambert et al. (2022), includes contaminated and low-sequence entries. trainingDataMarPRISM.csv.gz: training data with contaminated and low-sequence entries removed, training data used for MarPRISM.Transcriptomes were removed from the training data and used for testing that had less than 1200 total sequences, less than 500 total assigned Pfam domains, and/or greater than 50% contamination from non-target organisms.trainingDataMarPRISM_binary.csv.gz: trainingDataMarPRISM.csv.gz but TPM values greater than 0 were converted to 1 in order to determine whether a model could be built based on the binary expression of Pfams rathern than continuous expression values. trainingData_micromonasMixToPhot.csv.gz: trainingDataMarPRISM.csv.gz but with any transcriptomes from Micromonas strains that were originally labeled mixotrophic switched to phototrophic. This dataset was created to determine the effect of Micromonas in the training data, as well as the effect of permuting trophic mode labels in the training data as there may be errors in some of the trophic mode labels.The training datasets are unbalanced with more phototrophic transcriptomes than heterotrophic and mixotrophic transcriptomes.So feature selection and hyperparameter search was conducted on versions of the training datasets with phototrophic transcriptomes randomly undersampled. trainingData_withContam_withLowSeqs_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_140phot.zip: 140 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_contamLowSeqsRemoved_50phot.zip: 50 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_binary_50phot.zip: same as trainingData_contamLowSeqsRemoved_50phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_80phot.zip: same as trainingData_contamLowSeqsRemoved_80phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_100phot.zip: same as trainingData_contamLowSeqsRemoved_100phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_120phot.zip: same as trainingData_contamLowSeqsRemoved_120phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_micromonasMixToPhot_50phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 50 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_80phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 80 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_100phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 100 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_120phot.zip: after converting Micromonas mixotrophy labels in trainingDataMarPRISM.csv.gz to phototrophy, 120 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomesFeature selection was run on the training datasets after undersampling phototrophic transcriptomes. Feature selection was run for both XGBoost and Random Forest models.  MarPRISM_featurePfams.csv.gz: MarPRISM feature Pfams, XGBoost model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_rfModel_rfFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgRFFeatures.csv.gz: Union of XGBoost and Random Forest feature Pfams, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: XGBoost model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_binary.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, TPM values > 0 were converted to 1Extracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_micromonasMixToPhot.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, mixotrophy labels for any Micromonas strains were converted to phototrophyTranscriptomes not included in the training data and not from the MMETSP were used to test MarPRISM. testTranscriptomes.csv.gz: Transcript per million counts by Pfam ID for these test transcriptomesSome of these transcriptomes were processed and used for testing by Lambert et al., 2022, while other transcriptomes, from Pterosperma cristatum, Amphora coffeaeformis, Chaetoceros sp., and Cylindrotheca closterium, were newly added for testing. Transcripts per million for the latter transcriptomes derived from publicly salmon mappings.Pfam annotations for the newly added transcriptomes were generated through the following: the publicly available assembled transcriptomes for each species was six-frame translated with transeq , the longest reading frame (minimum 100 amino acid length) was selected for each contig, the longest contig was compared to the Pfam database version 34 with hmmsearch, and the Pfam annotation with the best bitscore for each contig was retained (e-value < 1e-05).Transcripts per million were summed by Pfam. testTranscriptomes.xlsx.gz: Accession IDs and references for these test transcriptomes; transcriptome culture conditions. Transcriptomes that were excluded from the training data for MarPRISM due to having high contamination or low-sequence abundance were also used to test MarPRISM. testTranscriptomes_MMETSP.csv.gz: Transcript per million counts by Pfam ID for these excluded transcriptomestestTranscriptomes_MMETSP.xlsx.gz: MMETSP IDs and references for these excluded transcriptomes. Contamination and low-sequence entries were identified by Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021 and curated by Groussman et al., 2023. Identity of ribosomal sequences was analyzed by Groussman et al., 2023.qc_flag_Groussman: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.num_sequences_Groussman: Number of sequences in original sequence file.num_pfams_Groussman: Number of Pfam domains identified in protein sequences.flag_Lasek: Flag notes from Lasek-Nesselquist and Johnson, 2019; CONTAM NOTED; ciliate samples reported as contaminated in this study.flag_VanVlierberghe: Flag for a high level of estimated contamination from 'flag_VanVlierberghe';CONTAM_50PCT; contamination percentages over 50%:flag_ribosomalContamination_Groussman: Flag for a high level of estimated contamination, from ‘ribosomal_contam_pct_Groussman'; CONTAM_50PCT; contamination percentages over 50%. ribosomal_contam_pct_Groussman: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity. Ribosomal taxonomy of most abundant contaminant: For entries with greater than 50% ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, the taxonomic identity of the most abundant ribosomal protein sequences not identified as the recorded identity of the transcriptome in the MMETSP.Expected taxonomy of transcriptome: recorded identity in MMETSP.

Authors

  • Thomas, Elaina ;
  • Groussman, Mora Jove ;
  • Coesel, Sacha ;
  • Armbrust, Virginia
2 Citations0 Mentions81% FAIR1.3 Dataset Index
10.5281/zenodo.145189022024

Gradients 1-3 polyA-selected transcripts per million, Gradients 3 depth profile polyA-selected processed metatranscriptomes

G3_depth.assembledReads.id99.fasta.gz G3 depth profiles amino acid assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the cluster representativesThe code for processing and assemblying reads can be found hereAmino acid formatNPac.G3PA_depth.bf100.id99.nt.fasta.gzG3 depth profiles nucleotide assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the nucleotide-encoded contigs that the amino acid-encoded cluster representatives are from The code for processing and assemblying reads can be found hereNucleotide formatNPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was taxonomically annotated using diamond last common ancestor and MarFERReTBash script used: diamond blastp --no-unlink -t ~/ -b 100 -c 1 -p 32 -d /mnt/nfs/projects/marferret/v1/data/marmicrodb/dmnd/MarFERReT.v1.1.MMDB.combined.dmnd -e 1e-5 --top 10 -f 102 -q G3_depth_assembledReads.id99.fasta -o NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tabG3PA_depth.Pfam34.domtblout.tab.gzPfam annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was functionally annotated against the Pfam database (version 34) using hmmsearchBash script used: hmmsearch --cut_tc --domtblout G3PA_depth.Pfam34.domtblout.tab Pfam_34.0/Pfam-A.hmm G3_depth_assembledReads.id99.fastaG3PA_depth.raw.est_counts.csv.gzEstimated counts of transcripts from G3 depth profile samples mapped to G3 depth profile assembly  (G3_depth.assembledReads.id99.fasta.gz)The code for processing, assemblying, and mapping reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded but the G3 depth profiles were assembled unstranded so the reads were mapped against the assembly unstranded The G3 depth profile reads were mapped against the G3 depth profile assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant -i NPac.G3PA_depth.id99.nt.idx -o ${SAMPLE} --threads=${N_THREADS} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> G3PA_depth.${SAMPLE}.kallisto.logG2PA_incubations.est_counts.csv.gzEstimated counts of transcripts from G2 incubation samples mapped to G2 surface assemblyTrimmed, quality-controlled reads from the G2 incubations were mapped to the G2 assembly (Gradients2.MGL1704.PA.assemblies.tar.gz)The code for processing reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded; the G2 surface samples were assembled reverse stranded; the G2 incubations were mapped reverse stranded against the G2 surface samplesThe G2 incubation reads were mapped against the G2 surface assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant --rf-stranded -i NPac.G2PA.bf100.id99.nt.idx -o ${SAMPLE} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> ${SAMPLE}.kallisto.logAt 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedchlorophyllA_g2Incubations.csvChlorophyll a measurements from the G2 incubations Samples for metatranscriptomes were collected after zero and 96 hours from the nutrient amendment experiments, along with onboard fluorometer measurements of chlorophyll a.Depth refers to the depth at which water was collected for the incubations. At 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedThe following datasets contain the transcripts per million per Pfam per taxonomic bin per sample:G1PA.tpm_counts.csv.gz G1 surfaceG2PA.tpm_counts.csv.gz G2 surfaceG3PA.tpm_counts.csv.gz G3 surfaceD1PA.tpm_counts.csv.gz Aloha dielG3PA_diel.tpm_counts.csv.gz G3 dielG2PA_incubations.tpm_counts.csv.gz G2 nutrient amendment incubationsG3PA_depth.tpm_counts.csv.gz G3 depth profilesTranscripts per million were calculated with the following stepsDivide the estimated number of reads mapped to each contig by its nucleotide length (in kilobases) to generate reads per kilobase (RPK).Sum the RPK by species and sample, then divide by one million to generate a conversion factor.Divide the RPK by the conversion factor to calculate transcripts per million per contig.Sum transcripts per million by Pfam for each species and sample.Transcript per million counts are only provided for annotated Pfams Transcripts per million were calculated using estimated counts outputted from kallistoEstimated count files for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereEstimated count files for the G2 nutrient amendment incubation and G3 depth profile samples can be found in this repository: G2PA_incubations.est_counts.csv.gz and G3PA_depth.raw.est_counts.csv.gz respectivelyTranscripts per million were calculated per taxonomic binTaxonomic annotations were generated using diamond last common ancestor and MarFERReTTaxonomic annotations for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the taxonomic annotations for the G2 surface assembly can also found in the previous repository: NPac.G2PA.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations the G3 depth profile samples can be found in this repository: NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTranscripts per million were calculated per PfamPfam annotations were generated using hmmsearchPfam annotations for the G1-G3 surface, Aloha diel, and G3 diel samples were generated using version 35 of the Pfam database and can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the Pfam annotations for the G2 surface assembly were generated using version 35 of the Pfam database and can also found in the previous repository: G2PA.Pfam35.domtblout.tab.gzPfam annotations the G3 depth profile samples were generated using version 34 of the Pfam database can be found in this repository: G3PA_depth.Pfam34.domtblout.tab.gz

Authors

  • Thomas, Elaina ;
  • Groussman, Mora Jove ;
  • Coesel, Sacha ;
  • Armbrust, Virginia
1 Citation0 Mentions85% FAIR0.9 Dataset Index
10.5281/zenodo.145190692024

MarPRISM

Data used to develop and test MarPRISM, a model to predict the in situ trophic mode of marine protists. To examine the in situ activity of protists, Lambert et al., 2022 developed a machine learning model to predict the trophic mode of marine protist species based on gene expression from metatranscriptomes.Recent studies (Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021) identified that a number of the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) transcriptomes used for training the Lambert model have a low number of sequences and/or high contamination.trainingData_withContam_withLowSeqs.csv.gz: training data used by Lambert et al. (2022), includes contaminated and low-sequence entries. trainingDataMarPRISM.csv.gz: training data with contaminated and low-sequence entries removed, training data used for MarPRISM.Transcriptomes were removed from the training data and used for testing that had less than 1200 total sequences, less than 500 total assigned Pfam domains, and/or greater than 50% contamination from non-target organisms.trainingDataMarPRISM_binary.csv.gz: trainingDataMarPRISM.csv.gz but TPM values greater than 0 were converted to 1 in order to determine whether a model could be built based on the binary expression of Pfams rathern than continuous expression values. trainingData_micromonasMixToPhot.csv.gz: trainingDataMarPRISM.csv.gz but with any transcriptomes from Micromonas strains that were originally labeled mixotrophic switched to phototrophic. This dataset was created to determine the effect of Micromonas in the training data, as well as the effect of permuting trophic mode labels in the training data as there may be errors in some of the trophic mode labels.The training datasets are unbalanced with more phototrophic transcriptomes than heterotrophic and mixotrophic transcriptomes.So feature selection and hyperparameter search was conducted on versions of the training datasets with phototrophic transcriptomes randomly undersampled. trainingData_withContam_withLowSeqs_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_withContam_withLowSeqs_140phot.zip: 140 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingData_withContam_withLowSeqs.csv.gztrainingData_contamLowSeqsRemoved_50phot.zip: 50 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_80phot.zip: 80 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_100phot.zip: 100 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_120phot.zip: 120 phototrophic transcriptomes, all of the mixotrophic and heterotrophic transcriptomes in trainingDataMarPRISM.csv.gztrainingData_contamLowSeqsRemoved_binary_50phot.zip: same as trainingData_contamLowSeqsRemoved_50phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_80phot.zip: same as trainingData_contamLowSeqsRemoved_80phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_100phot.zip: same as trainingData_contamLowSeqsRemoved_100phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_binary_120phot.zip: same as trainingData_contamLowSeqsRemoved_120phot.zip but TPM values > 0 were converted to 1trainingData_contamLowSeqsRemoved_micromonasMixToPhot_50phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 50 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_80phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 80 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_100phot.zip: after converting mixotrophy labels for any Micromonas strains in trainingDataMarPRISM.csv.gz to phototrophy, 100 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomestrainingData_contamLowSeqsRemoved_micromonasMixToPhot_120phot.zip: after converting Micromonas mixotrophy labels in trainingDataMarPRISM.csv.gz to phototrophy, 120 phototrophic transcriptomes were randomly selected along with all of the mixotrophic and heterotrophic transcriptomesFeature selection was run on the training datasets after undersampling phototrophic transcriptomes. Feature selection was run for both XGBoost and Random Forest models.  MarPRISM_featurePfams.csv.gz: MarPRISM feature Pfams, XGBoost model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_rfModel_rfFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgRFFeatures.csv.gz: Union of XGBoost and Random Forest feature Pfams, contaminated and low-sequence entries removed from training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: XGBoost model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsIncluded_xgModel_xgFeatures.csv.gz: Random Forest model, contaminated and low-sequence entries included in training dataExtracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_binary.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, TPM values > 0 were converted to 1Extracted_Pfams_contaminationLowSeqsRemoved_xgModel_xgFeatures_micromonasMixToPhot.csv.gz: XGBoost model, contaminated and low-sequence entries removed from training data, mixotrophy labels for any Micromonas strains were converted to phototrophyTranscriptomes not included in the training data and not from the MMETSP were used to test MarPRISM. testTranscriptomes.csv.gz: Transcript per million counts by Pfam ID for these test transcriptomesSome of these transcriptomes were processed and used for testing by Lambert et al., 2022, while other transcriptomes, from Pterosperma cristatum, Amphora coffeaeformis, Chaetoceros sp., and Cylindrotheca closterium, were newly added for testing. Transcripts per million for the latter transcriptomes derived from publicly salmon mappings.Pfam annotations for the newly added transcriptomes were generated through the following: the publicly available assembled transcriptomes for each species was six-frame translated with transeq , the longest reading frame (minimum 100 amino acid length) was selected for each contig, the longest contig was compared to the Pfam database version 34 with hmmsearch, and the Pfam annotation with the best bitscore for each contig was retained (e-value < 1e-05).Transcripts per million were summed by Pfam. testTranscriptomes.xlsx.gz: Accession IDs and references for these test transcriptomes; transcriptome culture conditions. Transcriptomes that were excluded from the training data for MarPRISM due to having high contamination or low-sequence abundance were also used to test MarPRISM. testTranscriptomes_MMETSP.csv.gz: Transcript per million counts by Pfam ID for these excluded transcriptomestestTranscriptomes_MMETSP.xlsx.gz: MMETSP IDs and references for these excluded transcriptomes. Contamination and low-sequence entries were identified by Groussman et al., 2023; Lasek-Nesselquist and Johnson, 2019; Van Vlierberghe et al., 2021 and curated by Groussman et al., 2023. Identity of ribosomal sequences was analyzed by Groussman et al., 2023.qc_flag_Groussman: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.num_sequences_Groussman: Number of sequences in original sequence file.num_pfams_Groussman: Number of Pfam domains identified in protein sequences.flag_Lasek: Flag notes from Lasek-Nesselquist and Johnson, 2019; CONTAM NOTED; ciliate samples reported as contaminated in this study.flag_VanVlierberghe: Flag for a high level of estimated contamination from 'flag_VanVlierberghe';CONTAM_50PCT; contamination percentages over 50%:flag_ribosomalContamination_Groussman: Flag for a high level of estimated contamination, from ‘ribosomal_contam_pct_Groussman'; CONTAM_50PCT; contamination percentages over 50%. ribosomal_contam_pct_Groussman: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity. Ribosomal taxonomy of most abundant contaminant: For entries with greater than 50% ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, the taxonomic identity of the most abundant ribosomal protein sequences not identified as the recorded identity of the transcriptome in the MMETSP.Expected taxonomy of transcriptome: recorded identity in MMETSP.

Authors

  • Thomas, Elaina ;
  • Groussman, Mora Jove ;
  • Coesel, Sacha ;
  • Armbrust, Virginia
1 Citation0 Mentions81% FAIR0.9 Dataset Index
10.5281/zenodo.145189012024

Gradients 1-3 polyA-selected transcripts per million, Gradients 3 depth profile polyA-selected processed metatranscriptomes

G3_depth.assembledReads.id99.fasta.gz G3 depth profiles amino acid assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the cluster representativesThe code for processing and assemblying reads can be found hereAmino acid formatNPac.G3PA_depth.bf100.id99.nt.fasta.gzG3 depth profiles nucleotide assemblyReplicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinityTranscripts are reverse stranded but were assembled unstrandedThe assembly from each set of replicates was concatenatedEach contig was six-frame translatedLongest reading frame was selected for each contig (minimum amino acid length 100)Longest reading frame amino acid sequences were clustered at 99% amino acid identityThe contigs in this file are the nucleotide-encoded contigs that the amino acid-encoded cluster representatives are from The code for processing and assemblying reads can be found hereNucleotide formatNPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was taxonomically annotated using diamond last common ancestor and MarFERReTBash script used: diamond blastp --no-unlink -t ~/ -b 100 -c 1 -p 32 -d /mnt/nfs/projects/marferret/v1/data/marmicrodb/dmnd/MarFERReT.v1.1.MMDB.combined.dmnd -e 1e-5 --top 10 -f 102 -q G3_depth_assembledReads.id99.fasta -o NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tabG3PA_depth.Pfam34.domtblout.tab.gzPfam annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz)G3_depth.assembledReads.id99.fasta.gz was functionally annotated against the Pfam database (version 34) using hmmsearchBash script used: hmmsearch --cut_tc --domtblout G3PA_depth.Pfam34.domtblout.tab Pfam_34.0/Pfam-A.hmm G3_depth_assembledReads.id99.fastaG3PA_depth.raw.est_counts.csv.gzEstimated counts of transcripts from G3 depth profile samples mapped to G3 depth profile assembly  (G3_depth.assembledReads.id99.fasta.gz)The code for processing, assemblying, and mapping reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded but the G3 depth profiles were assembled unstranded so the reads were mapped against the assembly unstranded The G3 depth profile reads were mapped against the G3 depth profile assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant -i NPac.G3PA_depth.id99.nt.idx -o ${SAMPLE} --threads=${N_THREADS} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> G3PA_depth.${SAMPLE}.kallisto.logG2PA_incubations.est_counts.csv.gzEstimated counts of transcripts from G2 incubation samples mapped to G2 surface assemblyTrimmed, quality-controlled reads from the G2 incubations were mapped to the G2 assembly (Gradients2.MGL1704.PA.assemblies.tar.gz)The code for processing reads can be found hereKallisto was used to map and outputs estimated counts (est_counts)The transcripts are reverse stranded; the G2 surface samples were assembled reverse stranded; the G2 incubations were mapped reverse stranded against the G2 surface samplesThe G2 incubation reads were mapped against the G2 surface assembly clustered at 99% amino acid identity in nucleotide spaceBash script used: kallisto quant --rf-stranded -i NPac.G2PA.bf100.id99.nt.idx -o ${SAMPLE} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> ${SAMPLE}.kallisto.logAt 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedchlorophyllA_g2Incubations.csvChlorophyll a measurements from the G2 incubations Samples for metatranscriptomes were collected after zero and 96 hours from the nutrient amendment experiments, along with onboard fluorometer measurements of chlorophyll a.Depth refers to the depth at which water was collected for the incubations. At 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe addedAt 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe addedAt 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe addedThe following datasets contain the transcripts per million per Pfam per taxonomic bin per sample:G1PA.tpm_counts.csv.gz G1 surfaceG2PA.tpm_counts.csv.gz G2 surfaceG3PA.tpm_counts.csv.gz G3 surfaceD1PA.tpm_counts.csv.gz Aloha dielG3PA_diel.tpm_counts.csv.gz G3 dielG2PA_incubations.tpm_counts.csv.gz G2 nutrient amendment incubationsG3PA_depth.tpm_counts.csv.gz G3 depth profilesTranscripts per million were calculated with the following stepsDivide the estimated number of reads mapped to each contig by its nucleotide length (in kilobases) to generate reads per kilobase (RPK).Sum the RPK by species and sample, then divide by one million to generate a conversion factor.Divide the RPK by the conversion factor to calculate transcripts per million per contig.Sum transcripts per million by Pfam for each species and sample.Transcript per million counts are only provided for annotated Pfams Transcripts per million were calculated using estimated counts outputted from kallistoEstimated count files for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereEstimated count files for the G2 nutrient amendment incubation and G3 depth profile samples can be found in this repository: G2PA_incubations.est_counts.csv.gz and G3PA_depth.raw.est_counts.csv.gz respectivelyTranscripts per million were calculated per taxonomic binTaxonomic annotations were generated using diamond last common ancestor and MarFERReTTaxonomic annotations for the G1-G3 surface, Aloha diel, and G3 diel samples can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the taxonomic annotations for the G2 surface assembly can also found in the previous repository: NPac.G2PA.MarFERReT_v1.1_MMDB.lca.tab.gzTaxonomic annotations the G3 depth profile samples can be found in this repository: NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gzTranscripts per million were calculated per PfamPfam annotations were generated using hmmsearchPfam annotations for the G1-G3 surface, Aloha diel, and G3 diel samples were generated using version 35 of the Pfam database and can be found hereTranscripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the Pfam annotations for the G2 surface assembly were generated using version 35 of the Pfam database and can also found in the previous repository: G2PA.Pfam35.domtblout.tab.gzPfam annotations the G3 depth profile samples were generated using version 34 of the Pfam database can be found in this repository: G3PA_depth.Pfam34.domtblout.tab.gz

Authors

  • Thomas, Elaina ;
  • Groussman, Mora Jove ;
  • Coesel, Sacha ;
  • Armbrust, Virginia
2 Citations0 Mentions85% FAIR1.3 Dataset Index
10.5281/zenodo.145190702024

The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts (Version: 0.92)

This data continues with the development of the NPEGC Trinity de novo metatranscriptome assemblies from the protein data repository of The North Pacific Eukaryotic Gene Catalog. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:NPac.G1PA.bf100.id99.nt.fasta.gzNPac.G2PA.bf100.id99.nt.fasta.gzNPac.G3PA.bf100.id99.nt.fasta.gzNPac.G3PA_diel.bf100.id99.nt.fasta.gzNPac.D1PA.bf100.id99.nt.fasta.gzA full description of this data is published in Scientific Data, available here: The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Please cite this publication if your research uses this data:Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Scientific Data, 11(1), 1161.These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3Key processing steps are sampled below with links to the detailed code on the main github code repository: https://github.com/armbrustlab/NPac_euk_gene_catalogCode used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: NPEGC.nt_kallisto_counts.shThere are two main steps:1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts2. Map the short reads from environmental samples back to the assembly indexAs generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.The code in this template script was used for each project: aggregate_kallisto_counts.RThe output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: G1PA.raw.est_counts.csv.gzG2PA.raw.est_counts.csv.gzG3PA.raw.est_counts.csv.gzG3PA_diel.raw.est_counts.csv.gzD1PA.raw.est_counts.csv.gz

Authors

  • Groussman, Ryan ;
  • Coesel, Sacha ;
  • Armbrust, E. Virginia
3 Citations0 Mentions69% FAIR1.5 Dataset Index
10.5281/zenodo.138268202024

The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts (Version: 0.92)

This data continues with the development of the NPEGC Trinity de novo metatranscriptome assemblies from the protein data repository of The North Pacific Eukaryotic Gene Catalog. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:NPac.G1PA.bf100.id99.nt.fasta.gzNPac.G2PA.bf100.id99.nt.fasta.gzNPac.G3PA.bf100.id99.nt.fasta.gzNPac.G3PA_diel.bf100.id99.nt.fasta.gzNPac.D1PA.bf100.id99.nt.fasta.gzA full description of this data is published in Scientific Data, available here: The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Please cite this publication if your research uses this data:Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Scientific Data, 11(1), 1161.These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3Key processing steps are sampled below with links to the detailed code on the main github code repository: https://github.com/armbrustlab/NPac_euk_gene_catalogCode used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: NPEGC.nt_kallisto_counts.shThere are two main steps:1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts2. Map the short reads from environmental samples back to the assembly indexAs generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.The code in this template script was used for each project: aggregate_kallisto_counts.RThe output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: G1PA.raw.est_counts.csv.gzG2PA.raw.est_counts.csv.gzG3PA.raw.est_counts.csv.gzG3PA_diel.raw.est_counts.csv.gzD1PA.raw.est_counts.csv.gz

Authors

  • Groussman, Ryan ;
  • Coesel, Sacha ;
  • Armbrust, E. Virginia
1 Citation0 Mentions79% FAIR0.8 Dataset Index
10.5281/zenodo.105704482024

Discrete Flow Cytometry of Depth Profile Samples from TN413 (2023) Using a BD Influx Cell Sorter (Version: 1.0)

The dataset consists of BD Influx-based analysis of phytoplankton populations from discrete flow cytometry data collected underway during the University of Washington School of Oceanography 2023 undergraduate senior thesis cruise (TN413) oceanographic research cruise from Hawaii to Fiji. The data consists of cell abundance, cell size (equivalent spherical diameter), carbon quota, and carbon biomass for heterotrophic bacteria, picophytoplankton populations, namely the cyanobacteria Prochlorococcus and Synechococcus, and small eukaryotic phytoplankton (<5 μm ESD). Time is in UTC format, latitude and longitude are in decimal degrees, and depth is in meters. Further information can be found here: https://github.com/fribalet/FCSplankton

Authors

  • Cain, Kelsy ;
  • Ribalet, François ;
  • Armbrust, Virginia
0 Citations0 Mentions79% FAIR0.6 Dataset Index
10.5281/zenodo.138212072024