Description
Descriptions of files generated in this study:- 2_members_subset.tsv : corresponding to members.tsv file from EggNOG database, but only containing the subset of gene families included in our analysis (i.e., those with at least 2000 genes and at least 10 taxa in each gene family). Contains no headers since the original TSV file did not. The tab-separated columns are taxonomic ID (NCBI taxonomic ID for bacteria is "2"), NOG ID (the identifier for the gene family in EggNOG database), number of genes, number of taxa, comma-separated list of gene IDs, and comma-separated list of taxonomic IDs. - 2_taxonomicGroupings.csv : Taxonomic grouping information for taxa in our dataset, extracted from NCBI taxonomy database. - 2_PPI.csv: PPI and COG mapping data for genes in the dataset, extracted from STRING database. - phylumID_log20_sit80.transfers.csv : the main CSV output file of the analysis, containing all the inferred inter-phylum HGT events with their associated information. Inferred direction column of the transfer is "1" if it is from taxa A to taxa B, "-1" if it is from taxa B to taxa A, and "0" if the direction could not be determined. In the filename log20 refers to the logging level for python logging module (20 corresponds to INFO level instead of DEBUG level), and sit80 refers to the filtering sequence identity threshold (80%) used for downstream analyses of sequence identities across the entire dataset. Other headers are self-explanatory, and the file is comma-separated. - classID_log20_sit80.transfers.csv and orderID_log20_sit80.transfers.csv are similar files for class and order-level HGT pairs respectively.- hgt_zscores.csv : contains the z-scores of the sequence identities, for each inferred inter-phylum HGT pair of genes, with respect to the distribution of sequence identities of all non-HGT pairs of genes between the same two phyla in the same gene family (NOG). - hgt_zscores_skipped.csv : information of inter-phylum HGT pairs from about 6% of the total dataset of gene families, for which the z-scores could not be calculated due to insufficient data (e.g., too few non-HGT pairs of genes between the same two phyla in the same gene family). For the sake of computational efficiency, once we don't find enough data to calculate the z-scores for a gene family for a pair of phyla, we skip the calculation of z-scores for all HGT pairs in that gene family for that phylum pair (marked as invalid_phylum_pair_cached in the file). - proteinid_assembly_mapping.csv : a mapping file for inter-phylum HGT pairs, that links each protein ID (corresponding to the gene IDs in our dataset) to its corresponding assembly ID, which can be used to retrieve additional information about the genome from which the protein was derived. We used this to check the completeness of the genomes involved in the inferred HGT events. - work_dir.tar.gz file contains the entire work directory of the analysis, including all the intermediate files, logs, and results of each step. -------------------------------------------------For further details on how the files were generated, please refer to the README.md file in the code repository of this project (link). The notebooks directory in the code repository contains the Jupyter notebooks that were used to analyze these files.For details on the structure of the work directory and the files contained in it, please refer to the readme file in the code repository of this project. Scripts used in this study (apart from the notebooks) are the same as those in the misc directory of the code repository, with documentation of how those scripts were run also present in the readme file.
Citations (0)
No citations found
Mentions (0)
No mentions found
Metrics Over Time
Publication Details
Subfield
Molecular Biology
Field
Biochemistry, Genetics and Molecular Biology
Domain
Life Sciences
Confidence Score
51%
Source
Scholar Data Model