Automated Author ProfileDias, Francisco
CHANGE - Global Change and Sustainability InstituteCentre for Ecology, Evolution and Environmental Changes (cE3c)
Dias, Francisco
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets for this author
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the author's datasets
Total Mentions
Total mentions of the author's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 1.0 (sum of 2 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
RECODE is a manually annotated corpus of ecological and taxonomic literature, aimed at training and fine-tuning models for automated extraction of occurrence and trait data from unstructured text. Documents present have been annotated and validated by experts familiar with the traits of the included taxa (currently spiders and insects). Furthermore, this dataset is goal-oriented and published simultaneously with a complementary R package (arete) and a practical case of its usage for both finetuning and validation.Table of ContentsThe following are all the elements contained in the archive ./recode.zip:recode, directory. Contains metadata.csv, containing the metadata for each entry in RECODE. Also subdivides into all available taxa: currently ./insecta and ./araneae. These directories subdivide by annotator if available. If not, .tsv files are placed under a ./all directory. Finally, annotation files are named by the focus taxa they belong to and their document ID.metadata.csv./araneae./insectar, directory. Contains the script used to calculate the values and plots present in the describing manuscript as well as all necessary files. These include two .csv tables containing taxonomic information on the species annotated in the dataset. These are supplied instead of being generated as part of the script as: 1) this data is prone to becoming outdated; 2) the method used to extract this information requires internet access.script_publish.Rtaxa_table_insects.csvtaxa_table_spiders.csvplots, directory. Contains all plots created in R that appear on the release version of the describing manuscript.FIG_1.pngFIG_3.pngFIG_4.pngFIG_5.pngFIG_6.png
Authors
- Veiga Branco, Vasco ;
- Cardoso, Pedro ;
- Correia, Luís ;
- Baranovičová, Lenka ;
- Lahin, Iiris ;
- Dias, Francisco ;
- Filipe, David
RECODE is a manually annotated corpus of ecological and taxonomic literature, aimed at training and fine-tuning models for automated extraction of occurrence and trait data from unstructured text. Documents present have been annotated and validated by experts familiar with the traits of the included taxa (currently spiders and insects). Furthermore, this dataset is goal-oriented and published simultaneously with a complementary R package (arete) and a practical case of its usage for both finetuning and validation.Table of ContentsThe following are all the elements contained in the archive ./recode.zip:recode, directory. Contains metadata.csv, containing the metadata for each entry in RECODE. Also subdivides into all available taxa: currently ./insecta and ./araneae. These directories subdivide by annotator if available. If not, .tsv files are placed under a ./all directory. Finally, annotation files are named by the focus taxa they belong to and their document ID.metadata.csv./araneae./insectar, directory. Contains the script used to calculate the values and plots present in the describing manuscript as well as all necessary files. These include two .csv tables containing taxonomic information on the species annotated in the dataset. These are supplied instead of being generated as part of the script as: 1) this data is prone to becoming outdated; 2) the method used to extract this information requires internet access.script_publish.Rtaxa_table_insects.csvtaxa_table_spiders.csvplots, directory. Contains all plots created in R that appear on the release version of the describing manuscript.FIG_1.pngFIG_3.pngFIG_4.pngFIG_5.pngFIG_6.png
Authors
- Veiga Branco, Vasco ;
- Cardoso, Pedro ;
- Correia, Luís ;
- Baranovičová, Lenka ;
- Lahin, Iiris ;
- Dias, Francisco ;
- Filipe, David