Automated Organization Profile

Telefonica Research

Current S-Index

50.2

Sum of Dataset Indices for all datasets

Average Dataset Index per Dataset

2.4

Average Dataset Index per dataset

Total Datasets

21

Total datasets in this organization

Average FAIR Score

66.8%

Average FAIR Score per dataset

Total Citations

42

Total citations to the organization's datasets

Total Mentions

1

Total mentions of the organization's datasets

S-Index Interpretation

S-Index Over Time

Cumulative Citations Over Time

Cumulative Mentions Over Time

Datasets

CRAWDAD telefonica/mobilephoneuse

Detailed logs of mobile phone usage of 342 people over the course of ca. 4 weeks. Contains events from 25 physical and virtual sensors (e.g. app in foreground, received notifications, ambient noise level, semantic location, ...), frequent self reports of emotions (valence and arousal), and for some participants self-reported traits (Big 5 Traits, Boredom Proneness [BPS], Patient Health Questionnaire [PHQ8]). Sensors: acceleration, airplane mode, app in foreground, audio, battery, battery drain, cell tower details, data consumption, ambient light, ambient noise, notifications, notification center access, notification dismissed, notifications cleared, device orientation, phone calls, photos, proximity (screen covered or not), screen events (on, off, unlocked), screen orientation (portrait, landscape), semantic location (home, work, ...), sms events, user activity (sitting, in vehicle, ...), wifi details. Details can be found at https://sites.google.com/view/mobile-phone-use-datasetlast modified : 2019-04-29nickname : mobilephoneuseinstitution : telefonicarelease date : 2019-04-29date/time of measurement start : 2016-04-15date/time of measurement end : 2016-07-12collection environment : 342 mobile phone users recruited by an agency to match the demographics of the country of study: SpainTracesettelefonica/mobilephoneuse/MobilePhoneUseDetailed mobile phone use data (events, sensors, emotion self-reports, traits) of 342 people over the course of 342files: mpu.zipdescription: Detailed logs of mobile phone usage of 342 people over the course of ca. 4 weeks. Contains events from 25 physical and virtual sensors (e.g. app in foreground, received notifications, ambient noise level, semantic location, ...), frequent self reports of emotions (valence and arousal), and for some participants self-reported traits (Big 5 Traits, Boredom Proneness [BPS], Patient Health Questionnaire [PHQ8]). Sensors: acceleration, airplane mode, app in foreground, audio, battery, battery drain, cell tower details, data consumption, ambient light, ambient noise, notifications, notification center access, notification dismissed, notifications cleared, device orientation, phone calls, photos, proximity (screen covered or not), screen events (on, off, unlocked), screen orientation (portrait, landscape), semantic location (home, work, ...), sms events, user activity (sitting, in vehicle, ...), wifi details. Details can be found at https://sites.google.com/view/mobile-phone-use-datasetmeasurement purpose: Usage Characterization, Human Behavior Modelingmethodology: This data was collected as part of a dedicated study. Volunteers installed an application onto their phones to collect a wide range of mobile phone use events, sensor data, and self-reports.telefonica/mobilephoneuse/MobilePhoneUse Traces:

Authors

  • Pielot, Martin
1 Citation0 Mentions58% FAIR0.9 Dataset Index
10.21227/n4qt-63892022

Tracking Knowledge Propagation Across Wikipedia Languages

We present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions are associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.

Authors

  • Valentim, Rodolfo ;
  • Comarela, Giovanni ;
  • Park, Souneil ;
  • Saez-Trumper, Diego
0 Citations0 Mentions69% FAIR0.4 Dataset Index
10.5281/zenodo.44331362021

Tracking Knowledge Propagation Across Wikipedia Languages

We present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions are associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.

Authors

  • Valentim, Rodolfo ;
  • Comarela, Giovanni ;
  • Park, Souneil ;
  • Saez-Trumper, Diego
2 Citations1 Mention73% FAIR1.5 Dataset Index
10.5281/zenodo.44331372021

Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" (Version: 1)

Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences. Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy

Authors

  • Antigoni-Maria Founta ;
  • Constantinos Djouvas ;
  • Chatzakou, Despoina ;
  • Leontiadis, Ilias ;
  • Blackburn, Jeremy ;
  • Stringhini, Gianluca ;
  • Vakali, Athena ;
  • Sirivianos, Michael ;
  • Kourtellis, Nicolas
0 Citations0 Mentions50% FAIR0.3 Dataset Index
10.5281/zenodo.36785762020

Public Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" (Version: 2)

Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only here. retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only here. UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide here the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy

Authors

  • Antigoni-Maria Founta ;
  • Constantinos Djouvas ;
  • Chatzakou, Despoina ;
  • Leontiadis, Ilias ;
  • Blackburn, Jeremy ;
  • Stringhini, Gianluca ;
  • Vakali, Athena ;
  • Sirivianos, Michael ;
  • Kourtellis, Nicolas
1 Citation0 Mentions77% FAIR0.9 Dataset Index
10.5281/zenodo.36785592020

Public Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" (Version: 2)

Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only here. retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only here. UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide here the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy

Authors

  • Antigoni-Maria Founta ;
  • Constantinos Djouvas ;
  • Chatzakou, Despoina ;
  • Leontiadis, Ilias ;
  • Blackburn, Jeremy ;
  • Stringhini, Gianluca ;
  • Vakali, Athena ;
  • Sirivianos, Michael ;
  • Kourtellis, Nicolas
0 Citations0 Mentions77% FAIR0.4 Dataset Index
10.5281/zenodo.36785582020

Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" (Version: 2)

Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy

Authors

  • Antigoni-Maria Founta ;
  • Constantinos Djouvas ;
  • Chatzakou, Despoina ;
  • Leontiadis, Ilias ;
  • Blackburn, Jeremy ;
  • Stringhini, Gianluca ;
  • Vakali, Athena ;
  • Sirivianos, Michael ;
  • Kourtellis, Nicolas
0 Citations0 Mentions50% FAIR0.3 Dataset Index
10.5281/zenodo.37068662020

Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" (Version: 2)

Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy

Authors

  • Antigoni-Maria Founta ;
  • Constantinos Djouvas ;
  • Chatzakou, Despoina ;
  • Leontiadis, Ilias ;
  • Blackburn, Jeremy ;
  • Stringhini, Gianluca ;
  • Vakali, Athena ;
  • Sirivianos, Michael ;
  • Kourtellis, Nicolas
0 Citations0 Mentions50% FAIR0.3 Dataset Index
10.5281/zenodo.36785752020

Dataset for: "Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children"

Dataset for paper: Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children The dataset consists of five files:
1. groundtruth_videos.json: This is the ground truth dataset. We have 4797 manually annotated videos (1513 suitable, 929 disturbing, 419 restricted, and 1936 irrelevant). You can distinguish among the different labels by observing the 'classification_label' field.
2. elsagate_related_videos.json: Contains the data for 233K elsagate-related YouTube videos (1K seed and 232K recommended) that were obtained as described in the paper.
3. other_child_related_videos.json: Contains the data for 155K other child-related YouTube videos (2K seed and 153K recommended) that were obtained as described in the paper.
4. random_videos.json: Contains the data for 482K random YouTube videos (8K seed and 474K recommended) that were obtained as described in the paper.
5. popular_videos.json: Contains the data for 11K popular YouTube videos (500 seed and 10.5K recommended) that were obtained between November 18 and November 21, 2018, as described in the paper. For each video in all sets, you can check the predicted label of our classifier by observing the 'prediction' field.

Authors

  • Papadamou, Kostantinos ;
  • Papasavva, Antonis ;
  • Zannettou, Savvas ;
  • Blackburn, Jeremy ;
  • Kourtellis, Nicolas ;
  • Leontiadis, Ilias ;
  • Stringhini, Gianluca ;
  • Sirivianos, Michael
0 Citations0 Mentions79% FAIR0.4 Dataset Index
10.5281/zenodo.36327802020

Dataset for: "Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children"

Dataset for paper: Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children The dataset consists of five files:
1. groundtruth_videos.json: This is the ground truth dataset. We have 4797 manually annotated videos (1513 suitable, 929 disturbing, 419 restricted, and 1936 irrelevant). You can distinguish among the different labels by observing the 'classification_label' field.
2. elsagate_related_videos.json: Contains the data for 233K elsagate-related YouTube videos (1K seed and 232K recommended) that were obtained as described in the paper.
3. other_child_related_videos.json: Contains the data for 155K other child-related YouTube videos (2K seed and 153K recommended) that were obtained as described in the paper.
4. random_videos.json: Contains the data for 482K random YouTube videos (8K seed and 474K recommended) that were obtained as described in the paper.
5. popular_videos.json: Contains the data for 11K popular YouTube videos (500 seed and 10.5K recommended) that were obtained between November 18 and November 21, 2018, as described in the paper. For each video in all sets, you can check the predicted label of our classifier by observing the 'prediction' field.

Authors

  • Papadamou, Kostantinos ;
  • Papasavva, Antonis ;
  • Zannettou, Savvas ;
  • Blackburn, Jeremy ;
  • Kourtellis, Nicolas ;
  • Leontiadis, Ilias ;
  • Stringhini, Gianluca ;
  • Sirivianos, Michael
0 Citations0 Mentions50% FAIR0.3 Dataset Index
10.5281/zenodo.36327812020