Automated Organization ProfileTelefonica Research
Telefonica Research
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets in this organization
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the organization's datasets
Total Mentions
Total mentions of the organization's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 50.2 (sum of 21 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
Detailed logs of mobile phone usage of 342 people over the course of ca. 4 weeks. Contains events from 25 physical and virtual sensors (e.g. app in foreground, received notifications, ambient noise level, semantic location, ...), frequent self reports of emotions (valence and arousal), and for some participants self-reported traits (Big 5 Traits, Boredom Proneness [BPS], Patient Health Questionnaire [PHQ8]). Sensors: acceleration, airplane mode, app in foreground, audio, battery, battery drain, cell tower details, data consumption, ambient light, ambient noise, notifications, notification center access, notification dismissed, notifications cleared, device orientation, phone calls, photos, proximity (screen covered or not), screen events (on, off, unlocked), screen orientation (portrait, landscape), semantic location (home, work, ...), sms events, user activity (sitting, in vehicle, ...), wifi details. Details can be found at https://sites.google.com/view/mobile-phone-use-datasetlast modified : 2019-04-29nickname : mobilephoneuseinstitution : telefonicarelease date : 2019-04-29date/time of measurement start : 2016-04-15date/time of measurement end : 2016-07-12collection environment : 342 mobile phone users recruited by an agency to match the demographics of the country of study: SpainTracesettelefonica/mobilephoneuse/MobilePhoneUseDetailed mobile phone use data (events, sensors, emotion self-reports, traits) of 342 people over the course of 342files: mpu.zipdescription: Detailed logs of mobile phone usage of 342 people over the course of ca. 4 weeks. Contains events from 25 physical and virtual sensors (e.g. app in foreground, received notifications, ambient noise level, semantic location, ...), frequent self reports of emotions (valence and arousal), and for some participants self-reported traits (Big 5 Traits, Boredom Proneness [BPS], Patient Health Questionnaire [PHQ8]). Sensors: acceleration, airplane mode, app in foreground, audio, battery, battery drain, cell tower details, data consumption, ambient light, ambient noise, notifications, notification center access, notification dismissed, notifications cleared, device orientation, phone calls, photos, proximity (screen covered or not), screen events (on, off, unlocked), screen orientation (portrait, landscape), semantic location (home, work, ...), sms events, user activity (sitting, in vehicle, ...), wifi details. Details can be found at https://sites.google.com/view/mobile-phone-use-datasetmeasurement purpose: Usage Characterization, Human Behavior Modelingmethodology: This data was collected as part of a dedicated study. Volunteers installed an application onto their phones to collect a wide range of mobile phone use events, sensor data, and self-reports.telefonica/mobilephoneuse/MobilePhoneUse Traces:
Authors
- Pielot, Martin
We present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions are associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.
Authors
- Valentim, Rodolfo ;
- Comarela, Giovanni ;
- Park, Souneil ;
- Saez-Trumper, Diego
We present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions are associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.
Authors
- Valentim, Rodolfo ;
- Comarela, Giovanni ;
- Park, Souneil ;
- Saez-Trumper, Diego
Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences. Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy
Authors
- Antigoni-Maria Founta ;
- Constantinos Djouvas ;
- Chatzakou, Despoina ;
- Leontiadis, Ilias ;
- Blackburn, Jeremy ;
- Stringhini, Gianluca ;
- Vakali, Athena ;
- Sirivianos, Michael ;
- Kourtellis, Nicolas
Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only here. retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only here. UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide here the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy
Authors
- Antigoni-Maria Founta ;
- Constantinos Djouvas ;
- Chatzakou, Despoina ;
- Leontiadis, Ilias ;
- Blackburn, Jeremy ;
- Stringhini, Gianluca ;
- Vakali, Athena ;
- Sirivianos, Michael ;
- Kourtellis, Nicolas
Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only here. retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only here. UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide here the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy
Authors
- Antigoni-Maria Founta ;
- Constantinos Djouvas ;
- Chatzakou, Despoina ;
- Leontiadis, Ilias ;
- Blackburn, Jeremy ;
- Stringhini, Gianluca ;
- Vakali, Athena ;
- Sirivianos, Michael ;
- Kourtellis, Nicolas
Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy
Authors
- Antigoni-Maria Founta ;
- Constantinos Djouvas ;
- Chatzakou, Despoina ;
- Leontiadis, Ilias ;
- Blackburn, Jeremy ;
- Stringhini, Gianluca ;
- Vakali, Athena ;
- Sirivianos, Michael ;
- Kourtellis, Nicolas
Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here. The Public version of the dataset can be found here hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Please cite the paper in any published work that uses any of these resources. @inproceedings{founta2018large,
title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},
author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},
booktitle={11th International Conference on Web and Social Media, ICWSM 2018},
year={2018},
organization={AAAI Press}
} For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy
Authors
- Antigoni-Maria Founta ;
- Constantinos Djouvas ;
- Chatzakou, Despoina ;
- Leontiadis, Ilias ;
- Blackburn, Jeremy ;
- Stringhini, Gianluca ;
- Vakali, Athena ;
- Sirivianos, Michael ;
- Kourtellis, Nicolas
Dataset for paper: Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children The dataset consists of five files:
1. groundtruth_videos.json: This is the ground truth dataset. We have 4797 manually annotated videos (1513 suitable, 929 disturbing, 419 restricted, and 1936 irrelevant). You can distinguish among the different labels by observing the 'classification_label' field.
2. elsagate_related_videos.json: Contains the data for 233K elsagate-related YouTube videos (1K seed and 232K recommended) that were obtained as described in the paper.
3. other_child_related_videos.json: Contains the data for 155K other child-related YouTube videos (2K seed and 153K recommended) that were obtained as described in the paper.
4. random_videos.json: Contains the data for 482K random YouTube videos (8K seed and 474K recommended) that were obtained as described in the paper.
5. popular_videos.json: Contains the data for 11K popular YouTube videos (500 seed and 10.5K recommended) that were obtained between November 18 and November 21, 2018, as described in the paper. For each video in all sets, you can check the predicted label of our classifier by observing the 'prediction' field.
Authors
- Papadamou, Kostantinos ;
- Papasavva, Antonis ;
- Zannettou, Savvas ;
- Blackburn, Jeremy ;
- Kourtellis, Nicolas ;
- Leontiadis, Ilias ;
- Stringhini, Gianluca ;
- Sirivianos, Michael
Dataset for paper: Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children The dataset consists of five files:
1. groundtruth_videos.json: This is the ground truth dataset. We have 4797 manually annotated videos (1513 suitable, 929 disturbing, 419 restricted, and 1936 irrelevant). You can distinguish among the different labels by observing the 'classification_label' field.
2. elsagate_related_videos.json: Contains the data for 233K elsagate-related YouTube videos (1K seed and 232K recommended) that were obtained as described in the paper.
3. other_child_related_videos.json: Contains the data for 155K other child-related YouTube videos (2K seed and 153K recommended) that were obtained as described in the paper.
4. random_videos.json: Contains the data for 482K random YouTube videos (8K seed and 474K recommended) that were obtained as described in the paper.
5. popular_videos.json: Contains the data for 11K popular YouTube videos (500 seed and 10.5K recommended) that were obtained between November 18 and November 21, 2018, as described in the paper. For each video in all sets, you can check the predicted label of our classifier by observing the 'prediction' field.
Authors
- Papadamou, Kostantinos ;
- Papasavva, Antonis ;
- Zannettou, Savvas ;
- Blackburn, Jeremy ;
- Kourtellis, Nicolas ;
- Leontiadis, Ilias ;
- Stringhini, Gianluca ;
- Sirivianos, Michael