Toxicity Detection Dataset in Twi Language

View Dataset
Ali, Gifty Suuk ;Mensah Kwabena, Patrick;Ayidzoe, Mighty Abra;Ayawli, Ben Belklisi Kwame

Description

This dataset contains 2,001 text entries labeled for toxicity classification. Each entry represents a user-generated comment along with an assigned toxicity label. The dataset is structured into two columns:COMMENT– A text field containing comments written primarily in Akan (Twi). These comments include expressions of gratitude, feedback, conversational messages, and general communication typical of social or online interactions.LABEL– A categorical variable indicating whether the comment is 'toxic' or 'non-toxic'.Current labels present in the dataset: 'non-toxic' (and any others present in the full file, if applicable).Key Features:• Total records: 2,001• Language: Primarily Akan (Twi)• Classification type: Binary toxicity classificationThere are no missing values (both columns have 2,001 non-null entries)Data types: ‘COMMENT’: string and ‘LABEL’`: stringThis dataset can support research in:• Toxic language detection in low-resource languages• Natural Language Processing (NLP) for African languages• Machine learning model training for text classification• Sociolinguistic analysis of online conversational contentThe File Format is CSV file: Toxicity_dataset.csvIt contains two columns: 'COMMENT' and ‘LABEL'

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.4

FAIR Score

69%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Mendeley Data

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Language and Linguistics

Field

Arts and Humanities

Domain

Social Sciences

Confidence Score

45%

Source

Scholar Data Model

Keywords

Toxicity

Normalization Factors

FT

63.46

CTw

1.00

MTw

1.00