NYtimes_train_test_set.hdf5

View Dataset
King Juan Carlos University

Description

The NYtimes dataset, part of the Bags of Words dataset from the UCI repository, comprises a collection of New York Times news articles represented as a bag of words. Each document in the dataset is associated with a set of word occurrences, where the dimensions represent unique words extracted from the articles. The dataset is organised as a document–word matrix, where each row corresponds to a document and each column corresponds to a word. The values in the matrix indicate the frequency of each word occurring in the respective document. Preprocessing steps include tokenization, removal of stopwords, and vocabulary truncation, with only words occurring more than ten times retained.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.5

FAIR Score

73%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Zenodo

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Artificial Intelligence

Field

Computer Science

Domain

Physical Sciences

Confidence Score

37%

Source

Scholar Data Model

Normalization Factors

FT

51.92

CTw

1.00

MTw

1.00