Sample of 376 article texts

View Dataset
Aubakirov, Shakarim

Description

The dataset comprises a collection of scientific articles, each represented by its full text and abstract, alongside the number of sentences in the abstract. The focus of the research utilizing this dataset is on optimizing the text summarization process, specifically honing in on the 'min_df' parameter, which is crucial for filtering terms in the summarization algorithm. Although the dataset contains various other fields, the analysis primarily utilized the article texts, abstract texts, and the count of sentences in the abstracts. This streamlined approach is aimed at enhancing the extractive summarization's effectiveness, judged by the ROUGE-1 score, a common metric for evaluating the quality of summarized texts. The objective is to fine-tune the summarization tool to produce high-quality summaries that are both informative and reflective of the original text, thereby improving the tool's utility in processing scientific documents.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.7

FAIR Score

65%

Citations

1

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Mendeley Data

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Biomedical Engineering

Field

Engineering

Domain

Physical Sciences

Confidence Score

90%

Source

Open Alex

Keywords

Text Extraction

Normalization Factors

FT

64.42

CTw

1.00

MTw

1.00