Dataset of Paper Titled: A Semi-Automated Approach for Detecting Ambiguities in Software Requirements Using SpanBERT and Named Entity Recognition

Touseef Tahir

Description

This dataset supports the study titled "A Semi-Automated Approach for Detecting Ambiguities in Functional Requirements Using SpanBERT." The dataset consists of 425 original functional requirements collected from 16 diverse software domains, such as finance, healthcare, education, and e-commerce. These requirements were curated to represent realistic and domain-relevant software specification statements written in natural language.Each requirement in the dataset has been manually annotated for three types of linguistic ambiguities:Anaphoric ambiguity (e.g., unclear references like "it" or "they"),Coordination ambiguity (e.g., ambiguous use of conjunctions like "and" or "or"),Missing condition ambiguity (e.g., implied conditions not explicitly stated).These annotations were used to train and evaluate a semi-automated ambiguity detection approach based on SpanBERT, a transformer-based natural language processing model. This dataset is intended to support further research on automated ambiguity detection, requirements engineering, and natural language understanding in software engineering. It is especially valuable for researchers and practitioners aiming to improve the clarity of software requirement specifications and reduce ambiguity-related issues during development. How to citeF. Talha, T. Tahir, and T. Nadeem, “ A Semiautomated Approach for Detecting Ambiguities in Software Requirements Using SpanBERT and Named Entity Recognition,” Journal of Software: Evolution and Process 37, no. 8 (2025): e70041, https://doi.org/10.1002/smr.70041.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.4

FAIR Score

79%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Zenodo

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Language and Linguistics

Field

Arts and Humanities

Domain

Social Sciences

Confidence Score

50%

Source

Scholar Data Model

Normalization Factors

FT

63.46

CTw

1.00

MTw

1.00