Domain-based Abbreviation Expansion using Topic Modelling for Data Cleaning

نویسندگان

1 Department of Software Engineering, Faculty of Computer, University of Isfahan, Isfahan, Iran.

2 Department of Software Engineering, Faculty of Computer, University of Isfahan, Isfahan, Iran.

3 Department of Software Engineering, Faculty of Computer, University of Isfahan, Isfahan, Iran.

doi
10.22108/jcs.2024.138180.1140
چکیده

Data cleaning is a necessary process in data analytics and management and plays an essential role in obtaining more reliable data for further analysis and business. The most important challenge with textual unstructured data is how to invest and activate data cleaning, where a large amount of textual data is rapidly generated that urgently needs to be cleaned to be able to use it to obtain knowledge. In textual data, abbreviations are appearing more and more often in different datasets. The abbreviation disambiguation process can be considered a critical issue in data analysis and obtaining more reliable data. Many abbreviation expansion approaches have been proposed to handle this issue, but these approaches have never paid attention to the domain to which these abbreviations belong. To overcome this drawback, a domain-based abbreviation expansion method using topic modeling is proposed in this paper. In this method, topic modeling is applied first to the text, then the domain is determined based on the contribution of topics. This will reduce the search space. Finally, expansion is applied to abbreviations according to their domains. The proposed method has been validated by applying it to a COVID-19 tweets dataset and employing a logistic regression classifier. Two types of comparisons have been made, with the dataset itself before using the proposed method and with the results of the CrowdCorrect approach. The results show a clear improvement in precision and recall as well as the accuracy of the classifier.