Automated Fuzzy Weighted Multi-Semantic Label Extraction of Persian News

نویسندگان

1 Department of Management, Faculty of Social Sciences and Economics, Alzahra University, Tehran, Iran

2 Associate Prof., Department of Management, Faculty of Social Sciences and Economics, Alzahra University, Tehran, Iran.

3 Department of Management, Faculty of Social Sciences and Economics, Alzahra University, Tehran, Iran

doi
10.22034/ijism.2025.1990502.1057
چکیده

 The vast number of online text documents and Semantic Web trends has increased researchers’ interest in semantic multi-label extraction. Research on the semantic extraction of multiple labels in Western and Eastern European languages is already well established. The challenge of machine reading and web-based knowledge extraction requires a scalable system to extract diverse information from large and heterogeneous collections. Hence, this study developed the multi-semantic fuzzy weight labeling system using natural language processing and supervised deep learning techniques. A long short-term memory (LSTM) was used for the extraction of labels, and the LSTM2 introduced by Yan, Wang, Gao, Zhang, Yang & Yin (2018) was used for the extraction of the label weights. To assess the degree of belonging of each document to each label, the resulting weights were modified according to their appearance in the document’s subject or in the Meta section of the web page, and the weights were normalized and fuzzified. Finally, the C-means fuzzy clustering algorithm was applied to the documents to assign each data point a degree of membership in relevant clusters. According to the results, the model's accuracy was 59.8%, indicating that the extraction of weighted key phrases and the semantic labeling of the text could be improved through supervised methods.