Modelling procedure while assessing the impact of news articles on cryptocurrency (Bitcoin) market movement

نویسندگان

1 Department of Statistics, Federal University, Otuoke, Nigeria

2 Department of Statistics, Nasarawa State University, Kefi, Nigeria.

3 Department of Statistics, Nasarawa State University, Kefi, Nigeria.

doi
10.22059/jcss.2025.393177.1139
چکیده

Background: Cryptocurrencies have a variety of unique qualities, from cutting-edge technology to highly secure architecture. Additionally, the ability to invest in cryptocurrency, as an asset or a function of its prosperity has made crypto-currencies attractive to venture capitalists, computer scientists, and statisticians.Aims: In this study, we concentrated on a collection of documents web-scrapped from the market section of CNBC, where each document is associated with a response variable.Methodology: These documents contain preprocessed words/terms of day-to-day reportage on cryptocurrency (Bitcoin). The corresponding response variables are the daily opening and closing price of Bitcoin prices. The Supervised Latent Dirichlet Allocation(sLDA), a statistical model of labeled documents, was used to analyze the textual data alongside their corresponding response variables, since our study aims to predict the response variable for unlabeled new documents.Results: Hidden Topics with their unique terms from the preprocessed articles were exposed through a Natural language processor. Mean absolute error (MAE), Mean absolute percentage error (MAPE), and Root mean square error (RMSE) graphs were constructed for the sLDA models with ‘k = 3,10,20,30,50,75,100 and 200 Topics’ values where the model with the best evaluation metric, was selected for prediction purpose.Conclusion: It was discovered that the sLDA model with k = 20. A posterior covariance matrix which shows the proportion of terms from the documents, making up a Topic. Coefficient values were generated in other to graphically visualize how important the discovered topics are and how they affect the market trend. Finally, the prediction of new labels (numeric-decoded closing prices) for the unlabeled documents was done and comparisons were made; the predicted labels follow a similar pattern to that of the time series closing price trend.