Scinovex
article Open Access

A smart system for identifying predatory publishing platforms using random forest

Abstract

The credibility and accuracy of scientific publications are under jeopardy due to predatory publishing houses that publish dubious studies. Their influence extends to areas of politics, society, the economy, and health, and they have brought about the shadow side of academic publication. In light of their spread and potential effects, many detection methods have been devised; nevertheless, these approaches are labor-intensive and rely on human intervention. In this study, we presented a smart framework that can automatically identify predatory venues and their infractions by using several AI approaches. This tool would be useful for researchers, students, and readers. This effort makes a difference by way of the following producing a database including 9,866 journals labeled as genuine or predatory, and suggesting a smart system for determining the legitimacy or predatoriness of a venue, supported by suitable logic. Various feature representation techniques, seven ML and DL models—including SVM, KNN, NNs, LSTM, CNN, BERT, A Lite BERT, and ALBERT—were used to rate our framework. The CNN model achieved an F1 score of 0.96, putting it ahead of the other models in the article categorization challenge. The SVM model got the greatest micro F1 score of 0.67 for the provisioning task's suitable reasoning.

Spam and Phishing DetectionRandom forestComputer sciencePublishingArtificial intelligenceArt
Citations
0
FWCI
0.00
field-weighted impact
References
0
Percentile
25%
vs. same field & year
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

A smart system for identifying predatory publishing platforms using random forest · Scinovex