Predictable by publication: discovery of early highly cited academic papers based on their own features
ISSN: 0737-8831
Article publication date: 6 February 2023
Issue publication date: 23 July 2024
Abstract
Purpose
Predicting highly cited papers can enable an evaluation of the potential of papers and the early detection and determination of academic achievement value. However, most highly cited paper prediction studies consider early citation information, so predicting highly cited papers by publication is challenging. Therefore, the authors propose a method for predicting early highly cited papers based on their own features.
Design/methodology/approach
This research analyzed academic papers published in the Journal of the Association for Computing Machinery (ACM) from 2000 to 2013. Five types of features were extracted: paper features, journal features, author features, reference features and semantic features. Subsequently, the authors applied a deep neural network (DNN), support vector machine (SVM), decision tree (DT) and logistic regression (LGR), and they predicted highly cited papers 1–3 years after publication.
Findings
Experimental results showed that early highly cited academic papers are predictable when they are first published. The authors’ prediction models showed considerable performance. This study further confirmed that the features of references and authors play an important role in predicting early highly cited papers. In addition, the proportion of high-quality journal references has a more significant impact on prediction.
Originality/value
Based on the available information at the time of publication, this study proposed an effective early highly cited paper prediction model. This study facilitates the early discovery and realization of the value of scientific and technological achievements.
Keywords
Acknowledgements
This study is supported by the Major Projects of National Social Science Foundation of China (Grant Number: 19ZDA349).
Citation
Tang, X., Zhou, H. and Li, S. (2024), "Predictable by publication: discovery of early highly cited academic papers based on their own features", Library Hi Tech, Vol. 42 No. 4, pp. 1366-1384. https://doi.org/10.1108/LHT-06-2022-0305
Publisher
:Emerald Publishing Limited
Copyright © 2023, Emerald Publishing Limited