RUS  ENG
Full version
JOURNALS // Computer Research and Modeling // Archive

Computer Research and Modeling, 2021 Volume 13, Issue 6, Pages 1317–1336 (Mi crm950)

MODELS OF ECONOMIC AND SOCIAL SYSTEMS

Bibliographic link prediction using contrast resampling technique

F. V. Krasnov, I. S. Smaznevich, E. N. Baskakova

NAUMEN R&D, 49A, Tatishcheva st., Yekaterinburg, 620028, Russian Federation

Abstract: The paper studies the problem of searching for fragments with missing bibliographic links in a scientific article using automatic binary classification. To train the model, we propose a new contrast resampling technique, the innovation of which is the consideration of the context of the link, taking into account the boundaries of the fragment, which mostly affects the probability of presence of a bibliographic links in it. The training set was formed of automatically labeled samples that are fragments of three sentences with class labels «without link» and «with link» that satisfy the requirement of contrast: samples of different classes are distanced in the source text. The feature space was built automatically based on the term occurrence statistics and was expanded by constructing additional features — entities (names, numbers, quotes and abbreviations) recognized in the text.
A series of experiments was carried out on the archives of the scientific journals «Law enforcement review» (273 articles) and «Journal Infectology» (684 articles). The classification was carried out by the models Nearest Neighbors, RBF SVM, Random Forest, Multilayer Perceptron, with the selection of optimal hyperparameters for each classifier.
Experiments have confirmed the hypothesis put forward. The highest accuracy was reached by the neural network classifier (95 %), which is however not as fast as the linear one that showed also high accuracy with contrast resampling (91–94 %). These values are superior to those reported for NER and Sentiment Analysis on comparable data. The high computational efficiency of the proposed method makes it possible to integrate it into applied systems and to process documents online.

Keywords: contrast resampling, citation analysis, data resampling, link prediction, text classification, artificial neural network.

UDC: 004.896, 004.584, 004.91, 519.688

Received: 30.07.2021
Revised: 14.09.2021
Accepted: 25.09.2021

DOI: 10.20537/2076-7633-2021-13-6-1317-1336



© Steklov Math. Inst. of RAS, 2024