USING ARTIFICIAL INTELLIGENCE TO AUTOMATE THE PROCESS OF COLLECTING AND ANALYZING DATA FROM ONLINE JOB POSTINGS
Keywords:
classification method, machine learning, neural network language models, natural language processing, information extraction, entity recognition.Abstract
The article discusses an approach to information extraction using online
learning based on determining the semantic proximity of sentence vectors and knowledge base
entities using neural network language models trained without a teacher on a large text corpus
of the subject area. A detailed review of modern supervised and unsupervised information
extraction methods is provided, which allow achieving acceptable quality in solving the problem
of analyzing current labor market requirements without the labor-intensive procedure of text
corpus tagging and without using rule-based approaches.
References
Al-Nabki W., Eduardo F., Enrique A., Laura F.-R. Improving Named Entity Recognition in Noisy User-Generated Text with Local Distance Neighbor Feature // Neurocomputing. – 2020. – Vol. 382. – P. 1–11.
Zhang Z., Iria J. A Novel Approach to Automatic Gazetteer Generation Using Wikipedia // Proceedings of the 2009 Workshop on the People’s Web Meets NLP, ACL-IJCNLP. – 2009. – P. 1–9.
Zahraa S.A., Mark C., Gholamreza H. Multi-Domain Evaluation Framework for Named Entity Recognition Tools // Computer Speech & Language. – 2017. – Vol. 43. – P. 34–55.
Rekia K., Yu Z., Weinan Zh., Ting L. CCG Supertagging via Bidirectional LSTM-CRF Neural Architecture // Neurocomputing. – 2017. – Vol. 283. – P. 31–37.
Wang Y., Tong H., Zhu Z., Li Y. Nested Named Entity Recognition: A Survey // ACM Transactions on Knowledge Discovery from Data. – 2022. – Online publication. – P. 1–29.
Hu X. et al. Location Reference Recognition from Texts: A Survey and Comparison // ACM Computing Surveys. – 2023. – Online publication. – P. 1–37.
Shao Y., Hardmeier C., Nivre J. Multilingual Named Entity Recognition Using Hybrid Neural Networks // The Sixth Swedish Language Technology Conference (SLTC). – 2016.
Zhiheng H., Wei X., Kai Yu. Bidirectional LSTM-CRF Models for Sequence Tagging // CoRR. – 2015. – arXiv:1508.01991.
Luan S., Anderson F. An Entity Resolution Approach Based on Word Embeddings and Knowledge Bases for Microblog Texts // Association for Computing Machinery. – 2021. – Article 53. – P. 1–8.
Kumarjeet P., Pramit M., Vaishali G. Named Entity Recognition Using Word2Vec // International Research Journal of Engineering and Technology (IRJET). – 2020. – Vol. 7, No. 9. – P. 1818–1820.
Debora N., Pikakshi M., Elisabetta F., Matteo P., Enza M. Learning to Adapt with Word Embeddings: Domain Adaptation of Named Entity Recognition Systems // Information Processing & Management. – 2021. – Vol. 58, No. 3.
Thien H.N., Barbara P., Ralph G. Semantic Representations for Domain Adaptation: A Case Study on the Tree Kernel-Based Method for Relation Extraction // Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing. – 2015. – P. 635–644.
Wenhao G., Xiao Y., Minhao Y., Kun H., Wenying P., Zexuan Z. MarkerGenie: An NLP-Enabled Text-Mining System for Biomedical Entity Relation Extraction // Bioinformatics Advances. – 2022. – Vol. 2, No. 1. – Article vbac035.
Wang Zc., Wang Zg., Li Jz. et al. Knowledge Extraction from Chinese Wiki Encyclopedias // Journal of Zhejiang University – Science C. – 2012. – Vol. 13. – P. 268–280.
Nayak T., Ng H. Effective Modeling of Encoder-Decoder Architecture for Joint Entity and Relation Extraction [Elektron resurs]. – 2019. DOI: 10.48550/arXiv.1911.09886.
Gensim Library [Elektron resurs]. – URL: – Murojaat sanasi: 10.02.2025.