МАТНЛИ МАЪЛУМОТЛАРНИ ТАСНИФЛАШ УСУЛ ВА АЛГОРИТМЛАРИ ТАҲЛИЛИ
Keywords:
матн таснифи, табиий тилни қайта ишлаш, машинали ўқитиш, Naive Bayes, Decision Tree, Random Forest, Passive Aggressive, таснифлаш алгоритмлари.Abstract
This article is devoted to the analysis of existing methods and algorithms for
automatic classification of text data, in which the theoretical foundations and practical
applications of machine learning algorithms such as Decision Tree (DT), Random Forest (RF),
Naive Bayes (NB) and Passive Aggressive (PA) are described. The results of the study showed the
importance of choosing the optimal algorithm depending on the type, volume and sources of data.
In addition, the work also covers the stages of text pre-processing, which can have a significant
impact on the final result.
References
Sebastiani F. Machine Learning in Automated Text Categorization // ACM Computing Surveys. – 2002. – Vol. 34, No. 1. – P. 1–47.
Aggarwal C.C., Zhai C. A Survey of Text Classification Algorithms // Mining Text Data. – Springer, 2012. – P. 163–222.
Jurafsky D., Martin J.H. Speech and Language Processing. – 3rd ed. – Draft, 2020.
Manning C.D., Raghavan P., Schütze H. Introduction to Information Retrieval. – Cambridge: Cambridge University Press, 2008.
Pedregosa F. et al. Scikit-Learn: Machine Learning in Python // Journal of Machine Learning Research. – 2011. – Vol. 12. – P. 2825–2830.