Performance Comparison of Turkish Web Pages Classification
Tarih
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Erişim Hakkı
Özet
Nowadays., web page classification is essential for efficient and fast search engines. There is an ever-increasing need for automatic classification techniques with higher classification accuracy. In this article., a performance comparison of existing Turkish language CNN models for web pages classification systems is performed. In more detail., the content of web pages is extracted first., then preprocessing steps that aim to detect the important parts and eliminate useless contents are used. Next., Bert word embedding is integrated to represent the texts by efficient numerical vectors. Finally., three state-of-the-art CNN models that fully support the Turkish language are investigated to find the best classifier. Overall., the three studied models obtained an acceptable performance while classifying the Turkish webpages., however., the third model was able to achieve slightly better than the other two models. © 2021 IEEE.










