Universiti Teknologi Malaysia Institutional Repository

Features discovery for web classification using support vector machine

Othman, Mohd. Shahizan and Mi Yusuf, Lizawati and J., Salim (2010) Features discovery for web classification using support vector machine. In: Proceedings - 2010 International Conference on Intelligent Computing and Cognitive Informatics, ICICCI 2010, 2010, Kuala Lumpur, Malaysia.

Full text not available from this repository.

Official URL: http://dx.doi.org/10.1109/ICICCI.2010.16


The ever fast-expanding web information resources pose a big challenge to internet users seeking the most relevant, latest and quality information. The sheer vast amount of web information has resulted in restructuring of the resources. Thus, an appropriate web classification method needs to be established in order for quality web information to be accessed. This paper intends to discuss the web document features that classify the web information resources. Six web document features have been identified which are text, meta tag and title (A), title and text (B), title (C), meta tag and title (D), meta tag (E) and text (F). The Support Vector Machine (SVM) method is used to classify the web document while four types of kernels namely: Radial Basis Function (RBF), linear, polynomial and sigmoid kernels was applied to test the accuracy of the classification. The studies show that the text, meta tag and title (A) features is the best features for classification of web document that employs the four kernels followed by the features on title and text (B) as well as the features on meta tag and title (C). The studies also found that the linear kernel is the best kernel in classifying the web document compared to the RBF, polynomial and sigmoid kernel.

Item Type:Conference or Workshop Item (Paper)
Uncontrolled Keywords:Support Vector Machine (SVM), web classification, web document
Subjects:Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Divisions:Computer Science and Information System (Formerly known)
ID Code:27993
Deposited By: Mrs Liza Porijo
Deposited On:30 Aug 2012 03:02
Last Modified:07 Feb 2017 07:21

Repository Staff Only: item control page