Classification of News Web Documents Based on Structural Features

Virach Sornlertlamvanich, Hitoshi Isahara, Shisanu Tongchim

doi:10.1007/11816508_17

The motivation of this work comes from the need of a Thai web corpus for testing our information retrieval algorithm. Two collections of news web documents are gathered from two different Thai newspaper web sites. Our goal is to find a simple yet effective method to extract news articles from these web collections. We explore the use of machine learning methods to distinguish article pages from non-article pages, e.g. table of contents, advertisements. Then, the selected web articles are compared in a fine-grained manner in order to find informative structures. Both steps of information extraction utilize the structural features of web documents rather than the extracted keywords or terms. Thus, the inherent errors of word segmentation, one of the major problems in Thai text processing, do not affect to this method.

Classification of News Web Documents Based on Structural Features

説明

詳細情報詳細情報について

書き出し

問題の指摘

Classification of News Web Documents Based on Structural Features

説明

詳細情報 詳細情報について

書き出し

問題の指摘

詳細情報詳細情報について