J4
• 论文 • 上一篇 下一篇
谷 峰,刘晨曦,吴扬扬
收稿日期:
修回日期:
出版日期:
发布日期:
通讯作者:
GU Feng,LIU Chen-xi,WU Yangyang
Received:
Revised:
Online:
Published:
Contact:
摘要: 提出了一种基于序列数据挖掘的中文网页候选特征的选择方法,并用于中文网页分类模型. 该方法运用改进的PAT树结构挖掘频繁出现在同一类中文网页中的字符串,通过净频率计算,挖掘出中文网页中频繁出现的有意义的词、短语、英文单词等,并结合CHI算法得到文本特征. 实验表明,该算法不仅能挖掘出传统方法所选择出的绝大部分特征,还能挖掘出一些有意义的、切词系统词库中没有的、能反映分类特点的人名,地名,新词、常用语、外文单词等.
关键词: 序列数据挖掘, pat树, 中文网页分类 , 频繁字串, 净频率
Abstract: Abstract: A method is proposed to select feature candidates from Chinese websites on the basis of sequential data mining, and it is used in the model of Chinese websites classification. This method uses improved PAT tree data structure to mine the frequent strings in the same class of Chinese websites, calculates the net frequency, mines frequent meaningful words, phrases, and English words from Chinese websites, and obtains text features with the help of the CHI algorithm. Experiments show that this algorithm not only mines most of the features selected by the traditional algorithm, but also mines some new meaningful personnames, placenames, new words, phrases, and foreign words.
Key words: chinese web page classification , frequent string, net frequency, pattree, sequential data mining
谷 峰,刘晨曦,吴扬扬 . 基于序列数据挖掘的中文网页特征选择方法[J]. J4, 2006, 41(3): 95-99 .
GU Feng,LIU Chen-xi,WU Yangyang . Chinese Web page feature selection method based on Sequential data mining[J]. J4, 2006, 41(3): 95-99 .
0 / / 推荐
导出引用管理器 EndNote|Reference Manager|ProCite|BibTeX|RefWorks
链接本文: http://lxbwk.njournal.sdu.edu.cn/CN/
http://lxbwk.njournal.sdu.edu.cn/CN/Y2006/V41/I3/95
Cited