您的位置:山东大学 -> 科技期刊社 -> 《山东大学学报(理学版)》

J4

• 论文 • 上一篇    下一篇

基于特征域词频的邮件过滤方法的研究

刘 慧1,2,马 军2,雷景生2,3,连 莉2   

  1. 山东经济学院计算机科学与技术学院,山东 济南 250014
  • 收稿日期:2006-03-29 修回日期:1900-01-01 出版日期:2006-10-24 发布日期:2006-10-24
  • 通讯作者: 刘 慧

Research on email filtering by the frequency of the terms in character fields

LIU Hui,MA Jun,LEI Jing-sheng,LIAN Li   

  1. School of Computer Science & Technology, Shandong Economic Univ., Jinan 250014, Shandong, China;
  • Received:2006-03-29 Revised:1900-01-01 Online:2006-10-24 Published:2006-10-24
  • Contact: LIU Hui

摘要: 出了根据邮件特征域信息和特征词频进行垃圾邮件过滤的新方法,并介绍在该方法中的文本特征选取、特征词典构造以及基于TF的权值计算等相关技术,以及改进的文本相似度计算概率模型.实验表明该方法在邮件过滤的查全率、查准率等几个性能评价指标上,比传统的Rocchio方法有了明显改善.

关键词: 垃圾邮件过滤, 特征域, 权值计算 , 词频, 特征词典

Abstract: A novel method for Email filtering is proposed based on the information of character fields and the frequency of the terms in the character fields. The techniques used in the method are discussed, which include selecting the characters of text documents, the constructing the character lexicons as well as the computation of the weights of the term frequency (TF). In addition, an improved probabilistic model for the computation of the similarity of among text documents is provided. Experiments show that the new method is better than traditional Rocchio method in terms of recall, precision and some other evaluation targets.

Key words: weight calculation , term frequency, character term lexicon, character field, spam filtering

No related articles found!
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!