摘要
为满足用户精确化和个性化获取信息的需要,通过分析Deep Web信息的特点,提出了一个可搜索不同主题Deep Web信息的爬虫框架。针对爬虫框架中Deep Web数据库发现和Deep Web爬虫爬行策略两个难题,分别提出了使用通用搜索引擎以加快发现不同主题的Deep Web数据库和采用常用字最大限度下载Deep Web信息的技术。实验结果表明了该框架采用的技术是可行的。
To satisfy people' s demand for getting precise and personal information, characteristics ofdeep web information are analyzed, and a framework of crawler for searching different subject information in deep web is put forward. To solve the difficult problems of deep web database discovery and deep web crawler crawling strategy, the technologies of discovering different subject deep web database quickly to use the universal search engine and downloading deep web information to the utmost by adopting the commonly used Chinese characters are proposed respectively. At last the experiment show that the framework is correct, and the technologies are feasible.
出处
《计算机工程与设计》
CSCD
北大核心
2010年第5期929-931,935,共4页
Computer Engineering and Design
基金
陕西省自然科学基金项目(2007F43)
关键词
深网
爬虫
搜索引擎
信息抽取
常用字
deep web
crawler
search engine
information extraction
commonly used Chinese characters