为解决情报采集过程中竞争企业名录的更新问题, 提出了一种基于网络的竞争企业名录自动更新方法。该方法首先利用产品名称从企业索引中检索出相关的企业名列表, 采用LCS(Longest Common Substring)算法抽取企业名模式, 以“产品名+企业名模式”的形式重构查询。然后, 使用搜索引擎进行网页搜索, 再利用贝叶斯分类算法对搜索的网页过滤, 将过滤后的企业信息更新到企业名录中。实验结果显示, 系统P@10、P@20、P@30分别为73.4%, 68.4%, 65.2%, MAP@10、MAP@20、MAP@30分别达到66.2%, 58.9%, 52.5%, 结果说明该方法可以有效的实现竞争企业名录的自动更新。
We propose a Web-based approach to solving the update problem of competitive business directory in the process of intelligence collection in this paper.Firstly, the approach retrieves related business list from business index by using product names, and adopts business name pattern extraction approach based on LCS algorithm., and we construct a new query in the form of product and business name pattern.Then, we obtain the pages by search engine, and use bayes classifier algorithm to filter the pages.We update the business information filtered into the business directory.Experiment results show that P@10, P@20, P@30 are 73.4%, 68.4%, 65.2%, and MAP@10, MAP@20, MAP@30 are 66.2%, 58.9%, 52.5%, respectively.The result shows that the approach can effectively realize automatic update of competitive business directory.
[1]包昌火, 赵刚, 李艳, 等.竞争情报的崛起[J].情报学报, 2005, 24(1):3-19.
[2]易聪.基于Web挖掘的企业竞争情报系统构建研究[D].广州:华南理工大学, 2011.
[3]乔林.基于多关键词检索的企业竞争情报搜集方法研究[D].合肥:中国科学技术大学, 2006.
[4]吴晓伟, 徐福缘, 吴伟昶.基于神经网络的企业竞争对手分析[J].情报学报, 2004, 23(4):502-506.
[5]吴晓伟, 徐福缘, 宋文官.基于人际网络节点中心度的竞争对手分析[J].情报学报, 2006, 25(1):122-128.
[6]吴晓伟, 刘仲英, 李丹.竞争情报研究的创新途径——基于社会网络分析的观点[J].情报学报, 2008, 27(2):295-302.
[7]龙青云, 吴晓伟, 娜日.基于社会网分析的竞争对手及实证研究[J].情报科学, 2013, 31(1):134-140.
[8]张红芹, 鲍志彦.基于专利地图的竞争对手识别研究[J].情报科学, 2011, 29(12):1825-1829.
[9]王知津, 周鹏, 韩正彪.基于决策树算法的竞争对手识别模型研究[J].情报理论与实践, 2013, 36(3):1-5.
[10]郑重.基于聚类分析的企业动态竞争对手辨识[J].情报杂志, 2010, 29(8):148-151.
[11]Bao S, R Li, Y Yu, et al.Competitor mining with the web[J].IEEE Transactions on Knowledge and Data Engineering, 2008, 20(10):1297-1310.
[12]Zhongming Ma, Gautam Pant, Olivia R.L.Sheng.Mining competitor relationships from online news:a network-based approach[J].Electronic Commerce Research and Applications, 2011(10):418-427.
[13]Vaughan L, Wu G.Links to commercial web sites as a source of business information[J].Scientometrics, 2004, 60(3):487-496.
[14]Vaughan L, Gao Y, Kipp M.Why are Hyperlinks to Business Websites Created? A Content Analysis [J].Scientometrics, 2006, 67(2):291-300.
[15]Liwen, vayggab, Justin Y.Comparing business competition positions based on web co-link data:the global market vs the chinese market [J].Scientometrics, 2006, 68(3):611-628.
[16]雷静.汉语机构名的构成模式[C].全国第七届计算语言学联合学术会议, 2003, 91-96.
[17]孙艳, 周学广.基于粗糙集与贝叶斯决策的不良网页过滤研究[J].中文信息学报, 2012, 26(1):67-71.