Collocation Extraction Using Web Statistics
Hsin-Hsi Chen, Yi-Cheng Yu, Chih-Long Lin
Department of Computer Science and Information Engineering, National Taiwan University, Taipei, Taiwan
This paper mines collocations from two different web usage corpora, NTU proxy log and TTS search log. The precisions for NTU and TTS test data are 61.76% and 57.50%, respectively, by human judgment for 2% sampling of extracted collocations. For automatic evaluation, we submit extracted collocation to Google search engine, and the resulting page counts are used to compute the mutual information of the collocation. Experimental results show that total 43.27% and 42.65% of collocations mined from NTU and TTS corpora passed the examination of MIs.
Collocation Extraction, Mutual Information, Proxy/Search Logs, Web Mining, Web Statistics