Show simple item record

dc.contributor.author Algiryage, N
dc.contributor.author Dias, G
dc.contributor.author Jayasena, S
dc.contributor.editor Chathuranga, D
dc.date.accessioned 2022-09-02T04:58:46Z
dc.date.available 2022-09-02T04:58:46Z
dc.date.issued 2018-05
dc.identifier.citation N. Algiryage, G. Dias and S. Jayasena, "Distinguishing Real Web Crawlers from Fakes: Googlebot Example," 2018 Moratuwa Engineering Research Conference (MERCon), 2018, pp. 13-18, doi: 10.1109/MERCon.2018.8421894. en_US
dc.identifier.uri http://dl.lib.uom.lk/handle/123/18862
dc.description.abstract Web crawlers are programs or automated scripts that scan web pages methodically to create indexes. Search engines such as Google, Bing use crawlers in order to provide web surfers with relevant information. Today there are also many crawlers that impersonate well-known web crawlers. For example, it has been observed that Google’s Googlebot crawler is impersonated to a high degree. This raises ethical and security concerns as they can potentially be used for malicious purposes. In this paper, we present an effective methodology to detect fake Googlebot crawlers by analyzing web access logs. We propose using Markov chain models to learn profiles of real and fake Googlebots based on their patterns of web resource access sequences. We have calculated log-odds ratios for a given set of crawler sessions and our results show that the higher the log-odds score, the higher the probability that a given sequence comes from the real Googlebot. Experimental results show, at a threshold log-odds score we can distinguish the real Googlebot from the fake. en_US
dc.language.iso en en_US
dc.publisher IEEE en_US
dc.relation.uri https://ieeexplore.ieee.org/document/8421894 en_US
dc.title Distinguishing real web crawlers from fakes: googlebot example en_US
dc.type Conference-Full-text en_US
dc.identifier.faculty Engineering en_US
dc.identifier.department Engineering Research Unit, University of Moratuwa en_US
dc.identifier.year 2018 en_US
dc.identifier.conference 2018 Moratuwa Engineering Research Conference (MERCon) en_US
dc.identifier.place Moratuwa, Sri Lanka en_US
dc.identifier.pgnos pp. 13-18 en_US
dc.identifier.proceeding Proceedings of 2018 Moratuwa Engineering Research Conference (MERCon) en_US
dc.identifier.email nilania@kln.ac.lk en_US
dc.identifier.email gihan@cse.mrt.ac.lk en_US
dc.identifier.email sanath@cse.mrt.ac.lk en_US
dc.identifier.doi 10.1109/MERCon.2018.8421894 en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record