The Center for Education and Research in Information Assurance and Security (CERIAS)

The Center for Education and Research in
Information Assurance and Security (CERIAS)

Dennis Fetterly - Microsoft

Students: Spring 2024, unless noted otherwise, sessions will be virtual on Zoom.

Using Statistical Analysis to Locate Spam Web Pages

Dec 08, 2004

Download: Video Icon MP4 Video Size: 215.6MB  
Watch on Youtube Watch on YouTube

Abstract

Commercial web sites are more dependant than ever on being placed prominently within the result pages returned by a search engine to be successful. "Spam" web pages are web pages that are created for the sole purpose of misleading search engines and misdirecting traffic to target sites. Certain classes of spam pages, in particular those that are machine-generated, diverge in some of their properties from the properties of web pages in general. As a result, these pages can be identified through statistical analysis. We have examined a variety of such properties, including linkage structure, page content, and page evolution, and have found that outliers in the statistical distributions of these properties are predominantly caused by web spam. Joint work with Mark Manasse and Marc Najork.

About the Speaker

Dennis Fetterly is a Technologist in Microsoft Research\'s Silicon Valley lab, which he joined in May, 2003. His research interests include a wide variety of web related topics including web crawling, the evolution and clustering of pages on the web, and identifying spam web pages.


Ways to Watch

YouTube

Watch Now!

Over 500 videos of our weekly seminar and symposia keynotes are available on our YouTube Channel. Also check out Spaf's YouTube Channel. Subscribe today!