Focused Crawling: A New Approach to Topic-Specific Web Resource Discovery

Explore this paper's citation graph

Summary

A new hypertext resource discovery system called a Focused Crawler that is robust against large perturbations in the starting set of URLs, and capable of exploring out and discovering valuable resources that are dozens of links away from the start set, while carefully pruning the millions of pages that may lie within this same radius.

Type
article
Published
1999-05-17
Cited by
1,839
References
39
Access
Open access

Keywords

Crawling, Web crawler, Computer science, Focused crawler, World Wide Web

References

Cited by

Related papers