Focused Crawling: A New Approach to Topic-Specific Web Resource Discovery
Explore this paper's citation graph
Summary
A new hypertext resource discovery system called a Focused Crawler that is robust against large perturbations in the starting set of URLs, and capable of exploring out and discovering valuable resources that are dozens of links away from the start set, while carefully pruning the millions of pages that may lie within this same radius.
- Type
- article
- Published
- 1999-05-17
- Cited by
- 1,839
- References
- 39
- Access
- Open access
- OpenAlex
- https://openalex.org/W1489992655
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:206134284
Keywords
Crawling, Web crawler, Computer science, Focused crawler, World Wide Web
References
- Internet Agents: Spiders, Wanderers, Brokers, and 'Bots
- Moving Up the Information Food Chain: Deploying Softbots on the World Wide Web
- Human Performance on Clustering Web Pages
- Distributed Hypertext Resource Discovery Through Examples
- A comparison of event models for naive bayes text classification
- A Technique for Measuring the Relative Size and Overlap of Public Web Search Engines
- Surfing the Web Backwards
- Authoritative sources in a hyperlinked environment
- Finding Related Pages in the World Wide Web
- Techniques for Disaggregating Centrality Scores in Social Networks
- The Shark-Search Algorithm. An Application: Tailored Web Site Mapping
- A Smart Itsy Bitsy Spider for the Web
- Automatic Resource Compilation by Analyzing Hyperlink Structure and Associated Text
- Web Watcher: A Tour Guide for the World Wide Web
- Social Network Analysis
- A new status index derived from sociometric analysis
- Efficient Crawling Through URL Ordering
- Dynamic Reference Sifting: A Case Study in the Homepage Domain
- Information Retrieval in the World-Wide Web: Making Client-Based Searching Feasible
- The Quest for Correct Information on the Web: Hyper Search Engines
Cited by
- Finding and fighting search engine spam
- Searching Entities on the Web by Sample
- Semi-automatic creation of domain ontologies with centroid based crawlers
- Monitoring RSS Feeds Based on User Browsing Pattern
- Using Probabilistiv Argumentation System to Search and Classify Web Sites.
- Studying, developing, and experimenting contextual advertising systems
- Search Engine-Crawler Symbiosis
- World news finder: How we cope without the semantic web
- Semantic linkage of the invisible geospatial web
- Semantic Co-Browsing System Based on Contextual Synchronization on Peer-to-Peer Environment
- Information Retrieval and Extraction from the Web: the CROSSMARC approach
- Are Related Links Effective for Contextual Advertising? - A Preliminary Study
- Scaling Information Extraction to Large Document Collections
- Webometric network analysis : mapping cooperation and geopolitical connections between local government administration on the web
- Walking on a graph with a magnifying glass: stratified sampling via weighted random walks
- Recherche des objets complexes dans le Web structuré. (Searching complex data on the structured Web)
- Researcher homepage classification using unlabeled data
- Parallel crawler architecture and web page change detection
- Modern Applications of Machine Learning
- Sound, Music and Textual Associations on the World Wide Web
Related papers
- Research and Simulation of Improved Topic Web Crawler Algorithm based on Deep Learning
- Design and Implementation of a University Focused Crawler
- Design and Implementation of the Topic-Focused Crawler Based on Scrapy
- Learnable topic-specific web crawler
- Adaptive focused crawler based on tunneling and link analysis
- Research and Implementation of Intelligent Focused Crawler
- Learnable Focused Meta Crawling Through Web