Scaling Information Extraction to Large Document Collections
Explore this paper's citation graph
Summary
This work reviews key approaches for scaling up information extraction, including using general-purpose search engines as well as indexing techniques specialized for information extraction applications, and highlights some of the opportunities and challenges.
- Type
- article
- Published
- 2005-01-01
- Cited by
- 48
- References
- 34
- OpenAlex
- https://openalex.org/W57058802
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:9887529
Keywords
Computer science, Scalability, Information extraction, Search engine indexing, Information retrieval
References
- BizTalk Server 2000 Business Process Orchestration.
- Focused Crawling: A New Approach to Topic-Specific Web Resource Discovery
- Foundations of Statistical Natural Language Processing
- Snowball: extracting relations from large plain-text collections
- Question Answering Using XML-Tagged Documents
- Maximum Entropy Markov Models for Information Extraction and Segmentation
- TREC Routing Experiments with the TRW/Paracel Fast Data Finder
- How to build a WebFountain: An architecture for very large-scale text analytics
- KnowItNow: Fast, Scalable Information Extraction from the Web
- A flexible learning system for wrapping tables and lists in HTML documents
- Querying text databases for efficient information extraction
- Question-answering by predictive annotation
- Indexing and Querying XML Data for Regular Path Expressions
- The Linguist’s Search Engine: An Overview
- Towards Terascale Semantic Acquisition
- Information extraction for enhanced access to disease outbreak reports
- Integrating DB and IR Technologies: What is the Sound of One Hand Clapping?
- Unsupervised named-entity extraction from the Web: An experimental study
- A search engine for natural language applications
- Semi-Markov Conditional Random Fields for Information Extraction
Cited by
- Big Data Quality Case Study Preliminary Findings, U.S. Army MEDCOM MODS
- Scalable Text Mining with Sparse Generative Models
- An arabic information management messaging system: Using information extraction for the proper flow of information within organizations
- Deriving a Web-Scale Common Sense Fact Database
- Information Extraction from Web-scale N-gram Data
- Geographically aware Web text mining
- Searching and ranking in entity-relationship graphs
- Quick-and-clean extraction of linked data entities from microblogs
- Harvesting, searching, and ranking knowledge on the web: invited talk
- Database and information-retrieval methods for knowledge discovery
- The YAGO-NAGA approach to knowledge discovery
- Information extraction as a filtering task
- Constructing efficient information extraction pipelines
- Entity categorization over large document collections
- YAGO: A Large Ontology from Wikipedia and WordNet
- From information to knowledge: harvesting entities and relationships from web sources
- Deriving a Web-Scale Common Sense Fact Knowledge Base
- Web Crawlers for Searching Hidden Pages: A Survey
- Object Search: Supporting Structured Queries in Web Search Engines
- Extraction of temporal facts and events from Wikipedia
Related papers
- Unsupervised named-entity extraction from the Web: An experimental study
- Information Extraction
- YAGO: A Core of Semantic Knowledge Unifying WordNet and Wikipedia
- Open information extraction from the web
- KnowItNow: Fast, Scalable Information Extraction from the Web
- Declarative Information Extraction Using Datalog with Embedded Extraction Predicates
- Web-scale information extraction in knowitall: (preliminary results)
- Querying text databases for efficient information extraction
- WordNet : an electronic lexical database