Keyword extraction from a single document using word co-occurrence statistical information
Explore this paper's citation graph
Summary
A new keyword extraction algorithm that applies to a single document without using a corpus and shows comparable performance to tfidf without using an corpus is presented.
- Type
- article
- Published
- 2004-03-01
- Cited by
- 886
- References
- 22
- Access
- Open access
- OpenAlex
- https://openalex.org/W64540378
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:2201076
Keywords
Keyword extraction, tf–idf, Term (time), Computer science, Word (group theory)
References
- A Study Using n-gram Features for Text Categorization
- Automatic text processing
- Statistical Models for Co-occurrence Data
- Similarity-Based Models of Word Cooccurrence Probabilities
- An Automatic Method of the Extraction of Important Words from Japanese Scientific Documents
- KeyGraph: automatic indexing by co-occurrence graph based on building construction metaphor
- A Statistical Approach to Mechanized Encoding and Searching of Literary Information
- A performance evaluation of similarity measures, document term weighting schemes and representations in a Boolean environment
- Computing Machinery and Intelligence
- Extraction of Lexical Translations from Non-Aligned Corpora
- An algorithm for suffix stripping
- METHODS OF AUTOMATIC TERM RECOGNITION : A REVIEW
- A statistical interpretation of term specificity and its application in retrieval
- Word Association Norms, Mutual Information, and Lexicography
- A Cooccurrence-Based Thesaurus and Two Applications to Information Retrieval
- Text Mining, knowledge extraction from unstructured textual data
- Word association norms
- Accurate Methods for the Statistics of Surprise and Coincidence
- Distributional clustering of English words
- Accurate methods for the statistics of surprise and coincidence
Cited by
- Evaluating the Use of Project Glossaries in Automated Trace Retrieval
- Narrative and Hypertext 2011 Proceedings: a workshop at ACM Hypertext 2011, Eindhoven
- Multi-view Clustering Trees for Search and Retrieval in Customer Product Forums
- Automated illustration of multimedia stories
- An Extensive Comparison of Metrics for Automatic Extraction of Key Terms
- Exploration of relationships from texts using self-organizing maps
- Mood-ex-Machina: Towards Automation of Moody Tunes
- A Language-Independent Approach to Keyphrase Extraction and Evaluation
- Social Media and Emergent Organizational Narratives
- Robust Estimation of Google Counts for Social Network Extraction
- Identifying the gist of conversational text: automatic keyword extraction and summarization
- Hashing-basierte Indizierung: Anwendungsszenarien, Theorie und Methoden
- Mining Large-scale Comparable Corpora from Chinese-English News Collections
- Visual document analysis: towards a semantic analysis of large document collections
- SPORK: A SUMMARIZATION PIPELINE FOR ONLINE REPOSITORIES OF KNOWLEDGE
- Traitor: Associating Concepts using the World Wide Web
- DETECTING NETFLIX SERVICE OUTAGES THROUGH ANALYSIS OF TWITTER POSTS
- Helmholtz Principle-Based Keyword Extraction
- A comparison of automated keyphrase extraction techniquesand of automatic evaluation vs. human evaluation
- Building Chemical Information Systems - the ViFaChem II Project