Using TF-IDF to Determine Word Relevance in Document Queries
Explore this paper's citation graph
Summary
This paper examines the results of applying Term Frequency Inverse Document Frequency to determine what words in a corpus of documents might be more favorable to use in a query and provides evidence that this simple algorithm efficiently categorizes relevant words that can enhance query retrieval.
- Published
- 2003-01-01
- Cited by
- 2,647
- References
- 7
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:14638345
References
- Reexamining tf.idf based information retrieval with Genetic Programming
- Term-Weighting Approaches in Automatic Text Retrieval
- Information retrieval as statistical translation
- Bridging the lexical chasm: statistical approaches to answer-finding
- Using Linear Algebra for Intelligent Information Retrieval
- A Statistical Approach to Machine Translation
Cited by
- WhatTheySaid: Enriching UK Parliament Debates with Semantic Web
- Search Queries in an Information Retrieval System for Arabic-Language Texts
- SPORK: A SUMMARIZATION PIPELINE FOR ONLINE REPOSITORIES OF KNOWLEDGE
- Intelligent Data Engineering and Automated Learning – IDEAL 2020: 21st International Conference, Guimaraes, Portugal, November 4–6, 2020, Proceedings, Part II
- An indexing weight for voice-to-text search
- Finding optimal text queries to cover Subjects in a taxonomy of a news article database
- Big smog meets web science: smog disaster analysis based on social media and device data on the web
- Parallelization of the Levenshtein distance algorithm
- New techniques for Arabic document classification
- Magnifico: A Platform For Expert Mining Using Metadata
- Who Are We Listening to? Detecting User-generated Content (UGC) on the Web
- The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems
- Anaphora resolution for Arabic machine translation : a case study of nafs
- Performance Evaluation of Search Engines Using Enhanced Vector Space Model
- Classification Method for Shared Information on Twitter Without Text Data
- Topic category analysis on twitter via cross-media strategy
- Short message service normalization for communication with a health information system
- Borsa Istanbul (BIST) daily prediction using financial news and balanced feature selection
- Process Improvement of LSA for Semantic Relatedness Computing
- Automated Fair Paper Reviewer Assignment for Conference Management System
Related papers
No related papers recorded.