Learning to Classify Texts Using Positive and Unlabeled Data
Explore this paper's citation graph
Summary
An effective technique to solve the problem of labeled negative document classification is proposed that combines the Rocchio method and the SVM technique for classifier building and outperforms existing methods significantly.
- Type
- article
- Published
- 2003-08-09
- Cited by
- 531
- References
- 20
- OpenAlex
- https://openalex.org/W22461475
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:59631679
Keywords
Computer science, Classifier (UML), Artificial intelligence, Support vector machine, Class (philosophy)
References
- Exploiting Relations Among Concepts to Acquire Weakly Labeled Training Data
- Active + Semi-supervised Learning = Robust Multi-View Learning
- Making large scale SVM learning practical
- Advances in kernel methods: support vector learning
- Enhancing Supervised Learning with Unlabeled Data
- Partially Supervised Classification of Text Documents
- Introduction to Modern Information Retrieval
- Efficient noise-tolerant learning from statistical queries
- A re-examination of text categorization methods
- PEBL: positive example based learning for Web page classification using SVM
- Combining labeled and unlabeled data with co-training
- Maximum likelihood from incomplete data via the EM - algorithm plus discussions on the paper
- A sequential algorithm for training text classifiers
- Text Classification from Labeled and Unlabeled Documents using EM
- Semi-supervised Clustering by Seeding
- Learning from positive and unlabeled examples
- Estimating the Support of a High-Dimensional Distribution
- The Nature of Statistical Learning Theory
- The effect of adding relevance information in a relevance feedback environment
- The Nature of Statistical Learning Theory
Cited by
- Learning to Identify Unexpected Instances in the Test Set
- Positive Unlabeled Learning for Data Stream Classification
- Personalised ontology learning and mining for web information gathering
- TPN 2 : Using positive-only learning to deal with the heterogeneity of labeled and unlabeled data
- Robust content-based image retrieval of multi-example queries
- Relevance feature discovery for text analysis
- Identification of Consumer Adverse Drug Reaction Messages on Social Media
- Partially Supervised Text Classification with Multi-Level Examples
- Automatic Generation of Background Text to Aid Classification
- Semi-Supervised Text Classification Using Positive and Unlabeled Data
- Learning from positive and unlabeled examples in biology. (Méthodes d'apprentissage statistique à partir d'exemples positifs et indéterminés en biologie)
- Improving Text Classification Using EM with Background Text
- Positive Unlabeled Learning Algorithm for One Class Classification of Social Text Stream with only very few Positive Training Samples
- POLYGLOT-NER: Massive Multilingual Named Entity Recognition
- Analysis of Mass Spectrometry Data for Protein Identification In Complex Biological Mixtures
- Analyse des propriétés stationnaires et des propriétés émergentes dans les flux d'informations changeant au cours du temps. (Analysis of stationary and emerging properties in information flows changing over time)
- A Novel Reliable Negative Method Based on Clustering for Learning from Positive and Unlabeled Examples
- Query-By-Multiple-Examples using Support Vector Machines
- Knowledge discovery using pattern taxonomy model in text mining
- Classifying biomedical citations without labeled training examples
Related papers
- PEBL: Web page classification without negative examples
- Building text classifiers using positive and unlabeled examples
- Learning with Positive and Unlabeled Examples Using Weighted Logistic Regression
- Estimating the Support of a High-Dimensional Distribution
- Analysis of Learning from Positive and Unlabeled Data
- Learning from positive and unlabeled examples
- Learning classifiers from only positive and unlabeled data
- Machine learning in automated text categorization
- PEBL: positive example based learning for Web page classification using SVM
- Partially Supervised Classification of Text Documents