NCBI Disease Corpus: A Resource for Disease Name Recognition and Concept Normalization
Explore this paper's citation graph
Summary
The results show that the NCBI disease corpus has the potential to significantly improve the state-of-the-art in disease name recognition and normalization research, by providing a high-quality gold standard thus enabling the development of machine-learning based approaches for such tasks.
- Type
- article
- Published
- 2014-01-03
- Cited by
- 965
- References
- 50
- Access
- Open access
- OpenAlex
- https://openalex.org/W2169099542
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:234064
Keywords
Computer science, Identifier, Annotation, Natural language processing, Information retrieval
References
- Analysis of biological processes and diseases using text mining approaches.
- Combining Terminology Resources and Statistical Methods for Entity Recognition: an Evaluation
- Technical Brief: Agreement, the F-Measure, and Reliability in Information Retrieval
- Information retrieval and knowledge discovery in biomedical text : papers from the AAAI Fall Symposium
- Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program
- Natural Language Processing - IJCNLP 2005, Second International Joint Conference, Jeju Island, Korea, October 11-13, 2005, Proceedings
- BioInfer: a corpus for information extraction in the biomedical domain
- A Reappraisal of Sentence and Token Splitting for Life Sciences Documents
- PubTator: a web-based text mining tool for assisting biocuration
- Pre-annotating Clinical Notes and Clinical Trial Announcements for Gold Standard Corpus Development: Evaluating the Impact on Annotation Speed and Potential Bias
- Anaphoric reference in clinical reports: Characteristics of an annotated corpus
- Exploring Two Biomedical Text Genres for Disease Recognition
- Getting Started in Text Mining
- Concept annotation in the CRAFT corpus
- Assessment of disease named entity recognition on a corpus of annotated sentences
- GENETAG: a tagged corpus for gene/protein named entity recognition
- The Human Phenotype Ontology
- A corpus of full-text journal articles is a robust evaluation tool for revealing differences in performance of biomedical natural language processing tools
- Overview of BioCreAtIvE task 1B: normalized gene lists
- The gene normalization task in BioCreative III
Cited by
- MetaMap Lite: an evaluation of a new Java implementation of MetaMap
- PKDE4J: Entity and relation extraction for public knowledge discovery
- PALM-IST: Pathway Assembly from Literature Mining - an Information Search Tool
- Crowdsourcing in biomedicine: challenges and opportunities
- Feature Engineering for Drug Name Recognition in Biomedical Texts: Feature Conjunction and Feature Selection
- Identifying named entities from PubMed® for enriching semantic categories
- BioC interoperability track overview
- Development of large-scale TCM corpus using hybrid named entity recognition methods for clinical phenotype detection: An initial study
- Community challenges in biomedical text mining over 10 years: success, failure and the future
- Processing biological literature with customizable Web services supporting interoperable formats
- tmBioC: improving interoperability of text-mining tools with BioC
- BC4GO: a full-text corpus for the BioCreative IV GO task
- Marky: A tool supporting annotation consistency in multi-user and iterative document annotation projects
- Natural language processing pipelines to annotate BioC collections with an application to the NCBI disease corpus
- Supporting the annotation of chronic obstructive pulmonary disease (COPD) phenotypes with text mining workflows
- Cadec: A corpus of adverse drug event annotations
- SimConcept: A Hybrid Approach for Simplifying Composite Named Entities in Biomedical Text
- Challenges in Clinical Natural Language Processing for Automated Disorder Normalization
- SimConcept: A Hybrid Approach for Simplifying Composite Named Entities in Biomedicine
- UTU: Disease Mention Recognition and Normalization with CRFs and Vector Space Representations
Related papers
- A peculiar aspect of patients' safety: the discriminating power of identifiers for record linkage.
- Object Identifiers, Keys, and Surrogates: Object Identifiers Revisited
- Identifiers in Libraries
- Unique Health Identifier for India: An algorithm and feasibility analysis on patient data
- Numbers to Identify Entities (ISADNs–International Standard Authority Data Numbers)
- Which are the best identifiers for record linkage?
- Facilitating linkage through universal patient identifiers: a difficult endeavor.