GENETAG: a tagged corpus for gene/protein named entity recognition
Explore this paper's citation graph
Summary
The annotation of GENETAG required intricate manual judgments by annotators which hindered tagging consistency, and the data were pre-segmented into words, to provide indices supporting comparison of system responses to the "gold standard", however, character- based indices would have been more robust than word-based indices.
- Type
- article
- Published
- 2005-05-24
- Cited by
- 270
- References
- 10
- Access
- Open access
- OpenAlex
- https://openalex.org/W2048140075
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:18074692
Keywords
DNA microarray, Computational biology, Gene, Named-entity recognition, Natural language processing
References
- Elements of Machine Learning
- Building a Large Annotated Corpus of English: The Penn Treebank
- Seventh Message Understanding Conference (MUC-7)
- BioCreAtIvE Task 1A: gene mention finding evaluation
- Disambiguating proteins, genes, and RNA in text: a machine learning approach
- Tagging gene and protein names in biomedical text
- GENIA corpus - a semantically annotated corpus for bio-textmining
- Boosting naïve Bayesian learning on a large subset of MEDLINE
- A critical assessment of text mining methods in molecular biology. Proceedings of a workshop. March 28-31, 2004. Granada, Spain.
Cited by
- Application of information extraction techniques to pharmacological domain : extracting drug-drug interactions
- Developing systems for gene normalisation
- SemCat: Semantically Categorized Entities for Genomics
- SimSem: Fast Approximate String Matching in Relation to Semantic Category Disambiguation
- Pathway Curation Support as an Information Extraction Task
- Biomedical Event Extraction with Machine Learning
- Generación automática de resúmenes con apoyo en ontologías aplicada al dominio Biomédico
- Online assessment of protein interaction information extraction systems
- Scaling up Biomedical Event Extraction to the Entire PubMed
- Text and Network Mining for Literature-Based Scientific Discovery in Biomedicine
- Combining Resources to Find Answers to Biomedical Questions
- What’s in a Name? Entity Type Variation across Two Biomedical Subdomains
- Information extraction from pharmaceutical literature
- Contribution à la construction d’ontologies et à la recherche d’information : application au domaine médical
- Assessment of software testing and quality assurance in natural language processing applications and a linguistically inspired approach to improving it
- PKDE4J: Entity and relation extraction for public knowledge discovery
- GoPubMed: ontology based literature search for the life sciences
- The development of a schema for semantic annotation: Gain brought by a formal ontological method
- Clinical Information Extraction: Lowering the Barrier
- BioTagger-GM: a gene/protein name recognition system.
Related papers
- Studying the impact of various features on the performance of Conditional Random Field-based Arabic Named Entity Recognition
- Named Entity Recognition and Classification for Punjabi Shahmukhi
- A Comparative Study of Dictionary-based and Machine Learning-based Named Entity Recognition in Pashto
- Named Entity Recognition on Morphologically Rich Language: Exploring the Performance of BERT with varying Training Levels
- Named Entity Recognition Using BERT Model for Kannada Language
- Named Entity Recognition for Hindi-English Code-Mixed Social Media Text
- Named Entity Recognition of Kumauni Language using Machine Learning (ML)
- Conditional Random Fields based Named Entity Recognition for Sinhala
- Thai Named Entity Recognition Using Bi-LSTM-CRF with Word and Character Representation