BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Explore this paper's citation graph
Summary
This article introduces BioBERT (Bidirectional Encoder Representations from Transformers for Biomedical Text Mining), which is a domain-specific language representation model pre-trained on large-scale biomedical corpora that largely outperforms BERT and previous state-of-the-art models in a variety of biomedical text mining tasks when pre- trained on biomedical Corpora.
- Type
- article
- Published
- 2019-01-25
- Cited by
- 7,868
- References
- 45
- Access
- Open access
- OpenAlex
- https://openalex.org/W2911489562
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:59291975
Keywords
Biomedical text mining, Computer science, Artificial intelligence, Natural language processing, Language model
References
- An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
- An attention‐based BiLSTM‐CRF approach to document‐level chemical named entity recognition
- Chemical–gene relation extraction using recursive neural network
- Introduction to the Bio-entity Recognition Task at JNLPBA
- The EU-ADR corpus: Annotated drugs, diseases, targets, and their relationships
- The SPECIES and ORGANISMS Resources for Fast and Accurate Identification of Taxonomic Names in Text
- LINNAEUS: A species name identification system for biomedical literature
- A neural probabilistic language model
- Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research
- The CHEMDNER corpus of chemicals and drugs and its annotation principles
- Distributed Representations of Words and Phrases and their Compositionality
- Overview of BioCreative II gene mention recognition
- 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text
- NCBI Disease Corpus: A Resource for Disease Name Recognition and Concept Normalization
- GloVe: Global Vectors for Word Representation
- BioCreative V CDR task corpus: a resource for chemical disease relation extraction
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Distributional Semantics Resources for Biomedical Text Processing
- ChimerDB 3.0: an enhanced database for fusion genes from cancer transcriptome and literature data mining
- A transition‐based joint model for disease named entity recognition and normalization
Cited by
- Enhancing Clinical Concept Extraction with Contextual Embedding
- Marie and BERT—A Knowledge Graph Embedding Based Question Answering System for Chemistry
- ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports
- Learning entity-oriented representation for biomedical relation extraction
- Clinical Concept Embeddings Learned from Massive Sources of Multimodal Medical Data
- Natural language processing
- NSEEN: Neural Semantic Embedding for Entity Normalization
- SciBERT: Pretrained Contextualized Embeddings for Scientific Text
- Unsupervised Domain Adaptation of Contextualized Embeddings: A Case Study in Early Modern English
- ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
- MetaPred: Meta-Learning for Clinical Risk Prediction with Limited Patient Electronic Health Records
- A Neural Named Entity Recognition and Multi-Type Normalization Tool for Biomedical Text Mining
- A Silver Standard Corpus of Human Phenotype-Gene Relations
- Towards reliable named entity recognition in the biomedical domain
- Transfer Learning for Causal Sentence Detection
- IITP at MEDIQA 2019: Systems Report for Natural Language Inference, Question Entailment and Question Answering
- Question Answering in the Biomedical Domain
- REflex: Flexible Framework for Relation Extraction in Multiple Domains
- Enhancing PIO Element Detection in Medical Text Using Contextualized Embedding
- Toward a clinical text encoder: pretraining for clinical natural language processing with applications to substance misuse
Related papers
- Wide-coverage relation extraction from MEDLINE using deep syntax
- Extraction of Medication and Temporal Relation from Clinical Text using Neural Language Models
- BioRel: towards large-scale biomedical relation extraction
- Establishing a baseline for literature mining human genetic variants and their relationships to disease cohorts
- Large-scale entity representation learning for biomedical relationship extraction
- Relation extraction from Traditional Chinese Medicine journal publication
- Improving chemical disease relation extraction with rich features and weakly labeled data