Large language models encode clinical knowledge
Explore this paper's citation graph
Summary
MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, is presented and a human evaluation framework for model answers is proposed, suggesting the potential utility of LLMs in medicine.
- Type
- preprint
- Published
- 2022-12-26
- Cited by
- 5,108
- References
- 113
- Access
- Open access
- OpenAlex
- https://openalex.org/W4313197536
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:255124952
Keywords
Computer science, Benchmark (surveying), Harm, Artificial intelligence, Key (lock)
References
- An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
- Health literacy interventions and outcomes: an updated systematic review.
- A Learning Algorithm for Boltzmann Machines
- The Reliability of AHRQ Common Format Harm Scales in Rating Patient Safety Events
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Development of the Patient Education Materials Assessment Tool (PEMAT): a new measure of understandability and actionability for print and audiovisual patient information.
- Bigger is not always better.
- Measuring Harm in Health Care: Optimizing Adverse Event Review
- Scale development: ten main limitations and recommendations to improve future research practices
- Updated systematic review
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Controlling Linguistic Style Aspects in Neural Language Generation
- Datasheets for datasets
- Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer
- Overview of the Medical Question Answering Task at TREC 2017 LiveQA
- emrQA: A Large Corpus for Question Answering on Electronic Medical Records
- Counterfactual Fairness in Text Classification through Robustness
- Model Cards for Model Reporting
- BioBERT: a pre-trained biomedical language representation model for biomedical text mining
- PubMedQA: A Dataset for Biomedical Research Question Answering
Cited by
- ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports
- A real-world test of artificial intelligence infiltration of a university examinations system: A “Turing Test” case study
- Factors influencing Chinese doctors to use medical large language models
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Variational Open-Domain Question Answering
- Putting ChatGPT's Medical Advice to the (Turing) Test
- The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
- Toward General Design Principles for Generative AI Applications 130-144
- ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models
- Do We Still Need Clinical Language Models?
- How Does In-Context Learning Help Prompt Tuning?
- The impending impacts of large language models on medical education
- Almanac: Knowledge-Grounded Language Models for Clinical Medicine
- Ground-Truthing in the European Health Data Space
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human Language
- Does Synthetic Data Generation of LLMs Help Clinical Text Mining?
- Artificial intelligence in oncology: chances and pitfalls
- An overview and a roadmap for artificial intelligence in hematology and oncology
- Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine
- Potential uses of AI for perioperative nursing handoffs: a qualitative study
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Solutions to the Third Benchmark Control Problem
- A Benchmark Characterization of the EEMBC Benchmark Suite
- The Performance Validation of Linear Programming Algorithm Based on Integrated Benchmark
- An empirical assessment of Bellon's clone benchmark
- On the use of a genetic algorithm in High Performance computer benchmark tuning