Large language models encode clinical knowledge

Explore this paper's citation graph

Summary

MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, is presented and a human evaluation framework for model answers is proposed, suggesting the potential utility of LLMs in medicine.

Type
preprint
Published
2022-12-26
Cited by
5,108
References
113
Access
Open access

Keywords

Computer science, Benchmark (surveying), Harm, Artificial intelligence, Key (lock)

References

Cited by

Related papers