Determining which biomedical concepts should be included in controlled vocabularies relies on expert clinical judgement, yet expert assessments often differ, making consistent evaluation challenging. This study examines how human evaluations compare with algorithmic approaches when assessing concept utility for incremental ontology expansion. Physician ratings were used to create multiple reference standards through a leave-one-out majority voting strategy, capturing both consensus and variability within expert opinions. We compared five individual physicians, five traditional Machine Learning (ML) models trained on expert-generated labels, and five pretrained large language models (LLMs) with zero-shot prompting against these standards. This framework allowed us to analyze patterns of agreement among human/algorithmic evaluators and quantify how closely algorithmic methods align with human expert judgements. Our findings indicate that for our small data set (making no general claims) ML models were superior to LLMs, which in turn were superior to single human experts relative to a set of experts.
Cite this work
Naren Khatwani, Nghia T. Bui, James Geller, Lijing Wang (2026). Human-AI Agreement in Concept Evaluation: A Benchmarking Framework to Support Incremental Ontology Expansion. Poster. In *2026 AMIA Annual Symposium*, 2026