Skip to Main Content (Press Enter)

Logo UNISS
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Professioni
  • Persone
  • Pubblicazioni
  • Strutture
  • Terza Missione
  • Competenze

Logo UNISS

|

UNIFIND

uniss.it
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Professioni
  • Persone
  • Pubblicazioni
  • Strutture
  • Terza Missione
  • Competenze
  1. Pubblicazioni

Large Language Models and Surgical Decision-Making: Evaluation of Generative Unimodal AI in Facial Traumatology Practice

Articolo
Data di Pubblicazione:
2025
Citazione:
Large Language Models and Surgical Decision-Making: Evaluation of Generative Unimodal AI in Facial Traumatology Practice / Benedetti, S., Frosolini, A., Catarzi, L., Vaira, L.A., Consorti, G., Paglianiti, M., Gennaro, P., Gabriele, G.. - In: JOURNAL OF MAXILLOFACIAL & ORAL SURGERY. - ISSN 0972-8279. - (2025). [10.1007/s12663-025-02556-7]
Abstract:
Introduction: Large language models (LLMs) offer remarkable potential in assisting healthcare professionals with diagnostic and therapeutic decision-making processes. However, the integration of LLMs into healthcare decision-making processes also introduces several doubts in the field of usefulness, reliability and ethical implications. Materials and Methods: To assess the potential of LLMs in managing intricate surgical situations, a cross-sectional study with 30 real-world cases of maxillofacial traumatology was designed. The cases were presented to ChatGPT-4, Google Bard, and maxillofacial surgery residents in a standardized manner, and the results of the subjects were evaluated by an expert surgeon panel of reviewers using the AIPI and QAMAI tools. Results: ChatGPT-4 and Bard showed comparable performances in patient feature consideration but differed significantly in their ability to suggest differential diagnoses. ChatGPT-4 outperformed Bard in proposing additional examinations and treatment plans. Compared to LLMs, human surgery residents consistently scored higher across all parameters of the QAMAI tool, indicating superior accuracy, clarity, relevance, completeness, quality of references, and overall usefulness. Discussion: Both LLMs demonstrated their potential to support clinical decision-making in facial traumatology, but they require further development to be sufficiently reliable for real-world clinical use. Conclusions: AIPI and QAMAI proved their utility as evaluation tools but highlighted the need for standardization in LLM-generated responses assessment.
Tipologia CRIS:
1.1 Articolo in rivista
Keywords:
AI; AIPI; Artificial intelligence; Bard; ChatGPT; Large language models; Maxillofacial; QAMAI; Surgery; Traumatology
Elenco autori:
Benedetti, S.; Frosolini, A.; Catarzi, L.; Vaira, L. A.; Consorti, G.; Paglianiti, M.; Gennaro, P.; Gabriele, G.
Autori di Ateneo:
VAIRA Luigi Angelo
Link alla scheda completa:
https://iris.uniss.it/handle/11388/367373
Pubblicato in:
JOURNAL OF MAXILLOFACIAL & ORAL SURGERY
Journal
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.7.0.0