| Naslov: | Extracting biomedical entities from clinical records : a comparative study of large language model approaches |
|---|
| Avtorji: | ID Calcina, Erik, Institut "Jožef Stefan" (Avtor) ID Novak, Erik, Institut "Jožef Stefan" (Avtor) ID Mladenić, Dunja, Institut "Jožef Stefan" (Avtor) ID Burger, Helena (Sodelavec pri raziskavi) ID Kuret, Zala (Sodelavec pri raziskavi) ID Matjačić, Zlatko (Sodelavec pri raziskavi) ID Vidmar, Gaj (Sodelavec pri raziskavi) |
| Datoteke: | URL - Izvorni URL, za dostop obiščite https://www.tandfonline.com/doi/full/10.1080/08839514.2026.2700905
PDF - Predstavitvena datoteka, prenos (2,91 MB) MD5: A8D2366096FF94F43F8953035F39A805
|
|---|
| Jezik: | Angleški jezik |
|---|
| Tipologija: | 1.01 - Izvirni znanstveni članek |
|---|
| Organizacija: | IJS - Institut Jožef Stefan
|
|---|
| Povzetek: | Medical institutions produce large volumes of unstructured medical data that require the extraction of relevant medical information used for performing clinical studies and statistical analysis. This paper examines the application of large language models (LLMs) for named entity recognition (NER) in the medical domain, with a focus on their practical usefulness in clinical settings. We evaluate both prompt-based and fine-tuned approaches using the MACCROBAT2020 dataset, which includes clinical case reports annotated with biomedical entities. Furthermore, we extend our evaluation of the fine-tuning methodology to three biomedical NER datasets: QUAERO, NCBI, and E3C. The study compares the performance of several open-source LLMs against baseline models, using exact and relaxed F1 scores across multiple entity types. Fine-tuned LLMs achieved higher strict-match accuracy and produced more reliable structured outputs than prompt-based methods, GLiNER variants, and supervised BERT baselines, particularly for complex medical entities. However, BERT-based encoders remained substantially faster and competitive under relaxed matching. Their performance across multiple datasets, although uneven across languages and annotation schemas, together with efficient operation when quantized, indicates potential for clinical pilot studies. Our code is publicly available on GitHub (https://github.com/erikcalcina/llm-medical-ner) under the MIT license. |
|---|
| Ključne besede: | medical data, named entity recognition, information extraction |
|---|
| Status publikacije: | Objavljeno |
|---|
| Verzija publikacije: | Objavljena publikacija |
|---|
| Datum objave: | 19.07.2026 |
|---|
| Založnik: | Taylor & Francis |
|---|
| Leto izida: | 2026 |
|---|
| Št. strani: | str. [1-23] |
|---|
| Številčenje: | Vol. 40, iss. 1, article no. 2700905 |
|---|
| PID: | 20.500.12556/DiRROS-32097  |
|---|
| UDK: | 004.6:004.8 |
|---|
| ISSN pri članku: | 1087-6545 |
|---|
| DOI: | 10.1080/08839514.2026.2700905  |
|---|
| COBISS.SI-ID: | 287379459  |
|---|
| Avtorske pravice: | © 2026 The Author(s). |
|---|
| Opomba: | Nasl. z nasl. zaslona;
Sodelavci pri raziskavi iz Slovenije: Helena Burger, Zala Kuret, Zlatko Matjačić, Gaj Vidmar;
Opis vira z dne 11. 8. 2026;
|
|---|
| Datum objave v DiRROS: | 27.08.2026 |
|---|
| Število ogledov: | 177 |
|---|
| Število prenosov: | 122 |
|---|
| Metapodatki: |  |
|---|
|
:
|
Kopiraj citat |
|---|
| | | | Objavi na: |  |
|---|
Postavite miškin kazalec na naslov za izpis povzetka. Klik na naslov izpiše
podrobnosti ali sproži prenos. |