Тип публикации: статья из журнала
Год издания: 2024
Идентификатор DOI: 10.3390/make6020064
Ключевые слова: machine learning, knowledge extraction
Аннотация: <jats:p>This study presents an integrated approach for automatically extracting and structuring information from medical reports, captured as scanned documents or photographs, through a combination of image recognition and natural language processing (NLP) techniques like named entity recognition (NER). The primary aim was to develПоказать полностьюop an adaptive model for efficient text extraction from medical report images. This involved utilizing a genetic algorithm (GA) to fine-tune optical character recognition (OCR) hyperparameters, ensuring maximal text extraction length, followed by NER processing to categorize the extracted information into required entities, adjusting parameters if entities were not correctly extracted based on manual annotations. Despite the diverse formats of medical report images in the dataset, all in Russian, this serves as a conceptual example of information extraction (IE) that can be easily extended to other languages.</jats:p>
Журнал: Machine Learning and Knowledge Extraction
Выпуск журнала: Т. 6, № 2
Номера страниц: 1361-1377
ISSN журнала: 25044990
Издатель: MDPI