Accurate ICD-11 coding is critical for billing, medical studies, and software communications and documentation in a clinical environment. Manual coding is not reliable and consumes a lot of time. There is limited data in the Persian language for training a language model. As such, a language model with an English dataset and a pre-trained model has been used to map sentences in Persian to model language for predicting ICD-11 codes. This reduces the reliance on Persian language datasets and greatly improves prediction accuracy for Persian. In this paper, a technique utilizing pre-trained large language models (LLMs) for ICD-11 automation in terms of patient grievances extracted from the MIMIC-IV corpus is proposed. Prompt-based training is adopted in our model, and it obtains satisfactory performance with an average F1-score of 0.90 for ICD-11 prediction. We also demonstrate the usability of our model through output in JSON format for integration into medical tools.
Tanhaei,M . (2026). Automated ICD-11 Coding with Pre-trained LLM Models: Leveraging Prompt-Based Learning. (e6227). AUT Journal of Mathematics and Computing, (), e6227 doi: 10.22060/ajmc.2025.24401.1417
MLA
Tanhaei,M . "Automated ICD-11 Coding with Pre-trained LLM Models: Leveraging Prompt-Based Learning" .e6227 , AUT Journal of Mathematics and Computing, , , 2026, e6227. doi: 10.22060/ajmc.2025.24401.1417
HARVARD
Tanhaei M. (2026). 'Automated ICD-11 Coding with Pre-trained LLM Models: Leveraging Prompt-Based Learning', AUT Journal of Mathematics and Computing, (), e6227. doi: 10.22060/ajmc.2025.24401.1417
CHICAGO
M Tanhaei, "Automated ICD-11 Coding with Pre-trained LLM Models: Leveraging Prompt-Based Learning," AUT Journal of Mathematics and Computing, (2026): e6227, doi: 10.22060/ajmc.2025.24401.1417
VANCOUVER
Tanhaei M. Automated ICD-11 Coding with Pre-trained LLM Models: Leveraging Prompt-Based Learning. AUT J Math Comput. 2026;():e6227. doi: 10.22060/ajmc.2025.24401.1417