LLM-Driven Knowledge Extraction from EHRs

Authors

  • Anumandla Mukesh Author

DOI:

https://doi.org/10.5281/de27zp27

Keywords:

Large language models, electronic health records, clinical informatics, natural language processing, named entity recognition, relation extraction, automatic reasoning, automatic summarization, region-based visual question answering, open information extraction.

Abstract

Large Language Models (LLMs) can facilitate structured knowledge extraction from unstructured clinical information sources, such as electronic health records (EHRs), clinical text notes, and clinical trial protocols. This capability bridging heterogeneous data modalities has the potential to power a broad range of downstream clinical decision support and data mining tasks.

Clinical informatics research often relies on specific types of information that are not readily available in explicitly structured form in the corresponding data repositories. These “structured knowledge” types are prevalent in a variety of clinical data sources, such as coded EHRs, but LLMs can offer a different, complementary approach. The task of mapping raw unlabeled text to structured knowledge types involves providing rich supervision during LLM training—using either fine-tuning or prompting mechanisms—to create and support structured knowledge extraction capabilities.

References

1. Li, Y., Rao, S., Solares, J. R. A., Hassaine, A., Ramakrishnan, R., Canoy, D., Zhu, Y., Rahimi, K., & Salimi-Khorshidi, G. (2020). BEHRT: Transformer for electronic health records. Scientific Reports, 10, 7155.

2. Fu, S., Chen, D., He, H., Liu, S., Moon, S., Peterson, K. J., Shen, F., Wang, L., Wang, Y., Wen, A., Zhao, Y., Sohn, S., & Liu, H. (2020). Clinical concept extraction: A methodology review. Journal of Biomedical Informatics, 109, 103526.

3. Mashetty, S., Malempati, M., Paleti, S., Adusupalli, B., & Singireddy, J. (2025). A Multidisciplinary Framework for AI and Data-Driven Transformation in Taxation, Insurance, Mortgage Financing, and Financial Advisory: Integrating Cloud Computing, Deep Learning, and Agentic AI for Community-Centric Economic Development. Insurance, Mortgage Financing, and Financial Advisory: Integrating Cloud Computing, Deep Learning, and Agentic AI for Community-Centric Economic Development.

4. Kraljevic, Z., Searle, T., Shek, A., Roguski, Ł., Noor, K., Bean, D., Mascio, A., Zhu, L., Folarin, A., Roberts, A., Bendayan, R., & Teo, J. (2021). Multi-domain clinical natural language inference. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3229–3241.

5. Rasmy, L., Xiang, Y., Xie, Z., Tao, C., Zhi, D., & others. (2021). Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digital Medicine, 4, 86.

6. Burugulla, J. K. R., Kannan, S., Malempati, M., Challa, H. A. H. S. R., Pandugula, C., & Goma, T. (2025, April). Federated Learning Based Cloud Computing Solutions with Privacy Preservation for Intelligent Utilities. In 2025 4th OPJU International Technology Conference (OTCON) on Smart Computing for Innovation and Advancement in Industry 5.0 (pp. 1-6). IEEE.

7. Mulyar, A., Dernoncourt, F., & Dockhorn, C. (2021). MT-clinical BERT: Scaling clinical information extraction with multitask learning. Journal of the American Medical Informatics Association, 28(10), 2108–2115.

8. Liu, F., Shareghi, E., Meng, Z., Basaldella, M., & Collier, N. (2021). Self-alignment pretraining for biomedical entity representations. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4228–4238.

9. Kummari, D. N., Singireddy, J., Sheelam, G. K., Nandan, B. P., Pandiri, L., Lakkarasu, P., & Dwaraka. (2025, August). Generative AI Models for Process Optimization in Semiconductor Wafer Design and Yield Prediction. In International Conference on Artificial Intelligence: Theory and Applications (pp. 220-233). Cham: Springer Nature Switzerland.

10. Gu, Y., Tinn, D., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., & Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23.

11. Pamisetty, V. (2019). Machine Learning Models for Real-Time Tax Fraud Detection and Risk Assessment in Digital Government Systems. Global Research Development (GRD) ISSN, 2455-5703.

12. Lewis, P., Ott, M., Du, J., & Stoyanov, V. (2020). Pretrained language models for biomedical natural language processing: A review. Briefings in Bioinformatics, 22(1), 1–14.

13. Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240.

14. Adusupalli, B., Malempati, M., Paleti, S., Mashetty, S., & Singireddy, J. (2025). Integrated financial ecosystems: AI-driven innovations in taxation, insurance, mortgage analytics, and community investment through cloud, big data, and advanced data engineering. Journal of Information Systems Engineering and Management, 10, 1103-1117.

15. Peng, Y., Yan, S., & Lu, Z. (2020). Transfer learning in biomedical natural language processing: An evaluation of BERT and ELMo on ten benchmarking datasets. Proceedings of the 18th BioNLP Workshop and Shared Task, 58–65.

16. Alsentzer, E., Murphy, J. R., Boag, W., Weng, W.-H., Jindi, D., Naumann, T., & McDermott, M. B. A. (2020). Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical Natural Language Processing Workshop, 72–78.

17. Suura, S. R., Chava, K., Chakilam, C., Nuka, S. T., Maguluri, K. K., & Goma, T. (2025, April). Blockchain-Based Secure and Scalable Models for Healthcare Network Traffic Monitoring and Optimization. In 2025 4th OPJU International Technology Conference (OTCON) on Smart Computing for Innovation and Advancement in Industry 5.0 (pp. 1-6). IEEE.

18. Si, Y., Wang, J., Xu, H., & Roberts, K. (2021). Enhancing clinical concept extraction with contextualized word representations. Journal of Biomedical Informatics, 117, 103754.

19. Yang, X., Chen, A., PourNejatian, N., Shin, H. C., Smith, K. E., Parisien, C., Compas, C., Martin, C., Flores, M. G., Zhang, Y., Magoc, T., Harle, C. A., Lipori, G., Mitchell, D. A., Hogan, W. R., Shenkman, E. A., Bian, J., & Wu, Y. (2022). GatorTron: A large clinical language model to unlock patient information from unstructured electronic health records. npj Digital Medicine, 5, 140.

20. Chakilam, C., Kannan, S., Recharla, M., Suura, S. R., & Nuka, S. T. (2025). The impact of big data and cloud computing on genetic testing and reproductive health management. American Journal of Psychiatric Rehabilitation, 28(1), 62-72.

21. Lu, Q., Dou, D., & Nguyen, T. (2022). ClinicalT5: A generative language model for clinical text. Findings of the Association for Computational Linguistics: EMNLP 2022, 5436–5443.

22. Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., & Liu, T.-Y. (2022). BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6), bbac409.

23. Sivanand, R., Kumar, D. P., Nagabhyru, K. C., Natarajan, E. P., Pamisetty, V., & Kapila, D. (2025, September). IoT and AI for Real-Time Monitoring in Substation Automation. In 2025 International Conference on Computing and Communications (COMPUTINGCON) (pp. 1-5). IEEE.

24. Yasunaga, M., Leskovec, J., & Liang, P. (2022). LinkBERT: Pretraining language models with document links. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 8003–8016.

25. Gu, Y., Tinn, D., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., & Poon, H. (2022). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23.

26. Yang, X., Wong, C., & others. (2022). A large language model for electronic health records. npj Digital Medicine, 5, 194.

27. Kalisetty, S., Lakkarasu, P., Singireddy, S., Burugulla, J. K. R., Challa, K., & Gadi, A. L. (2025, August). Next-Gen Payment Gateways: Leveraging Federated Learning for Fraud Detection in Cross-Border Transactions. In International Conference on Artificial Intelligence: Theory and Applications (pp. 200-211). Cham: Springer Nature Switzerland.

28. Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29, 1930–1940.

29. Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., ... Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620, 172–180.

30. Paleti, S., Baliyan, M., Aitha, A. R., Reddy, B. A., Bhadauria, G. S., & Sing, S. A. (2025, August). Graph—LSTM Hybrid Model for Improving Fraud Detection Accuracy in E-Commerce Financial Services. In 2025 2nd International Conference on Intelligent Algorithms for Computational Intelligence Systems (IACIS) (pp. 1-6). IEEE.

31. Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Hou, L., Clark, J., Pfohl, S., Cole-Lewis, H., Neal, D., Schaekermann, M., Wang, A., Amin, M., Lachmann, T., McKinney, S. M., Azizi, S., & Natarajan, V. (2023). Towards expert-level medical question answering with large language models. Nature Medicine, 29, 1–8.

32. Nori, H., King, N., McKinney, S. M., Carignan, D., & Horvitz, E. (2023). Capabilities of GPT-4 on medical challenge problems. arXiv preprint arXiv:2303.13375.

33. Yang, X., Chen, A., PourNejatian, N., Shin, H. C., Smith, K. E., Parisien, C., Compas, C., Martin, C., Flores, M. G., Zhang, Y., Magoc, T., Harle, C. A., Lipori, G., Mitchell, D. A., Hogan, W. R., Shenkman, E. A., Bian, J., & Wu, Y. (2023). A large language model for electronic health records: GatorTron and its clinical applications. Journal of the American Medical Informatics Association, 30(4), 1–12.

34. Bolton, E., Li, J., Hall, K., Yoder, N., Riley, A., & others. (2023). Stanford AlpacaMed: Instruction tuning language models for clinical applications. Proceedings of the Clinical NLP Workshop, 1–10.

35. Jeblick, K., Schachtner, B., Dexheimer, J., et al. (2024). ChatGPT makes medicine easy to swallow: An exploratory analysis of AI-generated medical explanations. European Radiology, 34, 1–10.

36. Huang, J., Yang, D. M., Rong, R., Nezafati, K., Treager, C., Chi, Z., Wang, S., Cheng, X., Guo, Y., Klesse, L. J., Xiao, G., Peterson, E. D., Zhan, X., & Xie, Y. (2024). A critical assessment of using ChatGPT for extracting structured data from clinical notes. npj Digital Medicine, 7, 106.

37. Mashetty, S. (2025). Technology-driven analytics in mortgage-backed securities for single-family mortgage financing. Available at SSRN 5236573.

38. Nerella, S., Bandyopadhyay, S., Zhang, J., Contreras, M., Siegel, S., Bumin, A., Silva, B., Sena, J., Shickel, B., Bihorac, A., Khezeli, K., & Rashidi, P. (2024). Transformers and large language models in healthcare: A review. Artificial Intelligence in Medicine, 154, 102900.

39. Meng, X., Yan, X., Zhang, K., Liu, D., Cui, X., Yang, Y., Zhang, M., Cao, C., Wang, J., Wang, X., Gao, J., Wang, Y.-G.-S., Ji, J., Qiu, Z., Li, M., Qian, C., Guo, T., Ma, S., Wang, Z., Guo, Z., Lei, Y., Shao, C., Wang, W., Fan, H., & Tang, Y.-D. (2024). The application of large language models in medicine: A scoping review. iScience, 27(5), 109713.

40. Wang, D., & Zhang, S. (2024). Large language models in medical and healthcare fields: Applications, advances, and challenges. Artificial Intelligence Review, 57, 299.

41. Omiye, J. A., Gui, H., Jasim, M., Chen, J. H., & Daneshjou, R. (2024). Large language models in medicine: The potentials and pitfalls. Annals of Internal Medicine, 177(2), 210–220.

42. Omar, M., Nadkarni, G. N., Klang, E., & Glicksberg, B. S. (2024). Large language models in medicine: A review of current clinical trials across healthcare applications. PLOS Digital Health, 3(11), e0000662.

43. Nori, H., King, N., McKinney, S. M., Carignan, D., & Horvitz, E. (2024). Capabilities of GPT-4 on medical challenge problems. npj Digital Medicine, 7, 1–10.

44. Shah-Mohammadi, F., & Finkelstein, J. (2024). Extraction of substance use information from clinical notes: A generative pretrained transformer-based investigation. JMIR Medical Informatics, 12, e56243.

45. Li, Y., Rao, S., Solares, J. R. A., Hassaine, A., Ramakrishnan, R., Canoy, D., Zhu, Y., Rahimi, K., & Salimi-Khorshidi, G. (2024). Clinical natural language processing for extracting structured information from electronic health records. Journal of Biomedical Informatics, 149, 104575.

46. Wang, Y., Sun, S., & others. (2024). Large language models for clinical information extraction and representation learning. Journal of Biomedical Informatics, 152, 104635.

47. Liu, X., Xu, Y., & others. (2024). Large language models for electronic health record summarization and clinical information extraction. Journal of the American Medical Informatics Association, 31(6), 1–12.

48. Wu, S., Zhang, X., & others. (2024). Clinical information extraction using transformer-based language models: A systematic evaluation. BMC Medical Informatics and Decision Making, 24, 1–15.

49. Kim, M.-S., Chung, P., & Aghaeepour, N. (2025). Information extraction from clinical texts with generative pre-trained transformer models. International Journal of Medical Sciences, 22(5), 1015–1028.

50. Ntinopoulos, V., & others. (2025). Large language models for data extraction from unstructured and semi-structured electronic health records: A multiple model performance evaluation. BMJ Health Care Informatics, 32, e101139.

51. Kim, J., Lee, S., & others. (2025). Scalable information extraction from free text electronic health records using large language models. BMC Medical Research Methodology, 25, 1–15.

52. Mahony Reategui-Rivera, C., & Finkelstein, J. (2025). Evaluation of the performance of a large language model to extract signs and symptoms from clinical notes. Studies in Health Technology and Informatics, 329, 1–5.

53. Fierens, A., Englebert, A., & Jodogne, S. (2025). Translating UMLS concepts to improve medical entity linking in French: A SapBERT-based approach. Studies in Health Technology and Informatics, 329, 1–5.

54. Shah, N. H., Chen, J. H., & others. (2025). Clinical entity augmented retrieval for clinical information extraction. npj Digital Medicine, 8, 45.

55. Lu, Y., Zhang, Y., & others. (2025). Large language models for clinical text processing: A systematic evaluation of information extraction and summarization. Journal of Biomedical Informatics, 162, 104760.

56. Chen, A., Yang, X., & others. (2025). Integrating large language models with human expertise for disease detection in electronic health records. Computers in Biology and Medicine, 186, 109773.

57. Zhang, Y., Liu, X., & others. (2025). Large language models for clinical text understanding and electronic health record analysis: A systematic review. Artificial Intelligence in Medicine, 160, 103024.

58. Li, J., Chen, Y., & others. (2025). Generative artificial intelligence for clinical natural language processing: Applications, evaluation, and challenges. Journal of the American Medical Informatics Association, 32(4), 1–12.

59. Truhn, D., Cuocolo, R., Adams, L. C., & Bressem, K. K. (2025). Current applications and challenges in large language models for patient care: A systematic review. Communications Medicine, 5, 26.

Additional Files

Published

2026-06-17

Data Availability Statement

none

How to Cite

LLM-Driven Knowledge Extraction from EHRs. (2026). American Data Science Journal for Advanced Computations (ADSJAC), 4(02). https://doi.org/10.5281/de27zp27