Enhancing Low-Resource Healthcare Chatbots via Multi-Stage Data Augmentation and IndoBERT Fine-Tuning

chatbot low-resource Data Augmentation Indobert IndoT5

Authors

Vol. 8 No. 3 (2026): August
Medical Informatics
February 10, 2026
August 16, 2026
August 29, 2026

Downloads

Indonesian healthcare question-answering (QA) systems often operate in low-resource settings and must handle substantial linguistic variability in real-world user queries, including paraphrasing, informal expressions, and implicit intent. These challenges are compounded by limited annotated healthcare data and the diverse ways patients express similar medical needs in everyday language, causing QA models to rely heavily on surface-level wording. This study proposes a new multi-stage data augmentation method to improve the robustness of a healthcare chatbot based on IndoBERT Fine-Tuned. The proposed method integrates domain-specific fine-tuning with an augmentation pipeline that introduces paraphrased question variants with IndoT5, normalizes informal language, and incorporates controlled lexical variation through Part-of-Speech (POS) based verb synonym replacement while preserving medical entities, thereby expanding linguistic coverage without requiring additional manual annotation. The augmentation process preserves medical intent while generating diverse surface forms, enabling the model to learn more flexible representations of user queries. Experimental results demonstrate that the proposed multi-stage augmentation substantially improves the robustness of the IndoBERT-based healthcare question answering system. Full augmentation expands the training data to 1,424 question–answer pairs and achieves 67.37 Exact Match (EM) and 85.59 F1. Compared with training without augmentation, this corresponds to relative improvements of approximately 11.8% in EM and 7.0% in F1, indicating more reliable answer span extraction under paraphrased, informal, and lexically varied queries. These gains reflect improved alignment between conversational user input and structured healthcare information. Overall, this work highlights the importance of integrated data augmentation for enhancing low-resource Indonesian healthcare question-answering systems. By exposing the model to broader linguistic variations during training, the proposed approach supports more stable real-world performance while maintaining medical intent consistency and provides insights into remaining challenges for reliable healthcare QA deployment.

How to Cite

Rosyadi, A. W., Rohman, T., Fajar, M. R., Zaman, M. Q., & Ma'shumah, S. (2026). Enhancing Low-Resource Healthcare Chatbots via Multi-Stage Data Augmentation and IndoBERT Fine-Tuning. Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics, 8(3), 380-390. https://doi.org/10.35882/ijeeemi.v8i3.327

Similar Articles

121-130 of 165

You may also start an advanced similarity search for this article.