Enhancing Low-Resource Healthcare Chatbots via Multi-Stage Data Augmentation and IndoBERT Fine-Tuning
Downloads
Indonesian healthcare question-answering (QA) systems often operate in low-resource settings and must handle substantial linguistic variability in real-world user queries, including paraphrasing, informal expressions, and implicit intent. These challenges are compounded by limited annotated healthcare data and the diverse ways patients express similar medical needs in everyday language, causing QA models to rely heavily on surface-level wording. This study proposes a new multi-stage data augmentation method to improve the robustness of a healthcare chatbot based on IndoBERT Fine-Tuned. The proposed method integrates domain-specific fine-tuning with an augmentation pipeline that introduces paraphrased question variants with IndoT5, normalizes informal language, and incorporates controlled lexical variation through Part-of-Speech (POS) based verb synonym replacement while preserving medical entities, thereby expanding linguistic coverage without requiring additional manual annotation. The augmentation process preserves medical intent while generating diverse surface forms, enabling the model to learn more flexible representations of user queries. Experimental results demonstrate that the proposed multi-stage augmentation substantially improves the robustness of the IndoBERT-based healthcare question answering system. Full augmentation expands the training data to 1,424 question–answer pairs and achieves 67.37 Exact Match (EM) and 85.59 F1. Compared with training without augmentation, this corresponds to relative improvements of approximately 11.8% in EM and 7.0% in F1, indicating more reliable answer span extraction under paraphrased, informal, and lexically varied queries. These gains reflect improved alignment between conversational user input and structured healthcare information. Overall, this work highlights the importance of integrated data augmentation for enhancing low-resource Indonesian healthcare question-answering systems. By exposing the model to broader linguistic variations during training, the proposed approach supports more stable real-world performance while maintaining medical intent consistency and provides insights into remaining challenges for reliable healthcare QA deployment.
Copyright (c) 2026 Ahmad Wahyu Rosyadi, Taufiqur Rohman, Moh. Rizki Fajar, Muhammad Qomaruz Zaman, Siti Ma'shumah (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlikel 4.0 International (CC BY-SA 4.0) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).






