Explainable XGBoost for Indonesian Hoax Detection under the Electronic Transactions Law
Keywords:
Indonesian hoax detection, IndoBERT, XGBoost, Electronic Information and Transaction Law, source-bias auditAbstract
The rapid circulation of misleading information in digital spaces creates a need for screening tools that are accurate, transparent, and suitable for human review. This study develops a reproducible Indonesian hoax-detection pipeline and examines whether model explanations can support cautious legal review. The experiment uses a political-hoax text corpus with fixed training, validation, and test splits. The primary classifier combines word- and character-level TF-IDF features with XGBoost, while a frozen-encoder IndoBERT-Lite model is evaluated as a CPU pilot. TreeSHAP summarizes global and local feature contributions, and LIME is used to inspect borderline predictions. On the held-out test set, XGBoost achieved 0.9725 accuracy, 0.9724 macro-F1, 0.9957 ROC-AUC, and a 0.0235 Brier score; the constrained IndoBERT pilot reached a validation macro-F1 of 0.3343 and is not treated as a final benchmark. The most influential features included source and article-genre markers such as “baca juga,” “referensi,” “Kompas,” “Facebook,” and “foto hoaks,” indicating that the model may learn publisher or writing-style shortcuts in addition to claim-related signals. The audit also identified normalized duplicate overlap across the training-validation and training-test splits. The resulting system should therefore support triage, explanation, and documentation by trained reviewers, not serve as a standalone basis for determining the truth of a claim or establishing an Electronic Information and Transactions Law violation.
References
[1] M. Fathin, Y. Sibaroni, and S. Prasetyowati, “Handling Imbalance Dataset on Hoax Indonesian Political News Classification using IndoBERT and Random Sampling,” Jurnal Media Informatika Budidarma, 8(1), 2024. Doi: https://doi.org/10.30865/mib.v8i1.7099.
[2] L. A. Pekandi, R. G. Widjaja, A. Ananta, J. Harefa, and K. Jingga, “Evaluating IndoBERT for Indonesian Hoax News Detection: A Comparative Study with Ensemble and CNN-LSTM Models,” Procedia Computer Science, vol. 269, pp. 1625–1633, 2025. Doi: https://doi.org/10.1016/j.procs.2025.09.105.
[3] R. Kozik, M. M. Ficco, A. Pawlicka, M. Pawlicki, F. Palmieri, and M. Choraś, “When Explainability Turns into a Threat—Using xAI to Fool a Fake News Detection Method,” Computers & Security, vol. 137, 103599, 2024. Doi: https://doi.org/10.1016/j.cose.2023.103599.
[4] V. Prisscilya and A. S. Girsang, ‘Classification of Indonesia False News Detection Using BERTopic and IndoBERT,’ Jurnal Indonesia Sosial Teknologi, vol. 5, no. 8, pp. 3913–3931, 2024. Doi: https://doi.org/10.59141/jist.v5i8.1310.
[5] S. M. T. Situmeang, D. Pudjiastuti, and Sutarjo, “Adequacy of Indonesian Cybercriminal Law Against Generative AI-Based Social Engineering Attacks,” Res Nullius Law Journal, vol. 6, no. 1, pp. 82–97, 2024. Doi: https://doi.org/10.34010/rnlj.v6i1.19927.
[6] N. P. R. Adiati, D. F. Priambodo, Girinoto, S. Indarjani, A. Rizal, A. Prayoga, and Y. Beatrix, “Comparative Study of Predictive Models for Hoax and Disinformation Detection in Indonesian News,” International Journal of Advances in Intelligent Informatics, vol. 10, no. 3, pp. 504–516, 2024. Doi: https://doi.org/10.26555/ijain.v10i3.878.
[7] V. Prisscilya and A. S. Girsang, “Classification of Indonesia False News Detection Using BERTopic and IndoBERT,” Jurnal Indonesia Sosial Teknologi, vol. 5, no. 8, pp. 3913–3931, 2024. Doi: https://doi.org/10.59141/jist.v5i8.1310.
[8] J. Chen, ‘Identification and Analysis of Real and Fake News by XGBoost Algorithm of Machine Learning,’ Applied and Computational Engineering, vol. 40, no. 1, pp. 255–262, 2024. Doi: https://doi.org/10.54254/2755-2721/40/20230661.
[9] P. A. Bhagaskara S., S. S. Prasetiyowati, and Y. Sibaroni, “Hoax Detection of Indonesian News Media on Twitter Using IndoBERT with Word Embedding Word2Vec,” Jurnal Media Informatika Budidarma, vol. 7, no. 3, 2023. Doi: https://doi.org/10.30865/mib.v7i3.6367.
[10] C. J. L. Tobing, I. G. N. Lanang Wijayakusuma, and L. P. I. Harini, “Perbandingan Kinerja IndoBERT dan MBERT untuk Deteksi Berita Hoaks Politik dalam Bahasa Indonesia,” Jurnal Sains dan Teknologi, vol. 14, no. 1, pp. 114–123, 2025. Doi: https://doi.org/10.23887/jstundiksha.v14i1.92126.
[11] M. Ahmed et al., “Explainable Text Classification Model for COVID-19 Fake News Detection,” JISIS, 12(2), 2022. Doi: https://doi.org/10.22667/JISIS.2022.05.31.051.
[12] N. P. R. Adiati, D. F. Priambodo, Girinoto, S. Indarjani, A. Rizal, A. Prayoga, and Y. Beatrix, “Comparative Study of Predictive Models for Hoax and Disinformation Detection in Indonesian News,” International Journal of Advances in Intelligent Informatics, vol. 10, no. 3, pp. 504–516, 2024. Doi: https://doi.org/10.26555/ijain.v10i3.878.
[13] S.-Y. Chien, C.-J. Yang, and F. Yu, “XFlag: Explainable Fake News Detection Model on Social Media,” International Journal of Human–Computer Interaction, vol. 38, no. 18–20, pp. 1808–1827, 2022. Doi: https://doi.org/10.1080/10447318.2022.2062113.
[14] J. Chen, “Identification and Analysis of Real and Fake News by XGBoost Algorithm of Machine Learning,” Applied and Computational Engineering, vol. 40, no. 1, pp. 255–262, 2024. Doi: https://doi.org/10.54254/2755-2721/40/20230661.
[15] C. Zhang et al., “A computational approach for real-time detection of fake news,” Expert Systems with Applications, 221, 119656, 2023. Doi: https://doi.org/10.1016/j.eswa.2023.119656.
[16] A. S. Dwiandari and R. Arifin, “Criminal Law Enforcement on Digital Identity Misuse in AI Era for Commercial Interests in Indonesia,” Indonesian Journal of International Clinical Legal Education, vol. 7, no. 1, 2025. Doi: https://doi.org/10.15294/iccle.v7i1.25525.
[17] S.-Y. Chien, C.-J. Yang, and F. Yu, ‘XFlag: Explainable Fake News Detection Model on Social Media,’ International Journal of Human–Computer Interaction, vol. 38, no. 18–20, pp. 1808–1827, 2022. Doi: https://doi.org/10.1080/10447318.2022.2062113.
[18] N. N. Hidayati, A. Santosa, and S. Shaleha, “Indonesia Political Hoax Dataset,” Mendeley Data, v1, 2025. Doi: https://doi.org/10.17632/mzx5rd397v.1.
[19] E. Mutmainnah, “The Phenomenon of Artificial Intelligence Hallucination,” Jurnal Rechtsvinding, 13(2), 2024. Doi: https://doi.org/10.33331/rechtsvinding.v13i2.1781.
[20] S. N. Dachlan, D. E. S. Karauwan, and N. Lahangatubun, ‘The Role of Artificial Intelligence in Law Enforcement: Towards a More Accurate and Efficient Justice System,’ Sinergi International Journal of Law, vol. 2, no. 3, pp. 198–207, 2024. Doi: https://doi.org/10.61194/law.v2i3.157.
[21] A. S. Dwiandari and R. Arifin, ‘Criminal Law Enforcement on Digital Identity Misuse in AI Era for Commercial Interests in Indonesia,’ Indonesian Journal of International Clinical Legal Education, vol. 7, no. 1, 2025. Doi: https://doi.org/10.15294/iccle.v7i1.25525.
[22] S. M. T. Situmeang, D. Pudjiastuti, and Sutarjo, ‘Adequacy of Indonesian Cybercriminal Law Against Generative AI-Based Social Engineering Attacks,’ Res Nullius Law Journal, vol. 6, no. 1, pp. 82–97, 2024. Doi: https://doi.org/10.34010/rnlj.v6i1.19927.
[23] N. P. R. Adiati, D. F. Priambodo, Girinoto, S. Indarjani, A. Rizal, A. Prayoga, and Y. Beatrix, ‘Comparative Study of Predictive Models for Hoax and Disinformation Detection in Indonesian News,’ International Journal of Advances in Intelligent Informatics, vol. 10, no. 3, pp. 504–516, 2024. Doi: https://doi.org/10.26555/ijain.v10i3.878.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Journal of Artificial Intelligence and Legal Technology

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.









