Evaluasi Perbandingan Model Machine Translation untuk Penerjemahan Dataset Etika Penggunaan AI

Authors

  • Caroline Angelia Setiawan Institut Teknologi Sepuluh Nopember
  • Aris Tjahyanto

DOI:

https://doi.org/10.35960/ikomti.v7i2.2516

Keywords:

bilingual evaluation understudy (BLEU), machine translation, ethical use of artificial intelligence, meteor

Abstract

The development of Large Language Models (LLMs) and Artificial Intelligence (AI)-based technologies has increased the demand for multilingual chatbots for AI ethics education. However, language differences between chatbot training data and user language remain a challenge that can affect the interaction quality.Although machine translation has been widely used to support multilingual chatbots, studies comparing the impact of different translation models on translation quality, particularly in the domain of AI ethics, remain scarce. This study aims to compare and select the best machine translation model in the field of artificial intelligence ethics. The dataset was obtained from UNESCO’s Recommendation on the Ethics of Artificial Intelligence document and generated using a Retrieval-Augmented Generation (RAG) approach based on LLMs. The dataset consisted of 1,000 English-language questions that were later translated into Indonesian using an LLM and manually validated. The Indonesian-language dataset was used as input for back-translation into English using several machine translation methods, namely Google Translate, MarianMT, and M2M-100. The evaluation was conducted using BLEU and METEOR metrics. The results indicate that Google Translate achieved the highest performance, with a BLEU score of 52.2% and a METEOR score of 81.7%, whereas the lowest performance was observed in MarianMT Multi-EN, with a BLEU score of 18.95% and a METEOR score of 56.19%. The findings also indicate that increasing the number of parameters in the M2M-100 model improved the translation quality. This study demonstrates that machine translation has significant potential for supporting multilingual chatbots, particularly in the field of AI ethics.

References

[1] I. Depounti and S. Natale, “Wild dreams and small routines: AI imaginaries and mundanity in the everyday experiences of genAI Replika bot users,” Inf Commun Soc, vol. 29, no. 5, pp. 1760–1778, Dec. 2025, doi: 10.1080/1369118X.2025.2604668.

[2] E. Ruane, A. Birhane, and A. Ventresque, “Conversational AI: Social and Ethical Considerations,” May 2019.

[3] J. Xue, Y.-C. Wang, C. Wei, X. Liu, J. Woo, and C.-C. J. Kuo, “Bias and Fairness in Chatbots: An Overview,” 2023. [Online]. Available: https://arxiv.org/abs/2309.08836

[4] E. Adamopoulou and L. Moussiades, “Chatbots: History, technology, and applications,” Machine Learning with Applications, vol. 2, p. 100006, 2020, doi: https://doi.org/10.1016/j.mlwa.2020.100006.

[5] K. H. Manurung, A. Shofia, and E. P. Nami, “Implementasi Chatbot Berbasis Kecerdasan Buatan untuk Mendukung Proses Rekognisi Pembelajaran Lampau pada Mahasiswa Jalur RPL,” IKOMTI, vol. 7, no. 1, pp. 1–8, 2026, doi: 10.35960/ikomti.v7i1.2178.

[6] F. S. O. Alhefeiti, M. Ezzat, N. A. A. el Azim, and H. A. Hefty, “A Comparative Study of AI-Powered Chatbot for Health Care,” Journal of Computer and Communications, vol. 13, no. 07, pp. 48–66, 2025, doi: 10.4236/jcc.2025.137003.

[7] T. B. Brown et al., “Language Models are Few-Shot Learners,” in NIPS’20: Proceedings of the 34th International Conference on Neural Information Processing Systems, Jul. 2020, pp. 1877–1901. doi: 10.48550/arXiv.2005.14165.

[8] P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury, “The State and Fate of Linguistic Diversity and Inclusion in the NLP World,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds., Online: Association for Computational Linguistics, Jul. 2020, pp. 6282–6293. doi: 10.18653/v1/2020.acl-main.560.

[9] A. Conneau, G. Lample, M. Ranzato, L. Denoyer, and H. Jégou, “Word Translation Without Parallel Data,” CoRR, vol. abs/1710.04087, 2017, [Online]. Available: http://arxiv.org/abs/1710.04087

[10] X. Liang, Z. Gu, Y. Xie, L. Wang, and Z. Tian, “MUSEDA: Multilingual Unsupervised and Supervised Embedding for Domain Adaption,” Knowl Based Syst, vol. 273, p. 110560, 2023, doi: https://doi.org/10.1016/j.knosys.2023.110560.

[11] P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching Word Vectors with Subword Information,” CoRR, vol. abs/1607.04606, 2016, [Online]. Available: http://arxiv.org/abs/1607.04606

[12] X. V. Lin et al., “Few-shot Learning with Multilingual Generative Language Models,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds., Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 9019–9052. doi: 10.18653/v1/2022.emnlp-main.616.

[13] S. Huang, Y. Ding, J. Pan, and Y. Zhang, “Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs,” 2025. [Online]. Available: https://arxiv.org/abs/2509.23657

[14] L. Ranaldi, G. Pucci, and A. Freitas, “Empowering cross-lingual abilities of instruction-tuned large language models by translation-following demonstrations,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Srikumar, Eds., Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 7961–7973. doi: 10.18653/v1/2024.findings-acl.473.

[15] H. Yoo, J. Jin, K. Cho, and A. Oh, “Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models,” 2025. [Online]. Available: https://arxiv.org/abs/2510.05678

[16] Z. Lin et al., “XPersona: Evaluating Multilingual Personalized Chatbot,” in Proceedings of the 3rd Workshop on Natural Language Processing for Conversational AI, A. Papangelis, P. Budzianowski, B. Liu, E. Nouri, A. Rastogi, and Y.-N. Chen, Eds., Online: Association for Computational Linguistics, Nov. 2021, pp. 102–112. doi: 10.18653/v1/2021.nlp4convai-1.10.

[17] D. Anastasiou et al., “A Machine Translation-Powered Chatbot for Public Administration,” in Proceedings of the 23rd Annual Conference of the European Association for Machine Translation, H. Moniz, L. Macken, A. Rufener, L. Barrault, M. R. Costa-jussà, C. Declercq, M. Koponen, E. Kemp, S. Pilos, M. L. Forcada, C. Scarton, J. den Bogaert, J. Daems, A. Tezcan, B. Vanroy, and M. Fonteyne, Eds., Ghent, Belgium: European Association for Machine Translation, Jun. 2022, pp. 329–330. [Online]. Available: https://aclanthology.org/2022.eamt-1.54/

[18] I. Rivera Trigueros, “Machine translation systems and quality assessment: a systematic review,” Lang Resour Eval, vol. 56, pp. 1–27, Jun. 2022, doi: 10.1007/s10579-021-09537-5.

[19] UNESCO, “Recommendation on the ethics of artificial intelligence,” https://www.unesco.org/en/artificial-intelligence/recommendation-ethics.

[20] A. Sagynbayeva, A. Pyo, S.-H. Yoon, and S.-B. Yang, “Evaluating user performance with RAG-based generative AI: A scenario-based experiment on AI-assisted information retrieval,” Comput Human Behav, vol. 180, p. 108952, 2026, doi: https://doi.org/10.1016/j.chb.2026.108952.

[21] M. M. H. Manik and G. Wang, “Gemma 4, Phi-4, and Qwen3: Accuracy-Efficiency Tradeoffs in Dense and MoE Reasoning Language Models,” 2026. [Online]. Available: https://arxiv.org/abs/2604.07035

[22] Y. Wu et al., “Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,” 2016. [Online]. Available: https://arxiv.org/abs/1609.08144

[23] M. Junczys-Dowmunt et al., “Marian: Fast Neural Machine Translation in C++,” 2018. [Online]. Available: https://arxiv.org/abs/1804.00344

[24] A. Fan et al., “Beyond English-Centric Multilingual Machine Translation,” 2020. [Online]. Available: https://arxiv.org/abs/2010.11125

[25] R. Menon, N. Tolani, G. Tolamatti, A. Ahuja, and P. R L, “Textlytic: Automatic Project Report Summarization Using NLP Techniques,” 2022, pp. 119–132. doi: 10.1007/978-981-16-7088-6_10.

[26] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a Method for Automatic Evaluation of Machine Translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, P. Isabelle, E. Charniak, and D. Lin, Eds., Philadelphia, Pennsylvania, USA: Association for Computational Linguistics, Jul. 2002, pp. 311–318. doi: 10.3115/1073083.1073135.

[27] A. Lavie and A. Agarwal, “Meteor: an automatic metric for MT evaluation with high levels of correlation with human judgments,” in Proceedings of the Second Workshop on Statistical Machine Translation, in StatMT ’07. USA: Association for Computational Linguistics, 2007, pp. 228–231.

Published

28-06-2026

How to Cite

[1]
C. A. Setiawan and A. Tjahyanto, “Evaluasi Perbandingan Model Machine Translation untuk Penerjemahan Dataset Etika Penggunaan AI”, IKOMTI, vol. 7, no. 2, Jun. 2026.