Assessing promise and limitations of ChatGPT: analysis of translation quality and translation errors of large language models
The rapid advancement of natural language processing (NLP) technologies and the increasing integration of automated tools into translation workflows necessitate a comprehensive exploration of the capabilities and constraints of large language models (LLMs), such as ChatGPT, across various stylistic domains and thematic areas. This study addresses the need for a comparative analysis of machine translation quality produced by ChatGPT, specifically focusing on specialized, literary, and scientific discourses. Central to this research is the identification of errors associated with contextual misinterpretation and the phenomenon of model "hallucinations."
Based on a comparative framework, the research delineates the operational parameters of ChatGPT, including its potential for context recognition, communicative situational awareness, and the resolution of specific translation challenges at the levels of stylistic congruence and lexical equivalence. This research employs a hybrid evaluation framework that integrates a linguistically-grounded metric with the automated BLEU scoring system. The findings reveal a significant correlation between translation quality, prompt precision, and the typological characteristics of the source text. Both linguistic and automated metrics indicate that translations of highly specialized technical content exhibit higher accuracy than those of literary or scientific texts, which require nuanced syntactic construction and terminological selection within specific linguistic conventions. Furthermore, the results highlight inherent risks of generative AI, such as semantic distortions and contextual errors, which can compromise the integrity of translated or edited content. The paper concludes by proposing practical strategies and prompt engineering techniques to enable translators to effectively leverage innovative neural network technologies, particularly the latest iterations of ChatGPT, in professional practice.
Moiseyeva, I.Yu., Relishsky, A. I. (2026). Assessing promise and limitations of ChatGPT: analysis of translation quality and translation errors of large language models, Research Result. Theoretical and Applied Linguistics, 12 (2), 84–110.


















While nobody left any comments to this publication.
You can be first.
Avetesyan, K. I. (2023). Method of detection of inter-language borrowings in texts, Abstract of Th. SC dissertation, Institute of Systems Programming by V. P. Ivannikova of the Russian Academy of Sciences, Moscow, Russia. (In Russian).
Bithymirov, A. R. (2023). Program to convert speech into text as an effective tool of translation, Abstract of Ph. SC dissertation, Military University named after Prince Alexander Nevsky, Moscow, Russia. (In Russian).
Blumfild, L. (1968). Yazyk [Language], Progress, Russia. (In Russian).
Efimova, O. V. (2017). Fundamentalnye osnovy sovremennogo mashinnogo perevoda v svete obshhey teorii perevodcheskoy deyatelnosti [Fundamental foundations of modern machine translation in the light of general theory of translation activity], Cheboksary, Russia, 205–212. (In Russian).
Zaytsev, D. V. (2024). Why big language models do not (always) reason as people?, Vestnik Moskovskogo universiteta,7, Filosofiya, 1, 76–93. (In Russian).
Lomov, P. A. (2025). The use of ontology for contextualization of requests to large language models, Ontologiya proektirovaniya, Volume 15, 2(56), 239-248. DOI: 10.18287/2223-9537-2025-15-2-239-248 (In Russian).
Liu, M., Shao, C., Ce, G. (2024). Automated translation of political discourse: from large language models to multi-agent system MAGIC-PTF, Litera, 11, 28–46. DOI: 10.25136/2409-8698.2024.11.72197 (In Russian).
Lyax, A. P. (2024). Classification and basic algorithms of emmending in the context of large language models, Scientific Notes TOGU, 15, 3, 79–83 (In Russian).
Rozenczvejg, V. Yu. (1974). Opyt lingvisticheskogo opisaniya leksiko-semanticheskix oshibok v rechi na nerodnom yazyke [Experience in the linguistic description of lexical and semantic errors in speech in a foreign language], Without publisher, Moscow, Russia. (In Russian)
Sushhin, M. A. (2024). CHATGPT and other intellectual assistants of the modern scientist, Nauka, Texnologiya i Obshhestvo, 2, 5–20. DOI: 10.31249/scis/2024.02.01 (In Russian).
Tyurina, D. A., Pal'mov, S. V. (2023). Application of neural networks in processing natural language, ZHurnal prikladnyh issledovanij, 7, 158–162. DOI 10.47576/2949-1878_2023_7_158 (In Russian).
Jacobson, R. and Halle, M. (1962). Phonology and its relation to phonetics, Novoe v lingvistike, Volume 2, 231–278. (In Russian).
Al Rousan, Rafat, Jaradat, Raghad S. and Malkawi, M. (2025) ChatGPT translation vs. human translation: an examination of a literary text, Cogent Social Sciences, 11 (1). https://doi.org/10.1080/23311886.2025.2472916(In English).
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu T., Chung, W., Do, Q. V., Xu, Y. and Fung P. (2023). A Multitask, multilingual, multimodal evaluation of CHATGPT on reasoning, hallucination, and interactivity, Arxiv preprint. https://doi.org/10.48550/arXiv.2302.04023(In English).
Bar-Hillel, Yehoshua (1953). Some Linguistic Problems Connected With Machine Translation, Philosophy of Science. 20 (3), 217–225. (In English).
Bentivogli, L., Cettolo, M., Federico, M. and Federmann, C. (2018). Machine Translation Human Evaluation: an investigation of evaluation based on Post-Editing and its relation with Direct Assessment, Proceedings of IWSLT, 62–69. (In English).
Brown, P. F., Della Pietra, S. A., Della Pietra, V. J. and Mercer, R. L. (1993). The Mathematics of Statistical Machine Translation: Parameter Estimation, Computational Linguistics, 19, 263–311.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T. and Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with Gpt-4, Arxiv preprint. https://doi.org/10.48550/arXiv.2303.12712(In English).
Chomsky, N. (1980). Rules and Representations, Columbia University Press, New York, US. (In English).
Dynel, M. (2023). Lessons in linguistics with ChatGPT: Metapragmatics, metacommunication, metadiscourse and metalanguage in human-AI interactions, Language & Communication, 93, 107–124. https://doi.org/10.1016/j.langcom.2023.09.002 (In English).
Jones, D. (1909). The Pronunciation of English, CUP, Cambridge, UK. (In English).
Hendy, Amr., Abdelrehim, M., Sharaf, Amr., Raunak, V., Gabr, M., Matsushita, H., Kim, Y. J., Afify, M. and Awadalla, H. H. (2023). How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation, Arxiv preprint. https://doi.org/10.48550/arXiv.2302.09210(In English).
Hutchins, J. (1999). Retrospect and prospect in computer-based translation, Proceedings of Machine Translation Summit, VII, 30–36. (In English).
Kade, O. (1968). Zufall und Gesetzmässigkeit in der Übersetzung, VEB Verlag Enzyklopädie, Leipzig, Germany (In German).
Lai, V. D., Ngo, N. T., Veyseh, A. P. B., Man, H., Dernoncourt, F., Bui, T. and Nguyen, T. H. (2023). ChatGPT beyond English: Towards a сomprehensive evaluation of large language models in multilingual learning, Arxiv preprint. https://doi.org/10.48550/arXiv.2304.05613 (In English).
Nababan, M., Nuraeni, A. and Sumardiono, S. (2012). Pengembangan Model Penilaian Kualitas Terjemahan, Kajian Linguistik Dan Sastra, 24 (1), 39–57. https://doi.org/10.23917/KLS.V24I1.101 (In English).
Nagao, М. (1989). Machine Translation, Oxford University Press, Oxford, UK. (In English).
Ouyang, F., Wu, M., Zhang, L., Xu, W., Zheng, L. and Cukurova, M. (2023). Making strides towards AI-supported regulation of learning in collaborative knowledge construction, Computers in Human Behavior, 142. https://doi.org/10.1016/j.chb.2023.107650
Peng, K., Ding, L., Zhong, Q., Shen, Li, Liu, X., Zhang, M., Ouyang, Y. and Tao, D. (2023). Towards Making the Most of ChatGPT for Machine Translation, Arxiv preprint. https://arxiv.org/pdf/2303.13780v3(In English).
Siu, Sai Cheong (2023) ChatGPT and GPT-4 for Professional Translators: Exploring the Potential of Large Language Models in Translation, SSRN Electronic Journal. http://dx.doi.org/10.2139/ssrn.4448091 (In English).
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. and Polosukhin, I. (2017), Attention Is All You Need, Arxiv preprint, 30. https://doi.org/10.48550/ARXIV.1706.03762(In English).
Wang, L., Lyu, C., Ji, T., Zhang, Z., Yu, D., Shi, S. and Tu, Z. (2023). Document-level machine translation with large language models, Arxiv preprint, 16646–16661. https://arxiv.org/pdf/2304.02210(In English).
Wu, H., Wang, W., Wan, Y., Jiao, W. and Lyu, M. (2023). ChatGPT or Grammarly? Evaluating ChatGPT on grammatical error correction benchmark, Arxiv preprint, https://arxiv.org/pdf/2303.13648(In English).
Zhou, D., Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, Ed. and Le, Q. (2022). Chain of Thought Prompting Elicits Reasoning in Large Language Models, Arxiv preprint, 3. https://arxiv.org/abs/2201.11903(InEnglish).