<?xml version='1.0' encoding='utf-8'?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<article article-type="research-article" dtd-version="1.2" xml:lang="ru" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><front><journal-meta><journal-id journal-id-type="issn">2313-8912</journal-id><journal-title-group><journal-title>Research Result. Theoretical and Applied Linguistics</journal-title></journal-title-group><issn pub-type="epub">2313-8912</issn></journal-meta><article-meta><article-id pub-id-type="doi">10.18413/2313-8912-2026-12-2-0-3</article-id><article-id pub-id-type="publisher-id">4247</article-id><article-categories><subj-group subj-group-type="heading"><subject>COMPARATIVE LINGUISTICS</subject></subj-group></article-categories><title-group><article-title>&lt;strong&gt;Semantic core identification as a method to overcome textoidness&lt;/strong&gt;</article-title><trans-title-group xml:lang="en"><trans-title>&lt;strong&gt;Semantic core identification as a method to overcome textoidness&lt;/strong&gt;</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Kovalchuk</surname><given-names>Aleksandr V.</given-names></name><name xml:lang="en"><surname>Kovalchuk</surname><given-names>Aleksandr V.</given-names></name></name-alternatives><email>kovalchuk_a_v@staff.sechenov.ru</email><xref ref-type="aff" rid="aff1" /></contrib></contrib-group><aff id="aff1"><institution>I.M. Sechenov First Moscow State Medical University, Moscow, Russia</institution></aff><pub-date pub-type="epub"><year>2026</year></pub-date><volume>12</volume><issue>2</issue><fpage>0</fpage><lpage>0</lpage><self-uri content-type="pdf" xlink:href="/media/linguistics/2026/2/Лингвистика_12_2_2026-62-84.pdf" /><abstract xml:lang="ru"><p>We believe that neural machine translation results intended to function as a text always have enough potential for a semantic core (i.e. a communicative center with text-forming properties) to be found and verbalized.

The relevance of this article is provided by two factors. On the one hand, machine translation software is widespread, easily available, and in active use; on the other hand, machine translation results have to be post-edited to the quality of a communicative text due to systematic disruption of its intra-textual connections in the machine translation results which turns out to be, in its raw, non-edited version, a set of separate sentences, in other words &amp;ndash; a &amp;lsquo;textoid&amp;rsquo; that should be fixed by an editor to function as a coherent text. Although frequent cross-checking between the original text and its translation helps eliminate occasional semantic errors and inaccuracies, the AI output in general still looks like a poor-quality text with a &amp;lsquo;machine DNA.&amp;rsquo; This brings us to the core problem: now, there is no reliable method to assess and achieve global semantic coherence in AI-generated translations.

That is why our study aims to lay the foundations of a linguistic method for overcoming textoid-quality of machine translation results by means of semantic core identification. Through a comprehensive approach that comprises such methods as abstraction, analysis, classification, synthesis, modeling, and measurement this study has achieved the following results: (a)&amp;nbsp;a unique tool for semantic core identification was proposed relying on such well-known linguistic concepts as subject, predicate, and object, as well as on a basic subject-logical typology of semantic relations; (b)&amp;nbsp;a need to adjust the initial core wording/formula was demonstrated in 46&amp;nbsp;% of cases; (c)&amp;nbsp;the median core volume (31&amp;nbsp;%) in a textoid was determined for medical news; (d)&amp;nbsp;basic principles of linguistic annotation (how to label specific linguistic, structural, or semantic features) were proposed as well as a system of notations; (e)&amp;nbsp;a principle for representing the semantic core by means of graphic formulae was proposed for illustrative purposes; (f)&amp;nbsp;ways for further scientific research were outlined.

Conclusion: 52 textoids were analyzed to demonstrate applicability of our method, intended to serve as a reliable linguistic tool for identifying a semantic core which, in its turn, can function as (1) a text-forming essence that can be used in converting a textoid into a text; (2) a subject-logical benchmark for controlling and verifying translation, both for specific segments of the machine translation and for the text as a whole; and (3) a tool for interpreting unclear or contradictory passages within the textoid (without direct need to check up with the source text).</p></abstract><trans-abstract xml:lang="en"><p>We believe that neural machine translation results intended to function as a text always have enough potential for a semantic core (i.e. a communicative center with text-forming properties) to be found and verbalized.

The relevance of this article is provided by two factors. On the one hand, machine translation software is widespread, easily available, and in active use; on the other hand, machine translation results have to be post-edited to the quality of a communicative text due to systematic disruption of its intra-textual connections in the machine translation results which turns out to be, in its raw, non-edited version, a set of separate sentences, in other words &amp;ndash; a &amp;lsquo;textoid&amp;rsquo; that should be fixed by an editor to function as a coherent text. Although frequent cross-checking between the original text and its translation helps eliminate occasional semantic errors and inaccuracies, the AI output in general still looks like a poor-quality text with a &amp;lsquo;machine DNA.&amp;rsquo; This brings us to the core problem: now, there is no reliable method to assess and achieve global semantic coherence in AI-generated translations.

That is why our study aims to lay the foundations of a linguistic method for overcoming textoid-quality of machine translation results by means of semantic core identification. Through a comprehensive approach that comprises such methods as abstraction, analysis, classification, synthesis, modeling, and measurement this study has achieved the following results: (a)&amp;nbsp;a unique tool for semantic core identification was proposed relying on such well-known linguistic concepts as subject, predicate, and object, as well as on a basic subject-logical typology of semantic relations; (b)&amp;nbsp;a need to adjust the initial core wording/formula was demonstrated in 46&amp;nbsp;% of cases; (c)&amp;nbsp;the median core volume (31&amp;nbsp;%) in a textoid was determined for medical news; (d)&amp;nbsp;basic principles of linguistic annotation (how to label specific linguistic, structural, or semantic features) were proposed as well as a system of notations; (e)&amp;nbsp;a principle for representing the semantic core by means of graphic formulae was proposed for illustrative purposes; (f)&amp;nbsp;ways for further scientific research were outlined.

Conclusion: 52 textoids were analyzed to demonstrate applicability of our method, intended to serve as a reliable linguistic tool for identifying a semantic core which, in its turn, can function as (1) a text-forming essence that can be used in converting a textoid into a text; (2) a subject-logical benchmark for controlling and verifying translation, both for specific segments of the machine translation and for the text as a whole; and (3) a tool for interpreting unclear or contradictory passages within the textoid (without direct need to check up with the source text).</p></trans-abstract><kwd-group xml:lang="ru"><kwd>Semantic core</kwd><kwd>From textoid to text</kwd><kwd>Cohesion</kwd><kwd>Neural machine translation</kwd><kwd>Isotopy</kwd></kwd-group><kwd-group xml:lang="en"><kwd>Semantic core</kwd><kwd>From textoid to text</kwd><kwd>Cohesion</kwd><kwd>Neural machine translation</kwd><kwd>Isotopy</kwd></kwd-group></article-meta></front><back><ref-list><title>Список литературы</title><ref id="B1"><mixed-citation>Belyayeva,&amp;nbsp;L.&amp;nbsp;N. (2022). Machine translation in modern technology of the translation process, Izvestiya Rossiyskogo gosudarstvennogo pedagogicheskogo universiteta im. A.&amp;nbsp;I.&amp;nbsp;Gertsena, 203, 22, available at: https://doi.org/10.33910/1992-6464-2022-203-22-30 (Accessed 02 September 2025).</mixed-citation></ref><ref id="B2"><mixed-citation>Boronin,&amp;nbsp;A.&amp;nbsp;A. (2016). On the issue of textoids, Vestnik Moskovskogo gosudarstvennogo oblastnogo universiteta. Seriya: Lingvistika, 2, 27, available at: https://doi.org/10.18384/2310-712X-2016-2-26-32 (Accessed 02 September 2025).</mixed-citation></ref><ref id="B3"><mixed-citation>Vishnyakova,&amp;nbsp;A.&amp;nbsp;I. (2017). The Idea of ​​Freedom as the Semantic Center of Dostoevsky&amp;#39;s Novel &amp;quot;Notes from the House of the Dead&amp;quot;, Science and Education, 11&amp;ndash;15.</mixed-citation></ref><ref id="B4"><mixed-citation>Galperin,&amp;nbsp;I.&amp;nbsp;R. (2006). Tekst kak object lingvisticheskogo issledovaniya [Text as an object of a linguistic study], KomKniga, Moscow, Russia.</mixed-citation></ref><ref id="B5"><mixed-citation>Goldenveyzer,&amp;nbsp;A.&amp;nbsp;B. (1922). Vblizi Tolstogo [Near Tolstoy], Cooperative Publishing House, Moscow, Russia. (In Russian)</mixed-citation></ref><ref id="B6"><mixed-citation>Gorina,&amp;nbsp;Ye.&amp;nbsp;V. (2021). Smyslovaya struktura zhurnalistskogo teksta: uchebno-metodicheskoe posobie [The Semantic Structure of Journalistic Text: A Textbook], Uralskiy federal&amp;#39;nyy universitet imeni pervogo Prezidenta Rossii B.&amp;nbsp;N.&amp;nbsp;Yeltsina, Yekaterinburg, Russia. (In Russian)</mixed-citation></ref><ref id="B7"><mixed-citation>Greimas,&amp;nbsp;A. (1985). V poiskakh transformatsionnykh modelei. Zarubezhnye issledovaniya po semiotike folklora [In Search of Transformational Models: International Research on the Semiotics of Folklore], Nauka, Moscow, Russia. (In Russian)</mixed-citation></ref><ref id="B8"><mixed-citation>Grigoryan,&amp;nbsp;V.&amp;nbsp;A. (2024). Textual Conceptualization and Interpretation in Modern Linguistics, Seventeenth Annual Scientific Conference. Social Sciences and Humanities, l, 513.</mixed-citation></ref><ref id="B9"><mixed-citation>Dzyaloshinskiy,&amp;nbsp;I.&amp;nbsp;M. (2019). Texts and textoids, or what happens to the author? PR i SMI v Kazakhstane: sbornik nauchnykh trudov [PR and Mass Media in Kazakhstan: A Collection of Scientific Papers], Almaty, Kazakhstan, 17, available at: https://publications.hse.ru/pubs/share/direct/306031376.pdf (Accessed 10 September 2025).</mixed-citation></ref><ref id="B10"><mixed-citation>Kobzeva,&amp;nbsp;O.&amp;nbsp;V. (2018). Violation of Linguistic Norms in Translation in В1-С1 Students, Vestnik Kemerovskogo gosudarstvennogo universiteta, 4, available at: https://doi.org/10.21603/2078-8975-2018-4-211-222 (Accessed 03&amp;nbsp;October&amp;nbsp;2025).</mixed-citation></ref><ref id="B11"><mixed-citation>Kovaleva,&amp;nbsp;L.&amp;nbsp;M. (1987). Problema strukturno-semanticheskogo analiza prostoi glagolnoi konstruktsii v sovremennom anglijskom yazyke [The Problem of Structural and Syntactic Analysis of the Simple Verb Constructions in Temporary English], Izd-vo Irkutstkogo un-ta, Irkutsk, Russia. (In Russian)</mixed-citation></ref><ref id="B12"><mixed-citation>Malyavina,&amp;nbsp;A.&amp;nbsp;N. (2024). Teaching Post-editing to Translaton Students, Aktualnyye problemy lingvistiki i metodiki prepodavaniya inostrannykh yazykov, 35.</mixed-citation></ref><ref id="B13"><mixed-citation>Moskalskaya,&amp;nbsp;O.&amp;nbsp;I. (1981). Grammatika teksta (posobie po grammatike nemetskogo yazyka dlya institutov i fakultetov inostrannykh yazykov) [Text Grammar (German Grammar Manual for Institutes and Faculties of Foreign Languages)], Vyssh. shkola, Moscow, Russia, available at: https://www.phantastike.com/linguistics/grammatika_teksta/djvu/view/ (Accessed 15&amp;nbsp;September&amp;nbsp;2025) (In Russian)</mixed-citation></ref><ref id="B14"><mixed-citation>Naumchik,&amp;nbsp;O.&amp;nbsp;S. (2020). The Image of a Mirror as the Semantic Center of Neil Gaiman&amp;#39;s Short Story Collection &amp;quot;Smoke and Mirrors&amp;quot;, Vestnik Baltiyskogo federalnogo universiteta im. I. Kanta. Seriya: Filologiya, pedagogika, psikhologiya, 1, 80&amp;ndash;87.</mixed-citation></ref><ref id="B15"><mixed-citation>Panasenkov,&amp;nbsp;N.&amp;nbsp;A. (2019). Experience of teaching linguistics students - how to post-edit machine-generated translation (based on English-to-Russian translations via Google Translate, Yandex Translate and Promt systems), Pedagogicheskoye obrazovaniye v Rossii, 1, available at: https://doi.org/10.26170/po19-01-08 (Accessed 07&amp;nbsp;September&amp;nbsp;2025).</mixed-citation></ref><ref id="B16"><mixed-citation>Perekhodko,&amp;nbsp;I.&amp;nbsp;V. (2017). Assessing the quality of computer translation, Vestnik Orenburgskogo gosudarstvennogo universiteta, 2&amp;nbsp;(202), 93.</mixed-citation></ref><ref id="B17"><mixed-citation>Psurtsev,&amp;nbsp;D.&amp;nbsp;V. (2001). The Meaning-Forming Aspect of Figurative-Associative Components of a Fiction Text (Based on English-Language Fiction), Ph.D. dissertation, Moscow State Linguistic University, Moscow, Russia.</mixed-citation></ref><ref id="B18"><mixed-citation>Sdobnikov,&amp;nbsp;V.&amp;nbsp;V. (2025). AI in Translation: How to Use it Effectively, Nauchnyi dialog, 14(3), available at: https://doi.org/10.24224/2227-1295-2025-14-3-62-80 (Accessed 01&amp;nbsp;September&amp;nbsp;2025).</mixed-citation></ref><ref id="B19"><mixed-citation>Sdobnikov,&amp;nbsp;V.&amp;nbsp;V. (2024). Artificial Intelligence in Translation: Clarification of Concepts, Voyenno-filologicheskiy zhurnal, 4, 42.</mixed-citation></ref><ref id="B20"><mixed-citation>Sirotinina,&amp;nbsp;O.&amp;nbsp;B. (1994). Texts, textoids, discourses in the zone of colloquial speech, Chelovek. Tekst. Kultura. Yekaterinburg, 109.</mixed-citation></ref><ref id="B21"><mixed-citation>Skrashchuk,&amp;nbsp;Ye.&amp;nbsp;I. (2019). The Semantic Dominant of V.&amp;nbsp;A.&amp;nbsp;Soloukhin&amp;#39;s Short Story &amp;quot;Back Alley&amp;quot;, Dni nauki studentov Vladimirskogo gosudarstvennogo universiteta imeni Aleksandra Grigoryevicha i Nikolaya Grigoryevicha Stoletovykh: Sbornik materialov nauchno-prakticheskikh konferentsiy [Days of Science for Students of Vladimir State University named after A.&amp;nbsp;G. and N.&amp;nbsp;G.&amp;nbsp;Stoletov: Collection of Materials from Scientific and Practical Conferences], Vladimir, Russia 2311&amp;ndash;2319.</mixed-citation></ref><ref id="B22"><mixed-citation>Fonova,&amp;nbsp;Ye.&amp;nbsp;G., Shitts,&amp;nbsp;O.&amp;nbsp;A. (2025). On the Issue of Professional Competencies of Translators in the Age of Artificial Intelligence, Vestnik Tomskogo gosudarstvennogo pedagogicheskogo universiteta, 1(237), available at: https://doi.org/10.23951/1609-624X-2025-1-148-156 (Accessed 01&amp;nbsp;September&amp;nbsp;2025).</mixed-citation></ref><ref id="B23"><mixed-citation>Chakyrova,&amp;nbsp;Yu.&amp;nbsp;I. (2013). Is post-editing a blessing or a curse? Industriya perevoda, 1, 137.</mixed-citation></ref><ref id="B24"><mixed-citation>Tarasti,&amp;nbsp;E. (2017). The Semiotics of A. J. Greimas: A European Intellectual Heritage Seen from the Inside and the Outside. Sign Systems Studies, 45(&amp;frac12;), available at: https://doi.org/10.12697/SSS.2017.45.1-2.03 (Accessed 19&amp;nbsp;May&amp;nbsp;2026). (In English)</mixed-citation></ref><ref id="B25"><mixed-citation>Corpus Material</mixed-citation></ref><ref id="B26"><mixed-citation>Linguistic Encyclopedic Dictionary (1990), available at: https://tapemark.narod.ru/les/392d.html (Accessed 03&amp;nbsp;September&amp;nbsp;2025).</mixed-citation></ref></ref-list></back></article>