Abstract: Artificial intelligence (AI) is becoming an unavoidable element of higher education, military training, and language learning. For future officers, however, English proficiency cannot be reduced to the efficient production of grammatically acceptable texts or rehearsed oral responses. Within the NATO environment, preparation for STANAG 6001 must remain anchored in independent, unrehearsed, general communicative ability across listening, speaking, reading, and writing. This article examines how AI can support English-language education for officer cadets preparing for STANAG 6001 examinations while preserving the validity, integrity, and professional purpose of language assessment.
Problem statement: How should future officers be trained to use powerful language technologies without becoming dependent on them?
So what?: AI-supported language learning must be governed by clear policies and task designs from military educational institutions that align with the descriptors of STANAG 6001 and the professional requirements of officer education. Instructors should harness AI to intensify practice, simulate communicative pressure, and provide formative feedback, but cadets need to learn to verify, revise, and eventually operate without technological aids. The conceptual shift is clear: AI should be positioned not as an answer generator but as a disciplined training environment for communicative readiness.

Introduction
The rapid diffusion of artificial intelligence into education has posed a difficult but necessary question for military educational institutions: How should future officers be trained to use powerful language technologies without becoming dependent on them? In civilian higher education, this question is often framed in terms of academic integrity, employability, or digital literacy. In military education, it has an additional operational dimension. Language proficiency supports coalition cooperation, briefing, reporting, staff work, negotiation, situational understanding, and participation in multinational command structures. For officer cadets, English is not merely an academic subject. It is part of professional readiness.
The issue is particularly important in relation to STANAG 6001, the NATO language proficiency standard used for curriculum, test development, and the recording and reporting of Standardised Language Profiles. BILC describes STANAG 6001, Edition 5, as the NATO-agreed standard for language curriculum, test development, and the reporting of SLPs. The NATO standard defines language proficiency as an individual’s unrehearsed general language communication ability and describes proficiency in listening, speaking, reading, and writing.[1], [2] If the construct to be assessed is unrehearsed communicative ability, then AI-supported preparation must not train candidates merely to outsource linguistic production to a machine.
Generative AI systems can draft essays, summarise texts, generate exam-style prompts, correct grammar, simulate interviews, and provide immediate feedback. Used carefully, these tools can expand access to practice and individualise language development. Used carelessly, they can create an illusion of competence. The learner may appear able to write, argue, or summarise at a higher level than they can sustain independently. In high-stakes military language examinations and in operational settings beyond the classroom, such false fluency is not a minor pedagogical problem; it is a risk to the credibility of certification and to the professional preparation of officers.
STANAG 6001 and the Purpose of Language Proficiency in the Military
STANAG 6001 plays a distinctive role within the Alliance by providing a common language proficiency scale across national systems. Its purpose is not to certify completion of a particular course but to describe what an individual can do in the target language. This distinction matters for pedagogy. A course may prepare a learner for a test. Still, the test is intended to reflect broader proficiency rather than familiarity with a single textbook, institutional syllabus, or set of rehearsed tasks.[3] Studies and test-development discussions on STANAG 6001 similarly show that score interpretation, test specifications, and national implementation must remain linked to the common proficiency construct rather than to local classroom routines.[4], [5]
The standard outlines six proficiency levels, ranging from Level 0 (no practical proficiency) to Level 5 (highly articulate native or bilingual proficiency). For most officer education contexts, the critical developmental trajectory spans Levels 1 (survival), 2 (functional), and 3 (professional), and, for selected posts or advanced roles, Level 4 (expert proficiency). BILC’s overview of STANAG 6001 levels emphasises that simplified descriptors support the interpretation of SLPs for job descriptions and positional requirements but do not replace full-level descriptors.[6] This warning is pedagogically significant. Preparation must not be reduced to memorising simplified tables; it must cultivate the ability to perform the communicative functions represented by the descriptors.
At Level 2, candidates are expected to manage routine social and work-related communication with sufficient accuracy and control, including narration, description, basic explanation, and participation in predictable exchanges. At Level 3, they must move beyond routine performance: they should handle unfamiliar topics, support opinions, explain abstract issues, follow extended discourse, and produce organised, paragraph-length speech or writing. At Level 4, demands are higher: candidates need to understand nuance, implication, and sophisticated argumentation, and to use language precisely, flexibly, and appropriately in a variety of complex professional and social contexts.[7]
This distinction becomes clearer when the four skills are considered together. In listening, progression moves from understanding concrete information to following extended, less predictable discourse and interpreting implied meaning. In speaking, progression moves from functional exchanges to sustained explanation, supported argument, flexible interaction, and the ability to repair communication under pressure. In reading, progression moves from locating information to interpreting the authorial stance, making inferences, understanding structure, and evaluating argument. In writing, progression moves from sentence-level control and simple organisation to coherent, paragraph-level organisation and then to extended discourse with appropriate register, lexical precision, and rhetorical control. AI-supported training is valuable only if it helps cadets progress along these developmental lines.
This creates a productive tension in military English education. On the one hand, cadets need professionally relevant training: briefings, reports, orders, multinational interaction, operational vocabulary, and military scenarios. On the other hand, STANAG 6001 assesses general language proficiency rather than narrow occupational performance. A candidate may need military English, but the examination is not a vocabulary test of military terminology. The strongest preparation, therefore, combines general proficiency development with military-professional communicative situations, consistent with the broader English for Specific Purposes tradition and with work on assessing languages for specific purposes.[8], [9] This is where AI can be useful, as it can generate varied topics, reformulate texts at different levels, simulate interactions, and provide targeted feedback. Yet the same flexibility can also distort preparation if instructors do not maintain descriptor alignment.
In language testing theory, validity depends on whether a test interpretation is justified by the construct being assessed. Bachman, Bachman, and Palmer; Fulcher; and McNamara have all, in different ways, stressed that language assessment must define which ability is being measured and under what conditions.[10], [11], [12], [13] For STANAG 6001, the construct is not the ability to co-author a fluent text with AI. It is the candidate’s own ability to comprehend and produce language. AI-supported instruction is therefore legitimate only when it helps the candidate internalise skills that can later be demonstrated without assistance.
AI as a Pedagogical Instrument in STANAG 6001 Preparation
The best way to leverage AI to support STANAG 6001 prep is through formative learning, not as a substitute for summative assessment. Formative use means AI helps learners identify gaps, practice effectively, receive feedback, compare performance, and reflect on improvement. It doesn’t mean submitting machine-generated output as evidence of competence. UNESCO calls for human-centred use, policy formulation, and safeguards to ensure privacy, equity, and quality in the use of generative AI in education and research.[14] Similarly, the European Commission’s ethical guidelines encourage educators to adopt a positive, critical, and ethical approach to AI rather than treating it as either good or bad.[15] A broader body of literature on AI in education echoes this cautionary note: the pedagogical benefits of AI lie less in technological novelty than in task design, teacher judgment, learner agency, and institutional governance.[16], [17]
In the listening area, AI can help teachers prepare pre-listening vocabulary exercises, post-listening comprehension questions, summaries of authentic audio transcripts, and targeted activities on gist, detail, speaker attitude, and implied meaning. Speech-to-text tools can also support delayed feedback by allowing cadets to compare what they heard with a transcript. However, listening competence cannot be developed solely through transcripts. Cadets must still confront authentic speed, accent variation, incomplete information, and cognitive load. So AI should facilitate metacognitive awareness of listening problems, not eliminate the difficulty that makes listening practice worthwhile.
AI chatbots and voice tools can simulate interviews, role plays, and follow-up questions to help you prepare to speak. They can help students practice giving opinions, structuring responses, clarifying meanings, and maintaining the interaction. They can also provide feedback on coherence, lexical range, and recurrent grammatical errors. The danger is rehearsed fluency. If memorised AI-generated scripts, they may sound competent on familiar topics but fail when the examiner shifts the angle, asks for justification, or moves into an unfamiliar domain. Speaking preparation should therefore include AI-assisted rehearsal followed by human-led disruption: unanticipated follow-up questions, counterarguments, role changes, and timed responses.
In the context of reading preparation, AI can scaffold complex texts by prompting learners to identify the thesis, argument structure, assumptions, evidence, tone, and implications. It can generate questions of varying difficulty levels and contrast literal and inferential comprehension. This is particularly valuable when moving from Level 2 to Level 3 and from Level 3 to Level 4, where candidates must increasingly manage abstract topics, argumentation, and nuance. Yet AI summaries should not replace reading. If a candidate reads only the AI-generated explanation, the core reading skill is bypassed. Sound design requires the learner to read first, answer independently, use AI for feedback or alternative interpretations, and then return to the text.
Writing is both the most promising and the most vulnerable domain. Generative AI can provide models, outline structures, offer corrective feedback, and explain register. It can help cadets compare a weak paragraph with a stronger version and understand why one is more coherent, precise, or appropriately formal. Research on computer-assisted language learning, AI-supported writing, and second-language writing suggests considerable potential for revision, noticing, and learner autonomy, especially when feedback is used reflectively rather than copied.[18], [19], [20], [21], [22] At the same time, writing is the easiest skill to outsource. If a cadet submits an AI-generated essay, the result may be linguistically polished yet provide little evidence of the cadet’s own ability. For STANAG-oriented writing, the central rule should be: AI may support planning, diagnosis, and revision, but the final assessed performance must be independently produced.
STANAG 6001 Descriptor Alignment
The most important design principle for AI-supported STANAG 6001 preparation is descriptor alignment. Instructors should begin not with a tool but with the target proficiency behaviour. Descriptor alignment also helps distinguish useful AI tasks from superficial ones. A prompt such as ‘write me an essay about leadership’ may produce fluent text, but it does not necessarily develop the learner’s capacity. A more pedagogically sound task would ask the cadet to write an initial response, compare it against descriptor-informed criteria, ask AI to identify weaknesses in organisation and lexical precision, revise the response independently, and then explain the revisions. The object of learning is not the generated answer but the cadet’s growing control over discourse.
A similar logic applies to speaking. AI can prompt a cadet to defend a position on compulsory military service, cyber defence, or multinational exercises. The task becomes educationally valuable only if the cadet must respond spontaneously, receive feedback aligned with relevant descriptors, reformulate the answer, and then attempt a new topic without support. AI is useful because it can multiply practice opportunities; it is dangerous if it replaces the productive struggle through which proficiency develops.
The following editable matrix summarises how AI affordances can be aligned with STANAG-oriented language development while preserving safeguards. It is intended as a planning tool for instructors rather than a prescriptive checklist.
| STANAG skill | AI-supported affordances | Main risk | Instructor safeguard |
| Listening | Gist/detail questions; accent and topic variation; transcript-based delayed feedback. | Transcript dependence: reduced tolerance for authentic speed, accent, and ambiguity. | First listen without a transcript; use AI only for analysis, reflection, and targeted re-listening. |
| Speaking | Interview simulation; follow-up questions; role-play; feedback on organisation and range. | Rehearsed fluency and memorised AI-generated responses. | Add human follow-up, counter-argument, time pressure, and spontaneous reformulation. |
| Reading | Inference questions; argument mapping; explanation of tone, stance, and implication. | Replacing reading with AI summaries. | Require independent answers first; verify interpretations against textual evidence. |
| Writing | Planning, revision feedback, error-pattern analysis, and register comparison. | Ghost-writing, false proficiency, and integrity breaches. | Use AI for diagnosis and revision; final assessed writing remains unassisted. |
Risks, Constraints, and Professional Safeguards
Any serious discussion of AI in military English education must address the following risks:
- Linguistic dependency – students might become more and more dependent on AI for sentence writing, choice of vocabulary, and argument structuring. Over time, it can erode retrieval, decrease tolerance for ambiguity, and inhibit independent formulation;
- False fluency – the output of AI may be accurate, idiomatic, and persuasive, even when the learner lacks the competence to reproduce it. This creates a gap between the training artefact and the test performance in the STANAG 6001 sense;
- Automation bias – for many years, research on automation has shown that users tend to over-trust automated systems, particularly when outputs appear confident and authoritative.[23], [24] In language learning, this can lead cadets to follow AI corrections even when they are stylistically inappropriate, factually incorrect, or misaligned with the task. In military education, such uncritical acceptance is particularly problematic because professional officers must learn to question sources, assess reliability, and remain accountable for their own communication;
- Academic integrity and assessment validity – higher education research has already documented the challenge of distinguishing legitimate AI support from cheating or misrepresentation.[25], [26], [27] In STANAG preparation, the boundary must be clear. AI can be part of learning; it cannot be part of a candidate’s performance in a controlled proficiency examination unless the assessment construct explicitly includes AI-mediated communication, which STANAG 6001 does not. Military educational institutions should therefore provide cadets with explicit categories for permitted, limited, and prohibited uses of AI; and
- Data protection and operational security – cadets and instructors should not enter sensitive information, personal data, internal testing materials, or restricted military scenarios into public AI systems. Even seemingly harmless classroom prompts may contain institutional patterns, personnel details, or examination content. UNESCO, OECD, and the European Commission all emphasise the need for governance, privacy protection, transparency, and responsible human oversight.[28], [29], [30] In military settings, these requirements are not optional administrative concerns; they are part of professional discipline.
The most effective safeguard is not prohibition but structured use. Cadets should learn when AI is useful, when it is inappropriate, how to verify its output, and how to document its role in learning. Instructors should design tasks that require independent attempts first, AI-supported analysis and revision, and oral or written reflection. Assessment scholarship in the digital environment increasingly emphasises authentic assessment, assessment security, feedback literacy, and long-term learning rather than simply detecting misconduct.[31], [32], [33] STANAG-oriented courses should therefore include unassisted performance under timed or supervised conditions. In this way, AI becomes a means of developing competence rather than concealing its absence.
| Category | Permitted example | Limited/conditional example | Prohibited example |
| Formative learning | Asking AI to explain recurring errors after an independent attempt. | AI-suggested structures followed by independent rewriting and reflection. | Submitting AI-written paragraphs as the cadet’s own work. |
| Speaking practice | Generating practice questions on non-sensitive topics. | AI rehearsal before a human-led mock interview. | Memorising AI-generated scripts for examination use. |
| Reading practice | Generate inference questions after reading the source text. | Explaining difficult passages after independent comprehension work. | Replacing reading with AI summaries. |
| Security and data | Using generic, non-sensitive prompts. | Using approved institutional tools under policy. | Uploading test items, personal data, or restricted content to public systems. |
A Responsible Framework: DARE-HS
To operationalise responsible AI use in the preparation of STANAG 6001, this article proposes the DARE-HS framework: Descriptor Alignment, Authentic Performance, Responsible Use, Evaluation and Reflection, Human Supervision, and Security. The framework is intentionally simple because military education requires models that can be communicated, implemented, and audited. It is not a replacement for institutional policy or assessment regulations; rather, it serves as a pedagogical bridge between AI affordances and language proficiency requirements.
Descriptor Alignment requires mapping AI activities to the relevant STANAG proficiency behaviours rather than to generic English practice. Authentic performance occurs when the cadet can demonstrate independent listening, speaking, reading, or writing, especially after preparation using AI. Responsible use means transparency, academic integrity, and awareness of AI limitations. Evaluation and reflection mean learners explain what changed, why, and how they can do better without AI. Human supervision retains the instructor’s responsibility for task design, interpretation of feedback, and professional judgment. Security reminds users that public AI systems are not neutral notebooks for sensitive military information.
This framework facilitates a gradual-release model. In the early stages, AI could provide more scaffolding: vocabulary support, grammar explanations, model comparisons, and guided practice. As the candidate approaches the examination, AI support should be gradually reduced. The final objective is unassisted performance. This mirrors the logic of military training more broadly: simulation, rehearsal, and feedback are valuable because they prepare individuals to act without the simulator at the decisive moment.
| DARE-HS element | Purpose | Operationalisation in STANAG 6001 preparation |
| Descriptor alignment | Connect AI tasks to proficiency behaviours. | Map prompts and feedback to descriptors for listening, speaking, reading, or writing. |
| Authentic performance | Protect independent communicative ability. | Use unassisted first attempts and unassisted final performances. |
| Responsible use | Maintain integrity and transparency. | Declare AI use where required; distinguish support from substitution. |
| Evaluation and reflection | Convert feedback into learning. | Explain revisions, rejected suggestions, and remaining weaknesses. |
| Human supervision | Keep professional judgment central. | Instructors validate materials, interpret feedback, and assess performance. |
| Security | Prevent inappropriate disclosure. | Exclude sensitive data, live test material, and restricted military scenarios. |
Implications for Military Educational Institutions and Lectors
Military educational institutions should integrate AI literacy into language education and not as a separate digital skills module. Cadets preparing for STANAG 6001 must learn not just how to use AI effectively, but also why some uses undermine competence. Institutional clarity required. An AI use policy for language courses should be brief and cover acceptable formative uses, prohibited assessment uses, data privacy regulations, and transparency requirements. Such a policy should comply with national regulations, institutional academic integrity policies, and NATO-related security requirements.[34]
Professional development for instructors is also needed. Effective use of AI in preparing STANAGs is more than just writing a prompt. They require knowledge of language assessment, proficiency descriptors, feedback literacy, task authenticity, and ethical constraints. UNESCO’s AI competency framework for teachers is relevant here because it frames teacher competence around human-centred mindsets, ethics, AI foundations, pedagogy, and professional learning. For military English instructors, these dimensions should be supplemented by awareness of test security, military data protection, and the specific construct of STANAG 6001 proficiency.
A practical course design might include four recurring activity types. First, diagnostic tasks in which cadets complete unassisted listening, reading, speaking, or writing tasks and then use AI-supported feedback to identify patterns of weakness. Second, descriptor comparison tasks in which learners compare Level 2 and Level 3 versions of a response and discuss differences in organisation, range, accuracy, and communicative effectiveness. Third, scenario-based tasks in which AI generates varied but non-sensitive prompts for briefings, interviews, and professional discussions. Fourth, reflection tasks in which cadets explain what they accepted, rejected, or modified based on AI feedback.
Such a design preserves the instructor’s professional authority. AI may generate materials, but instructors determine their appropriateness. The AI can provide feedback, but the instructor teaches learners how to interpret it. AI can simulate interaction, but human examiners and teachers are still needed to assess communicative competence in context. This is especially true for speaking and writing, where language performance means coherence, pragmatic appropriateness, register, reasoning, flexibility, and the ability to respond to pressure.
Military educational institutions should also maintain a clear separation between training data and evaluation materials. Do not upload internal examination prompts, live test items, candidate responses, or any confidential institutional materials to publicly available AI systems. Procurement and governance need to address privacy, data retention, access control, and accountability when using AI tools in an institutional setting. NATO’s recent emphasis on digital transformation demonstrates that adapting to technology is not only a technical matter, but also an organisational and cultural one.[35] Language education should be part of that transformation, but in a disciplined, professionally governed form.
Recommendations
Several recommendations regarding the use of AI in English language education and the preparation of cadets for the STANAG 6001 English language examination were identified:
- Military educational institutions should adopt a concise AI use policy for language education that distinguishes four contexts: informal self-study, formative classroom work, coursework submitted for feedback, and formal assessment. The policy should define permitted support, limited support, and prohibited substitution in terms that cadets can understand and that instructors can enforce;
- Each STANAG 6001 preparation module should be designed backwards from descriptors. Course teams should first identify the target communicative behaviour, then select AI-supported activities only if they support that behaviour. For example, a listening lesson should specify whether the aim is the gist, details, attitude, or implications before using AI to generate questions or transcripts;
- Cadets should be required to submit an independent first attempt before receiving AI feedback. This principle protects productive struggle and provides both learner and instructor with evidence of actual ability. In writing, the first draft should be human-generated; in speaking, the first response should be spontaneous; in reading and listening, first answers should be completed before consulting any AI-generated explanation;
- AI feedback should be accompanied by reflection and disclosure. Cadets should be able to explain which AI suggestions they accepted or rejected and how the changes relate to clarity, accuracy, coherence, register, task fulfilment, or descriptor-relevant performance. A simple AI-use declaration for formative writing tasks would normalise transparency without turning every task into a disciplinary issue;
- English language instructors should receive targeted professional development in AI literacy, prompt design, language assessment, STANAG descriptor interpretation, and data protection. The required competence is not merely technical. It is the ability to determine when an AI-generated task is pedagogically valid, linguistically appropriate, ethically acceptable, and professionally safe;
- Speaking preparation should preserve unpredictability. AI-generated practice questions are useful for building frequency and confidence. Still, human-led follow-up, interruption, challenge, reformulation, topic shifts, and time pressure remain necessary to develop spontaneous communicative control. The closer cadets move to the examination, the more support should be withdrawn;
- Writing instruction should treat AI as a revision tutor, not a ghostwriter. Cadets may use AI to diagnose weaknesses, compare alternatives, and understand register, but the final assessed writing must be independently produced. When institutional policy permits AI-supported drafts in formative work, instructors should require process evidence: first draft, feedback, revision, and reflection;
- Military educational institutions should establish a protected-materials rule. Internal test items, live prompts, candidate scripts, personal data, non-public scenarios, and restricted or sensitive military content should never be uploaded to public AI systems. When approved institutional tools are used, procurement and governance should address data retention, access control, auditability, and accountability; and
- Observable improvements in unassisted performance should evaluate the effectiveness of AI-supported STANAG preparation. Military educational institutions should monitor whether AI-supported practice improves independent listening, reading, speaking, and writing using diagnostic baselines, periodic unassisted checks, mock-examination performance, and instructor judgment. The decisive question is not whether cadets can produce better English with AI, but whether AI-supported practice helps them communicate more accurately, coherently, and confidently when AI is unavailable.
Conclusion
Artificial intelligence will remain part of the educational environment where tomorrow’s officers are trained. The question is not whether cadets will encounter AI, but whether military educational institutions will shape its use to strengthen professional competence. In English language education, particularly in preparation for STANAG 6001, the stakes are high because the examination is intended to certify independent communicative proficiency across four skills. A candidate’s ability to produce polished language with technological assistance is not equivalent to the ability to listen, read, speak, and write under authentic professional conditions.
The most defensible approach is neither prohibition nor unrestricted adoption. AI should be integrated as a controlled training aid: useful for diagnosis, feedback, scenario variation, revision, and reflection, yet always subordinate to descriptor alignment, human supervision, assessment validity, and security. When used this way, AI can help military English instructors provide more individualised, intensive, and varied preparation. It can also help cadets become more reflective learners who better understand their language development.
The final aim must therefore remain unchanged. Future officers need communicative competence that is reliable without assistance, adaptable under pressure, and credible in multinational military environments. AI can support that aim only if it is used to build capability, not to simulate it. For STANAG 6001 preparation, the central measure of success is not the sophistication of AI-mediated output but the officer candidate’s ability to understand, speak, read, and write independently when professional communication matters.
[1] Bureau for International Language Co-ordination, “STANAG 6001,” accessed May 15, 2026, https://natobilc.org/stanag-6001/.
[2] NATO Standardization Office, NATO Standard ATrainP-5: Language Proficiency Levels, Edition A, Version 2 (Brussels: NATO Standardization Office, 2016), 1–2, https://natobilc.org/wp-content/uploads/2024/11/ATrainP-5-EDA-V2-E.pdf.
[3] Bureau for International Language Co-ordination, Best Practices in STANAG 6001 Testing (BILC, August 2025), https://natobilc.org/wp-content/uploads/2025/08/Best_practices_in_stanag_6001_testing_august_2025-upd.pdf.pdf.
[4] Richard J. Tannenbaum and E. Caroline Wylie, Mapping TOEIC Test Scores to the STANAG 6001 English Language Proficiency Levels, ETS Research Memorandum RM-10-11 (Princeton, NJ: Educational Testing Service, 2010).
[5] Ekrem Solak, “NATO STANAG Language Proficiency Levels for Joint Missions,” 2013, https://files.eric.ed.gov/fulltext/ED553405.pdf.
[6] Bureau for International Language Co-ordination, NATO STANAG 6001, Ed. 5: Overview of Language Proficiency Levels (BILC, February 2019), https://natobilc.org/wp-content/uploads/2024/11/STANAG-6001-Overview-Feb-2019.pdf.
[7] Bureau for International Language Co-ordination, STANAG 6001 Level 4 Testing – Tutorial: Testing Receptive Skills (BILC, 2018), https://natobilc.org/wp-content/uploads/2024/11/STANAG-6001-L4-Tutorial-Testing-Receptive-Skills1.pdf.
[8] Dan Douglas, Assessing Languages for Specific Purposes (Cambridge: Cambridge University Press, 2000).
[9] Helen Basturkmen, Developing Courses in English for Specific Purposes (Basingstoke: Palgrave Macmillan, 2010).
[10] Lyle F. Bachman, Fundamental Considerations in Language Testing (Oxford: Oxford University Press, 1990).
[11] Lyle F. Bachman and Adrian S. Palmer, Language Testing in Practice (Oxford: Oxford University Press, 1996).
[12] Glenn Fulcher, Practical Language Testing (London: Routledge, 2010).
[13] Tim McNamara, Language Testing (Oxford: Oxford University Press, 2000).
[14] Fengchun Miao and Wayne Holmes, Guidance for Generative AI in Education and Research (Paris: UNESCO, 2023), https://unesdoc.unesco.org/ark:/48223/pf0000386693.
[15] European Commission, Directorate-General for Education, Youth, Sport and Culture, Ethical Guidelines on the Use of Artificial Intelligence (AI) and Data in Teaching and Learning for Educators (Luxembourg: Publications Office of the European Union, 2022), https://op.europa.eu/en/publication-detail/-/publication/d81a0d54-5348-11ed-92ed-01aa75ed71a1/language-en.
[16] Wayne Holmes, Maya Bialik, and Charles Fadel, Artificial Intelligence in Education: Promises and Implications for Teaching and Learning (Boston: Center for Curriculum Redesign, 2019).
[17] Ben Williamson and Rebecca Eynon, “Historical Threads, Missing Links, and Future Directions in AI in Education,” Learning, Media and Technology 45, no. 3 (2020): 223–235.
[18] Carol A. Chapelle, Computer Applications in Second Language Acquisition: Foundations for Teaching, Testing and Research (Cambridge: Cambridge University Press, 2001).
[19] Carol A. Chapelle and Dan Douglas, Assessing Language through Computer Technology (Cambridge: Cambridge University Press, 2006).
[20] Robert Godwin-Jones, “Partnering with AI: Intelligent Writing Assistance and Instructed Language Learning,” Language Learning & Technology 26, no. 2 (2022): 5–24.
[21] Lucas Kohnke, Benjamin Luke Moorhouse, and Di Zou, “ChatGPT for Language Teaching and Learning,” RELC Journal 54, no. 2 (2023): 537–550.
[22] Shaofeng Li, “Generative AI and Second Language Writing,” Digital Studies in Language and Literature, 2025.
[23] Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human Factors 52, no. 3 (2010): 381–410.
[24] Mary L. Cummings, “Automation Bias in Intelligent Time Critical Decision Support Systems,” AIAA 1st Intelligent Systems Technical Conference (2004).
[25] Debby R. E. Cotton, Peter A. Cotton, and J. Reuben Shipway, “Chatting and Cheating: Ensuring Academic Integrity in the Era of ChatGPT,” Innovations in Education and Teaching International 61, no. 2 (2024): 228–239.
[26] Phillip Dawson, Defending Assessment Security in a Digital World: Preventing E-Cheating and Supporting Academic Integrity in Higher Education (London: Routledge, 2021).
[27] Mike Perkins, “Academic Integrity Considerations of AI Large Language Models in the Post-Pandemic Era: ChatGPT and Beyond,” Journal of University Teaching & Learning Practice 20, no. 2 (2023).
[28] OECD, OECD Digital Education Outlook 2023: Towards an Effective Digital Education Ecosystem (Paris: OECD Publishing, 2023), https://www.oecd.org/en/publications/oecd-digital-education-outlook-2023_c74f03de-en.html.
[29] OECD, AI and the Future of Skills, Volume 2: Methods for Evaluating AI Capabilities (Paris: OECD Publishing, 2023).
[30] UNESCO, AI Competency Framework for Teachers (Paris: UNESCO, 2024), https://www.unesco.org/en/articles/ai-competency-framework-teachers.
[31] Margaret Bearman, Phillip Dawson, Rola Ajjawi, Joanna Tai, and David Boud, “Re-imagining University Assessment in a Digital World,” in Re-imagining University Assessment in a Digital World, ed. Margaret Bearman et al. (Cham: Springer, 2020).
[32] David Boud and Nancy Falchikov, eds., Rethinking Assessment in Higher Education: Learning for the Longer Term (London: Routledge, 2007).
[33] David Carless, Excellence in University Assessment: Learning from Award-winning Practice (London: Routledge, 2015).
[34] NATO, NATO’s Digital Transformation Implementation Strategy, October 17, 2024, https://www.nato.int/en/about-us/official-texts-and-resources/official-texts/2024/10/17/natos-digital-transformation-implementation-strategy.
[35] NATO Standardization Office, AJP-01: Allied Joint Doctrine, Edition F, Version 1 (Brussels: NATO Standardization Office, 2022).








