Speakers
Description
This poster presents Phase 1 of a multi-year study (Kaken) evaluating three AI models’ capacity to generate level-appropriate Elicited Imitation sentences for Japanese EFL learners at CEFR A2-B1. Using systematic prompt engineering and vocabulary constraints based on the New General Service List, a MANOVA comparing ChatGPT, Claude, and Gemini revealed significant differences between models. Claude outperformed competitors in vocabulary compliance and sentence length precision, establishing it as the optimal tool for subsequent phonetic training phases.
Abstract section 1: Relevance
Improving speaking comprehensibility and intelligibility remains a persistent challenge for Japanese EFL learners. While High Variability Phonetic Training and Elicited Imitation offer evidence-based approaches, creating level-appropriate materials at scale is resource-intensive. Meta-analytic evidence confirms HVPT’s effects on L2 perception (Uchihara et al., 2025), and AI-generated speech shows promise for HVPT delivery (Al-Shami & Cardoso, 2025). Systematic reviews document rapid growth in generative AI for language education (Li et al., 2025), yet research on AI-generated speaking assessment materials remains limited. This study investigates whether large language models can reliably produce calibrated EI sentences matching predefined linguistic criteria for lower-proficiency learners.
Abstract section 5: References
Al-Shami, A., & Cardoso, W. (2025). Text-to-speech in high-variability phonetic training: Focus on L2 phonological awareness. Computer-Assisted Language Learning Electronic Journal, 26(6), 21–42.
Li, B., Tan, Y. L., Wang, C., & Lowell, V. (2025). Two years of innovation: A systematic review of empirical generative AI research in language learning and teaching. Computers and Education: Artificial Intelligence, 9, 100445. https://doi.org/10.1016/j.caeai.2025.100445
Uchihara, T., Karas, M., & Thomson, R. I. (2025). High variability phonetic training (HVPT): A meta-analysis of L2 perceptual training studies. Studies in Second Language Acquisition, 47(3), 794–827. https://doi.org/10.1017/S0272263125100879
Abstract section 4: Outcomes/results
MANOVA revealed significant multivariate effects for AI model type (Wilks’ Lambda = 0.42, F(9, 238) = 12.45, p < .001). Follow-up univariate ANOVAs showed Claude significantly outperformed ChatGPT and Gemini in vocabulary compliance (p < .01) and sentence length precision (p < .05), while no significant differences emerged for grammatical complexity. These findings demonstrate that carefully prompted AI models can produce calibrated EI materials, though model selection matters substantially. The poster details prompt engineering methodology, examines sample materials across CEFR levels, and explains how AI-generated content integrates into a larger HVPT research program targeting Japanese EFL learners’ speaking development.
Abstract section 2: Contribution/research questions
This poster presents a systematic methodology for using AI to generate pedagogical materials. The central research question is: To what extent can AI assistants generate level-appropriate elicited imitation sentences for CEFR A2-B1 Japanese EFL learners? We contribute a replicable prompt engineering framework that controls grammatical complexity, restricts vocabulary to established word lists, and calibrates sentence length to learner processing capacity, enabling students to efficiently produce validated assessment materials.
Abstract section 3: Content/method
Three AI models (ChatGPT, Claude, Gemini) generated EI sentences (n=150 per model, 50 per CEFR level) using identical prompts incorporating expert role assignment, empirical output examples, and constraints targeting the New General Service List and NGSL-Speaking list. Three expert raters evaluated materials on grammatical complexity, vocabulary appropriateness, and sentence length suitability. A one-way MANOVA assessed model effects across all dependent variables, with follow-up univariate ANOVAs and Tukey HSD comparisons identifying specific differences.
| Title | Comparing AI Models for Scalable EI Materials in EFL |
|---|---|
| Teaching Context | College and university education |