Speakers
Description
This study evaluates three AI assistants' ability to generate level-appropriate Elicited Imitation materials for Japanese EFL learners. This phase of our research utilized the New General Service List and NGSL-Speaking list to define linguistic parameters. A One-Way MANOVA comparing models (Wilks’ Lambda = 0.42, p < .001) revealed that Claude significantly outperformed ChatGPT and Gemini in vocabulary compliance and sentence length precision. These findings establish a foundation for our future phonetic training interventions.
Short summary
This study evaluates three AI assistants' ability to generate level-appropriate Elicited Imitation materials for Japanese EFL learners. This phase of our research utilized the New General Service List and NGSL-Speaking list to define linguistic parameters. A One-Way MANOVA comparing models (Wilks’ Lambda = 0.42, p < .001) revealed that Claude significantly outperformed ChatGPT and Gemini in vocabulary compliance and sentence length precision. These findings establish a foundation for our future phonetic training interventions.
Abstract
Improving speaking comprehensibility and intelligibility remains a critical challenge for Japanese EFL learners. This poster details Phase 1 of a multi-year Kaken research project evaluating the capacity of AI assistants to generate level-appropriate Elicited Imitation (EI) materials for learners at CEFR A2 to B1 levels. The investigation focused on establishing linguistic parameters to ensure materials matched learner processing capacities, specifically targeting grammatical complexity and sentence length while restricting vocabulary to the New General Service List of 2,818 words and the NGSL-Speaking list of 721 words. The methodology involved a systematic approach to prompt engineering by assigning three AI models the role of Expert Linguist and Materials Creator and providing empirical examples to guide output consistency. A comparative analysis was conducted across ChatGPT, Claude, and Gemini to determine which model most reliably adhered to these constraints. A One-Way MANOVA was utilized to measure the effect of the AI model on grammatical accuracy, vocabulary compliance, and length precision. Results revealed a statistically significant difference between models: Wilks Lambda = 0.42, F(9, 238) = 12.45, p < .001. Follow-up univariate ANOVAs indicated that Claude significantly outperformed competing models in vocabulary compliance (p < .01) and sentence length precision (p < .05). These results establish Claude as the most effective tool for producing calibrated educational content in our context, providing a good foundation for subsequent experimental phases involving High Variability Phonetic Training.
Keywords
Elicited Imitation
High Variability Phonetic Training
AI Assistants
Materials Assessment
Special scheduling requests
I prefer Saturday poster session if possible since I have limited time on Sunday; but will present on Sunday if that is the only poster day. Sorry for the request.
| Scheduling preference | Anytime on Saturday |
|---|---|
| Title | Evaluating AI Models for Level-Appropriate Elicited Imitation Materials |