May 23 – 24, 2026
Chukyo University - Nagoya Campus
Asia/Tokyo timezone

Evaluating AI Models for Level-Appropriate Elicited Imitation Materials

May 23, 2026, 12:45 PM
45m
0号building/2-1 - Yamate Hall (posters) (Chukyo University)

0号building/2-1 - Yamate Hall (posters)

Chukyo University

100
C. Poster Session CALL: Computer Assisted Language Learning Posters - 2F Yamate Hall

Speakers

Omar Massoud (Meiji Gakuin University) Robert Cvitkovic (Teikyo University) Yoko Kita (Kyoto Notre Dame University)

Description

This study evaluates three AI assistants' ability to generate level-appropriate Elicited Imitation materials for Japanese EFL learners. This phase of our research utilized the New General Service List and NGSL-Speaking list to define linguistic parameters. A One-Way MANOVA comparing models (Wilks’ Lambda = 0.42, p < .001) revealed that Claude significantly outperformed ChatGPT and Gemini in vocabulary compliance and sentence length precision. These findings establish a foundation for our future phonetic training interventions.

Short summary

This study evaluates three AI assistants' ability to generate level-appropriate Elicited Imitation materials for Japanese EFL learners. This phase of our research utilized the New General Service List and NGSL-Speaking list to define linguistic parameters. A One-Way MANOVA comparing models (Wilks’ Lambda = 0.42, p < .001) revealed that Claude significantly outperformed ChatGPT and Gemini in vocabulary compliance and sentence length precision. These findings establish a foundation for our future phonetic training interventions.

Abstract

Improving speaking comprehensibility and intelligibility remains a critical challenge for Japanese EFL learners. This poster details Phase 1 of a multi-year Kaken research project evaluating the capacity of AI assistants to generate level-appropriate Elicited Imitation (EI) materials for learners at CEFR A2 to B1 levels. The investigation focused on establishing linguistic parameters to ensure materials matched learner processing capacities, specifically targeting grammatical complexity and sentence length while restricting vocabulary to the New General Service List of 2,818 words and the NGSL-Speaking list of 721 words. The methodology involved a systematic approach to prompt engineering by assigning three AI models the role of Expert Linguist and Materials Creator and providing empirical examples to guide output consistency. A comparative analysis was conducted across ChatGPT, Claude, and Gemini to determine which model most reliably adhered to these constraints. A One-Way MANOVA was utilized to measure the effect of the AI model on grammatical accuracy, vocabulary compliance, and length precision. Results revealed a statistically significant difference between models: Wilks Lambda = 0.42, F(9, 238) = 12.45, p < .001. Follow-up univariate ANOVAs indicated that Claude significantly outperformed competing models in vocabulary compliance (p < .01) and sentence length precision (p < .05). These results establish Claude as the most effective tool for producing calibrated educational content in our context, providing a good foundation for subsequent experimental phases involving High Variability Phonetic Training.

Keywords

Elicited Imitation
High Variability Phonetic Training
AI Assistants
Materials Assessment

Special scheduling requests

I prefer Saturday poster session if possible since I have limited time on Sunday; but will present on Sunday if that is the only poster day. Sorry for the request.

Scheduling preference Anytime on Saturday
Title Evaluating AI Models for Level-Appropriate Elicited Imitation Materials

Author

Robert Cvitkovic (Teikyo University)

Co-authors

Omar Massoud (Meiji Gakuin University) Yoko Kita (Kyoto Notre Dame University)

Presentation materials

There are no materials yet.