20–22 Nov 2026
The WINC Aichi
Asia/Tokyo timezone

Attenuation of Spoken Features in ASR/LLM-Mediated Output

22 Nov 2026, 12:00
1h
The WINC Aichi

The WINC Aichi

Poster Presentation (60-minutes) CALL: Computer Assisted Language Learning Room 1001

Speaker

Yuichi Ishikawa (Kokusai Junior College)

Description

Thirty Japanese junior college students produced monologues of up to one minute. Manual transcripts were compared with Whisper ASR output and GPT-4o model texts generated from the ASR transcripts. Disfluency, spoken grammar, and register features were examined across the three text types. Disfluencies largely disappeared in ASR and model texts, ellipsis decreased, and lexical density increased in AI output. The findings suggest that ASR/LLM mediation attenuates spoken-language features, raising pedagogical concerns for interpreting automated feedback.

Abstract section 1: Relevance

Spoken grammar differs systematically from written grammar, including features such as heads and tails, discourse markers, and ellipsis (Carter & McCarthy, 1995, 2006). Disfluency phenomena such as filled pauses and repetitions are integral to the speech production process (Levelt, 1989; Shriberg, 1994), and may also serve interactional functions in discourse (Clark & Fox Tree, 2002). Research on register variation has demonstrated systematic differences between spoken and written discourse (Biber et al., 1999; Halliday & Matthiessen, 2014). However, little is known about how ASR and subsequent AI-mediated reformulation represent spoken discourse and how linguistic features may change across these representations.

Abstract section 5: References

Biber, D., Johansson, S., Leech, G., Conrad, S., & Finegan, E. (1999). Longman grammar of spoken and written English. Longman.
Carter, R., & McCarthy, M. (1995). Grammar and the spoken language. Applied Linguistics, 16(2), 141–158. https://doi.org/10.1093/applin/16.2.141
Carter, R., & McCarthy, M. (2006). Cambridge grammar of English. Cambridge University Press.
Clark, H. H., & Fox Tree, J. E. (2002). Using uh and um in spontaneous speaking. Cognition, 84(1), 73–111. https://doi.org/10.1016/S0010-0277(02)00017-3
Halliday, M. A. K., & Matthiessen, C. M. I. M. (2014). Halliday’s introduction to functional grammar (4th ed.). Routledge.
Levelt, W. J. M. (1989). Speaking: From intention to articulation. MIT Press.
Shriberg, E. E. (1994). Preliminaries to a theory of speech disfluencies (Doctoral dissertation, University of California, Berkeley).

Abstract section 4: Outcomes/results

Disfluencies showed strong attenuation across representations. Friedman tests indicated significant differences (p < .001), and pairwise comparisons showed that filled pauses, repetitions, and self-repairs were greatly reduced in ASR output and largely eliminated in GPT-4o model texts. Among spoken-grammar features, ellipsis decreased significantly in model texts (p < .001), whereas heads and tails and discourse markers showed no significant differences. Register analysis revealed increased lexical density in GPT-4o output (p < .001). These findings suggest that AI-generated texts reshape learner speech toward more written-like forms, requiring careful interpretation in speaking instruction.

Abstract section 3: Content/method

Thirty students at a Japanese junior college completed up to one-minute monologue tasks. Speech was manually transcribed and compared with Whisper ASR output and GPT-4o model texts generated from the ASR transcripts. Analytical categories included filled pauses, repetitions, and self-repairs; heads/tails, ellipsis, and discourse markers; and lexical density, noun–verb ratio, and finite-verb clause length. Feature frequencies were normalized per 100 words. Differences were analyzed using Friedman tests with Wilcoxon pairwise comparisons.

Abstract section 2: Contribution/research questions

This study compares spoken-language features across three textual representations: manual transcripts, ASR output, and AI-generated model texts. The research questions are: (1) How do disfluencies differ across these representations? (2) How do spoken-grammar features differ across these representations? (3) To what extent do register-related measures vary among the three representations?

Title Attenuation of Spoken Features in ASR/LLM-Mediated Output
Teaching Context College and university education

Author

Yuichi Ishikawa (Kokusai Junior College)

Presentation materials

There are no materials yet.