20–22 Nov 2026
The WINC Aichi
Asia/Tokyo timezone

Improving Automated Transcription of L2 Learner Texts

22 Nov 2026, 14:20
30m
The WINC Aichi

The WINC Aichi

Research-oriented Presentation (30-minutes) VOCAB: Vocabulary Room 1209

Speakers

Gavin Brooks (Kyoto University of Foreign Studies) Christopher HOLLIS (Tottori University (Tottori JALT))

Description

This presentation addresses the problem of creating accurate spoken L2 learner corpora by introducing a hybrid transcription approach combining top-down LLM-mediated ASR (Whisper) with bottom-up wav2vec. Findings from an analysis of word error rate in 200 L2 speech samples show that ASR accuracy varies by task and proficiency. The presenters will demonstrate a prototype tool that aligns ASR and wav2vec transcripts to enable efficient human review and to export a learner corpus for further analysis.

Abstract section 3: Content/method

The presenters report findings from an analysis of word error rate (WER) in 200 samples of L2-generated speech, including 100 presentations and 100 discussions from learners at CEFR A2 to B2. ASR transcripts were combined with wav2vec output using a modified Needleman-Wunsch alignment algorithm. The resulting transcript was compared against professionally transcribed and manually cleaned versions. These findings were then used to identify problematic phonemes in L2 learner speech to further improve the wav2vec models.

Abstract section 5: References

Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449–12460.

Brezina, V., Gablasova, D., & Lenko-Szymanska, A. (2022). Learner corpus research: Past, present and future. In V. Brezina, D. Gablasova, & A. Lenko-Szymanska (Eds.), The Routledge handbook of corpora and English language teaching and learning (pp. 107–123). Routledge.

Kyle, K., & Eguchi, M. (2024). Assessing the accuracy of automatic speech recognition for L2 English. Language Learning & Technology, 28(1), 1–23.

McGuire, M., & Larson-Hall, J. (2025). Assessing Whisper automatic speech recognition and WER scoring for elicited imitation: Steps toward automation. Research Methods in Applied Linguistics, 4(1), 100197. https://doi.org/10.1016/j.rmal.2025.100197

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2022). Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356.

Wang, Y., Shen, J., & Jia, Y. (2021). Non-native English speaker speech recognition challenges. Proceedings of Interspeech 2021, 3820–3824.

Abstract section 4: Outcomes/results

Results indicate that while ASR performs comparably to professional transcription for single-speaker presentations, it produces a significantly higher WER for multi-speaker discussions due to overlapping speech, varying accents, and background noise. Higher-proficiency participants' speech was transcribed with greater accuracy. Based on these findings and insights from prior literature, the presenter introduces a prototype tool that aligns ASR and word2vec transcripts at the word level, highlights discrepancies for manual review, and exports time-stamped gold-standard transcripts suitable for corpus analysis in tools such as PRAAT. We will also provide recommendations for recording and using ASR in L2 research contexts.

Abstract section 2: Contribution/research questions

This presentation investigates a hybrid transcription approach that combines top-down LLM-mediated ASR (Whisper) with bottom-up wav2vec for learner corpus development. Drawing on the findings from previous projects that have examined ASR accuracy in L2 speech (McGuire & Larson-Hall, 2025), the presenters investigate two research questions:
1. How accurate are these approaches compared to a human transcriber in capturing vocabulary used by learners?
2. Can the system's performance be improved by training wav2vec on learner-generated speech?

Abstract section 1: Relevance

Recent studies have explored the use of Automated Speech Recognition (ASR) models like Whisper for transcribing L1 speech (Baevski et al., 2020; Radford et al., 2022), but there are challenges with applying ASR to L2 English learners due to pronunciation errors, disfluencies, and atypical grammatical constructions (Kyle & Eguchi, 2024; Wang et al., 2021). Because of this, there is a need for better tools that integrate ASR output with human oversight to create reliable transcripts, as a means of building spoken corpora to facilitate research on learner vocabulary usage across different domains and language contexts (Brezina et al., 2022).

Title Improving Automated Transcription of L2 Learner Texts
Teaching Context College and university education

Authors

Gavin Brooks (Kyoto University of Foreign Studies) Christopher HOLLIS (Tottori University (Tottori JALT))

Presentation materials

There are no materials yet.