Speaker
Description
This project trained Open-Source Large Language Models (LLMs) using large CEFR Datasets in order to have them generate materials appropriate for each of the six discrete CEFR levels. Training methods will be discussed, and steps for teachers to create their own systems will be introduced.
Abstract
Vocabulary is an essential element of language learning. Understanding 95-98% of words in a text is a key level to ensure meaningful comprehension in second language learning, and a gateway to learning more through context and inference. However, it is both difficult and time-consuming for teachers to create texts that meet these requirements for a majority of students, let alone all students in a class. Many teachers have tried to use Large Language Models (LLMs) to generate materials for the classroom. However, they have had mixed results with restricting the language used in these materials. This project has sought to train open source LLMs on Common European Framework of Reference for Languages (CEFR) datasets. The models were trained on vocabulary and grammar suitable for each of the six discrete CEFR levels. The materials generated by each model were compared and analyzed for appropriateness and accuracy. The presentation shares reproducible methods, and evaluation framework (coverage targets, syntactic complexity, and task suitability). The ultimate goal of the project is to offer tools for teachers to generate level-appropriate texts or for them to create their own tools using this project as a template.
Short summary
This project trained Open-Source Large Language Models (LLMs) using large CEFR Datasets in order to have them generate materials appropriate for each of the six discrete CEFR levels. Training methods will be discussed, and steps for teachers to create their own systems will be introduced.
Keywords
Vocabulary
AI
Materials
CEFR
| Scheduling preference | Anytime on Saturday |
|---|---|
| Title | Level Appropriate Materials Generation Using CEFR-Trained Models |