Towards Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions
Organizations: Carnegie Mellon University Pittsburgh, PA, USA · Renaissance Philanthropy New York, NY, USA · Vanderbilt University Nashville, TN, USA
Abstract
Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new affordances. An example of this is remote tutoring programs, where human tutors support students who use learning systems while video conferencing. Toward better platform-general modeling of learning, we introduce an AI-driven multimodal transcription system that processes screen-recording videos into unified screenplay-style transcripts containing audio dialogue and annotated learning log actions. We describe a planned method for temporally aligning AI-generated multimodal transcripts with MATHia learning logs and for identifying and classifying student learning processes to align with MATHia logs. Lastly, we highlight challenges and potential solutions in capturing learning processes in one system, offering initial steps towards generalizing log data across diverse systems.
Figures & tables
| Time | AI-generated multimodal transcript (Fig. 2 ) | Temporally-merged transcript (Fig. 3 ) |
|---|---|---|
| 13:21–13:27 | Tutor : I think it will be easy for me to explain. So, yeah, I think, okay. | Tutor : Your screen might stop sharing, because I think it would be easy for me to explain. |
| 13:28–14:11 | [view change event: The tutor’s whiteboard appears on the screen.] | |
| 14:12–14:30 | [drawing event: tool=pen, Tutor writes the equation ‘74 = + - 10’ on the whiteboard.] | |
| 14:33–14:39 | Tutor : Now we have a simple equation and we just need to find the value of . | Tutor: Now we have a simple equation and we just need to find the value of . |
| 14:41–14:48 | Tutor : Um, I think even you can write. So if you want to write on the whiteboard itself, you can feel free to go ahead and write as well. | Tutor : I think even you can write. So if you want to write on the whiteboard itself, you can be free to go ahead and write as well. |
| 15:44–15:47 | [drawing event: tool=pen, The student starts to write ‘84’.] |
| Event Type | Input/AI-Generated Multimodal Transcription | Output/Ground Truth (MATHia) |
|---|---|---|
| attempt | Visual detection of keyboard input, mouse drags, or system feedback while using MATHia containing “type”, “types”, “plots”, “drags”, “selects”, etc. (e.g., “Student types 30 into the input box” ; “text ‘13’ into the reflection line value box” ). Correct : “green checkmark appears”, “page advances”, “next screen loads”; Incorrect : “Try again modal appears”, “Error message appears”, “Orange (or red) input highlights”, “try again”, “incorrect”, etc. | Action = “Attempt” contains input text data (e.g., {value = x}) and correctness data (e.g., OK or ERROR). |
| hint_request | Visual detection of initial hint interaction while using MATHia containing keywords such as “hint” or “hints” (e.g., “Student clicks Hints button” ; “Hint pop-up appears, showing Hint 1 of 3” ; “The next button hint appears” ). Distinguished from hint_level_change by being the first event in the hint sequence. | Action = “Hint Request”, just-in-time “JIT”, or outcome = “INITIAL_HINT”. |
| hint_level_change | Visual detection of advancing hint states (e.g., “Student clicks Next button in hint pop-up” ; “Hint pop-up changes to show Hint 2 of 3” ). | Action = “Hint Level Change” or outcome = “HINT_LEVEL_CHANGE”. |
| done | Visual detection of problem-level completion indicated by “done” or “I’m done” (e.g., “Student clicks the I’m Done button” ). | Action = “Attempt” with input = “Done” or step_name = “Done”. |