cs.HCSep 30, 2026

Characterizing Questioning Patterns and Student Engagement Through Contextual Analysis of Real-Time Classroom Interactions

Authors: Rohit Sharma, Pavani Ayinampudi, Aditya B. M. V., Jinal Gupta, Prakash Hegade, Sakshi Sharma, Meenakshi V, SRS Iyengar

Organizations: Indian Institute of Technology Ropar, Rupnagar, Punjab 140001, India · ANNAM.AI, Rupnagar, Punjab, India

Abstract

Real-time classroom polling is now routine, yet the data it produces is usually read narrowly, as a correctness score or a headcount. Such readings say little about what a poll is doing within a lecture or how it shapes engagement. This is particularly relevant for short-response formats such as True/False, where the same question format can be used to test recall, check comprehension, or direct students' attention to a deliberately misleading statement. This study asks whether a poll's answer and instructional function can be determined by reading it against its lecture transcript, what cognitive levels of Bloom's taxonomy and instructional-function clusters the corpus contains, and how student engagement relates to answering correctly. We analyse a naturalistic corpus of 47 live sessions over 39 days, comprising 604 poll questions and 340,668 responses from 2,807 learners, most items True/False, read against time-aligned lecture transcripts and attendance. Reading each poll in context proves essential: the answer to 89% of polls is locatable in the lecture, and a recurring attention-checking device is visible only through context. Questioning is overwhelmingly lower-order and falls into seven instructional functions, and a poll's response follows its function rather than its wording. Engagement is broad but concentrated, and the class majority answers correctly 88.5% of the time, though a small set of high-consensus yet incorrect answers cannot be detected by agreement alone. An independent survey of 579 students agrees on what the polls are and on their participation, but reveals a gap between perception and reality: students cannot judge their own correctness, and the polls they find hardest are not those they answer worst.

Figures & tables

Explore similar work

May 28, 2026cs.CY

Context-Aware Prediction of Student Quiz Performance with Multimodal Textbook Features

Educational platforms often predict student performance from prior interactions, but the assessment content itself also varies in linguistic and visual complexity. This paper studies whether lightweight content features extracted from CourseKata chapter-review questions improve prediction of end-of-chapter quiz scores beyond a student's average prior exercise performance. The study combines 2023 CourseKata student response data with chapter-level text features from review-question wording and image features from textbook visuals. Across 4,742 student-chapter observations from 562 class-student IDs, adding content features improves student-grouped five-fold quiz prediction performance by 9.1% relative to a prior-performance baseline. In leave-chapter-out validation, text features reduce prediction error relative to the baseline, while image-containing models have higher error. This paper suggests that a context-aware model adds useful signal about the text and visual features of questions to better predict student quiz performance compared with using past student performance alone.
Sep 21, 2026cs.CL

You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From

Questions scraped from the web are used across academia and industry as a proxy for what people want to know. Across QA training data, retrieval benchmarks, and content strategy, questions on a page are assumed to reflect human intent. We test this assumption at scale by extracting 13.4B question occurrences across 110 FineWeb snapshots (2013-2025), and report three findings. First, you can tell who is asking: provenance (the host/page of questions) leaves a signal in question form, and a logistic model can separate genuine user questions from templated/manufactured ones at AUC 0.725 via length and surrounding context rather than question type, though only 0.554 against commerce FAQ writing. Second, question frequency does not measure demand: the most-frequent questions are boilerplate/templated (over 70% of the top thousand), so occurrence counts measure how often a string was published and not how often it was asked. Third, over twelve years the genuine share of occurrences fell by 79% (42-56% after controlling for crawl composition), with question length and context decreasing. We present the first diachronic, occurrence-level measurement of web question provenance, and find the crawlable web's questions have shifted from being asked by humans toward manufactured for machines to read.
Apr 26, 2026cs.CL

Your Students Don't Use LLMs Like You Wish They Did

Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation of pedagogical alignment in student-AI dialogue. We validate our metrics through analysis of 12,650 messages across 500 conversations from four courses. Using our metrics, we identify a fundamental misalignment: educators design conversational tutors for sustained learning dialogue, but students mainly use them for answer-extraction. Deployment context is the strongest predictor of usage patterns, outweighing student preference or system design: when AI tools are optional, usage concentrates around deadlines; when integrated into course structure, students ask for solutions to verbatim assignment questions. Whole-dialogue evaluation misses these turn-by-turn patterns. Our metrics will enable researchers building educational dialogue systems to measure whether they are achieving their pedagogical goals.