Abstract This paper addresses a critical safety gap in the use Automated Verbal Response Scoring (AVRS). We present a novel hybrid framework for troubled student detection that combines a text classifier, trained to detect responses based on their content, and an audio classifier, trained to detect responses using prosodic markers. This approach overcomes key limitations of traditional AVRS systems by considering both content and prosody of responses, achieving enhanced performance in identifying potentially concerning responses. This system can expedite the review process by humans, which can be life-saving particularly when timely intervention may be crucial.
Explore similar work Aug 4, 2026 · Anil Sharma, Sarthak Ahuja, Mayank Gautam +1 Early Warning Direction-Of-Arrival Estimation
Jun 10, 2026 · Michael L. Chrzan, Meghavarshini Krishnaswamy, Robert Gibboni +3 Classrooms Semantic Anchor
Jul 16, 2026 · Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir +3 Speaker Verification Performance Large Audio Language Models
Aug 4, 2026 · cs.SD J/K move · Enter open · S save
Anil Sharma, Sarthak Ahuja, Mayank Gautam, Sanjit Kaul
IIIT-Delhi, India
We investigate an unobtrusive and
24 × 7 24\times7 24 × 7 human distress detection and signaling system, Always Alert, that requires the smartphone, and not its human owner, to be on alert. The system leverages the microphone sensor, at least one of which is available on every phone, and assumes the availability of a data network. We propose a novel two-stage supervised learning framework, using support vector machines (SVMs), that executes on a user's smartphone and monitors natural vocal expressions of fear---screaming and crying in our study---when a human being is in harm's way. The challenge is to achieve a high distress detection rate while ensuring that the false alarm rate is a manageable overhead, while a typical smartphone user goes about living life as usual. We train the learning framework with carefully selected audio fingerprints of distress and of varied environmental contexts. The audio is used to tune the learning framework to obtain a desirable distress detection rate and false alarm rate (FAR). The ability of the proposed framework to detect distress in rather challenging audio environments is demonstrated. Exploiting the time contiguous nature of false alarms further allows us to reduce the FAR. We show the feasibility of using our framework anytime and anywhere by testing it over many hours of audio fingerprints recorded by volunteers on their smartphones, as they went about their daily routines. We are able to achieve high distress detection rates at an average overhead that is equivalent to about 1 facebook post every 3 to 4 hours.