cs.CLDec 25, 2025

Context-Aware Classification and Grading of Sensitive Information in Online Conversational Health Data

Authors: Yiwei Yan, Guanfeng Liu

Organizations: School of Computing, Macquarie University, Australia

Abstract

Online medical consultations contain sensitive health information whose privacy implications depend not only on the entities mentioned but also on how those entities are described in context. Existing classification and grading approaches often map health-information entities directly to predefined sensitivity levels, potentially overlooking whether a condition is confirmed, suspected, negated, hypothetical, or merely planned for investigation. In this study, we formulate sensitive-information grading in online medical dialogues as a context-aware evaluation task. We develop a standard-informed operational framework that incorporates assertion status, experiencer, test-result status, and information granularity. We further design a naturalistic evaluation setting together with contrastive cases that minimally alter negation, uncertainty, experiencer, or granularity, and compare large language models under mention-only and full-context conditions. The study aims to quantify the contribution of contextual information to sensitivity grading and to characterize safety-critical over- and under-grading errors. Our framework provides a reproducible basis for evaluating whether LLMs can distinguish sensitive entity mentions from contextually established sensitive disclosures.

Figures & tables

Explore similar work

CardsList
  1. Clinically Grounded Privacy Evaluation of Medical LMs

    Jun 8, 2026Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle +6Data LeakagePrivacy

  2. MIRA: A Bilingual Benchmark for Medical Information Response Audit

    May 27, 2026Mengyu Xu, Qiaoxin Yang, Qianqian Wang +3Multilingual Healthcare Q&AMultilingual Benchmark

  3. Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

    Jun 10, 2026Hongjian Zhou, Xinyu Zou, Jinge Wu +19Large Language Model EvaluationDiagnostic Benchmark