PII Detection

PII: Personally Identifiable Information

Latest papers 24

All topics
CardsList
  1. TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes

    Oct 5, 2026Jose D. Posada, Somalee Datta, Priya DesaiPII Detection

  2. The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond

    Sep 27, 2026Andreas Maier, Monica Hinrichs-Mayer, Franziska Weber +4Privacy AuditingPII Detection

  3. ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction

    Sep 24, 2026Sudha Priyadarshini, Mohamed Chahine GhanemPrivacy-Preserving Language ModelsLLM Agents

  4. MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

    Sep 9, 2026Stig Hellemans, Tom Stroobants, Elyne Scheurwegs +3Clinical NLPPII Detection

  5. Mind the Gap: Robustness Risks in PII Detection Systems

    Sep 3, 2026Adeel Zafar, Slawomir NowaczykDistribution Shift RobustnessPII Detection

  6. LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images

    Sep 2, 2026Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. YamshchikovPII DetectionData Leakage

  7. Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study

    Aug 4, 2026Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio +1Privacy AuditingPII Detection

  8. DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

    Jun 29, 2026Roland Roller, Vera Czehmann, Derya Erman +13PII Detection

  9. RedactionBench

    Jun 17, 2026Sean Brynjólfsson, Shashvat Jayakrishnan, Esha Sali +2LLM EvaluationContextual Integrity

  10. Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification

    Jun 16, 2026Haocheng Zhang, Zhuqian Zhou, Kirk Vanacore +2AI in EducationPrivacy-Preserving ML

  11. Detecting Sensitive Personal Information in Japanese Pre-Training Corpora for Large Language Models

    Jun 10, 2026Rei Minamoto, Yusuke Oda, Daisuke KawaharaLegal NLPPrivacy-Preserving Language Models

  12. Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation

    May 16, 2026Osama Zafar, Alexander Nemecek, Yiqian Zhang +5Data LeakageRAG Security

  13. The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

    May 8, 2026William Brach, Federico Torrielli, Stine Lyngsø Beltoft +3Privacy Leakage in Language ModelsMulti-Agent Systems

  14. SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification

    May 5, 2026Jose D. Posada, David Love, Somalee Datta +1HealthcareClinical Information Extraction

  15. Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering

    Apr 13, 2026Vladimir StantchevAI Agent SecuritySafety Filtering

  16. BodhiPromptShield: Pre-Inference Prompt Mediation for Surface-Form Privacy Propagation in LLM Agent Pipelines

    Date pendingBo Ma, Jinsong Wu, Weiqi YanData LeakagePrivacy-Preserving Language Models