cs.CLJul 28, 2026

Shieldstral

Authors: Antonia CalviAvinash SooriyarachchiGiada PistilliGuillaume LampleMaarten BuylMaximilian AugustinMaximilian MüllerPierre Stock+268 more

Abstract

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7×\times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

Explore similar work

CardsList
  1. GLiGuard: Schema-Conditioned Classification for LLM Safeguard

    May 8, 2026Urchade Zaratiana, Mary Newhauser, George Hurn-Maloney +1Large Language Model SafetyGuardrail