cs.CLSep 28, 2026

Nürnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access

Authors: Philipp Steigerwald, Eric Rudolph, Jens Albrecht

Organizations: Technische Hochschule Nürnberg Georg Simon Ohm

Abstract

We describe the Nürnberg NLP system for ChildSafeAds 2026. The shared task asks what a monitoring system for commercial content in child-facing YouTube videos can achieve at a given level of data access. We answer with per-subtask ensembles of nine voters, organised into three branches that differ in backbone, adaptation method and class scope. Selection rests on channel-disjoint cross-validation, with the development set as a transfer check. The system wins two of the three subtasks. Its product-category score (ST2, 0.8243) and its compliance-flag score (ST3, 0.6530) are the best of the 22 final entries, and it places third on the task mean (0.7079). We further compare four access levels and report the cost at test-set scale.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 2, 2026cs.CY

Auditing Engagement Incentives in the Kidfluencer Ecosystem: A Multimodal Weak Supervision Approach

The rise of `kidfluencers' on YouTube has raised ethical concerns about child digital labor and exploitation. While emerging legislation attempts to regulate this ecosystem, empirical evidence linking exploitation to engagement remains scarce, given the difficulty of operationalizing exploitation at scale. This study presents a multimodal AI audit of 5,051 videos across 79 kidfluencer channels, using weak supervision to detect exploitation signals without large-scale manual labels. We aggregate noisy labeling functions -- including LLM-based classification of titles and GPT-4 Vision analysis of thumbnails and descriptions across six literature-grounded dimensions -- to assign a probabilistic exploitation score to each video. A multi-annotator validation study (N=107) shows strong agreement with human judgment (macro-average F1 =0.911= 0.911) and high sensitivity for overall exploitation risk (recall =0.960= 0.960, F1 =0.793= 0.793). Our findings reveal a significant engagement premium for performative labor, emotional bait, and privacy violations. Exploitation scores correlate with view counts (Spearman ρ=0.229ρ= 0.229, p<10−50p < 10^{-50}), and mixed-effects regression controlling for channel-level variation shows that a one-unit increase in exploitation score yields a 4.4×4.4\times increase in views (p<0.001p < 0.001). Within-channel analyses indicate median view boosts of +65.6%+65.6\% for emotional bait and +56.0%+56.0\% for performative content (FDR-corrected p<0.001p<0.001), with effects holding in same-year robustness checks (p=0.030p=0.030). Explicit commercial content (product placement), by contrast, shows no premium (−3.8%-3.8\%, n.s.), suggesting the platform rewards commodification of the child's identity and labor over traditional advertising. These findings challenge policy frameworks focused solely on financial trusts, showing that engagement is systematically tied to the intensive, performative labor of children.
Jun 29, 2026cs.LG

CAREBench: A Child-Safety Risk Benchmark for Language Models

How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earlier: in model assistance that helps adults manipulate, impersonate, profile, or isolate minors, and in model responses that deepen children's emotional dependence on AI systems rather than redirecting them toward human support. We introduce CAREBench (Child AI Risk Evaluation), a benchmark to assess such upstream child-safety risks in language models. CAREBench contains 500 prompts spanning twelve risk categories, including grooming and relationship engineering, deception and impersonation, surveillance and privacy, sextortion and sexual abuse, AI anthropomorphization, emotional dependency, and mental illness sensitivity. Developed with response annotations from parents and clinicians, the benchmark excludes explicit abuse material and imagery; instead, it evaluates whether models recognize, refuse, de-escalate, or redirect risky interactions before harm becomes overt. Evaluating seven frontier models on our benchmark, we find failure rates ranging from 2% to 58%, with failure patterns that vary across risk categories. CAREBench provides a responsibly scoped evaluation for LLM developers to identify and close gaps in child safety policies.
Aug 5, 2026cs.CL

Social Pressure Breaks Majority Voting in LLM Safety Panels

Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the same misleading context before voting. We study this risk in a controlled two-round experiment. Each model first judges an item alone, then judges it again after six simulated peers either assert the wrong label or abstain. We combine the final judgments by majority vote. Across six open-weight LLMs and six datasets, we find that the wrong-label peer message raises the average reviewer false-alarm rate from 56.5% under silent peers to 87.5%, and majority voting raises the panel false-alarm rate to 100%. Without an asserted label, the same panel outperforms its average member. The effect is strongly asymmetric: reviewers follow pushes toward "unsafe" far more than pushes toward "safe" (about 75% versus 17%), so the panel's false-alarm rate rises sharply while its harmful-miss rate changes little. The proprietary-model probe shows substantial variation across models. These results identify susceptibility to shared social cues as a failure mode of safety panels and provide a simple pre-deployment diagnostic.