Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier
Organizations: Billennium S.A.
Abstract
We develop a multi-label Polish content-safety classifier by fine-tuning allegro/herbert-base-cased (124M) across five categories (hate, vulgarity, sexual content, crime, self-harm) using a Focal + R-Drop objective, and evaluate the resulting model against Bielik Guard (Sójka) on the shared out-of-distribution Gadzi Język benchmark. Both systems are given per-category threshold tuning on the same calibration split. Under that matched protocol our model holds a small but statistically significant lead in micro F1, while an apparent macro-F1 lead does not survive: it was an artifact of comparing a tuned model against an untuned one. We also report what that micro figure is worth. Because Gadzi Język is 97% crime-positive, a classifier that flags crime on every input and nothing else already scores 0.910 micro F1 on the same test split, so micro separates neither system from a degenerate strategy and macro is the column that does. Per-category and per-protocol figures are reported in Section 4. The residual out-of-distribution gap is one of calibration rather than discrimination. Ranking quality stays high while positive probabilities collapse, and per-category temperature scaling recovers the loss where Platt scaling and isotonic regression do not. That recovery turns out to be conditional on the calibration set containing safe text. Gadzi Język contains almost none, so thresholds fitted on it flag crime on every safe input, and a balanced refit buys a deployable operating point at the cost of adversarial recall. We report both operating points rather than only the flattering one. Two changes that are standard practice, per-class cost-sensitive weighting and mean pooling, each raise in-distribution macro F1 while lowering the out-of-distribution figure, which indicates that robustness has to be selected for directly rather than inherited from in-distribution accuracy.
Figures & tables
| Component | Configuration |
|---|---|
| Encoder | allegro/herbert-base-cased (124M) |
| Head | Linear(768 → 5), dropout 0.22 |
| Loss | Focal (γ = 1.5) + R-Drop (α = 0.20) |
| Optimizer | AdamW (lr = 9.4e-5, wd = 5e-4) |
| Schedule | Cosine, 8% warmup |
| Frozen layers | 4 / 12 encoder layers |
| Encoder | OOD macro | OOD macro (augmented) |
|---|---|---|
| HerBERT ( allegro/herbert-base-cased ) | 0.601 | 0.611 |
| MMLW-RoBERTa (Sójka’s backbone, sweep-tuned) | 0.275 | - |
| Polish RoBERTa | 0.282 | 0.269 |
| XLM-RoBERTa | 0.158 | 0.211 |
| Name | Configuration |
|---|---|
| Baszta 0.1 | Clean v1 pool, no augmentation |
| Baszta 0.2 | Balanced pool, online augmentation |
| Baszta 0.3 | Balanced pool, aggressive augmentation |
| Baszta 0.4 | Adds style-matched synthetic crime data |
| Baszta 0.5 | Adds synthetic crime , sex and self-harm data |
| Baszta 1.0 | Sweep recipe (§3.4), low-γ focal, v2 synthetic data, per-category temperature scaling |
| Category | Train support |
|---|---|
| hate | ~3,146 |
| vulgar | ~4,037 |
| sex | ~4,500 (incl. 501 synthetic) |
| crime | ~4,250 (incl. 286 synthetic) |
| self-harm | ~1,246 (incl. 507 synthetic) |
| Category | Definition used |
|---|---|
| hate | Hostility, dehumanisation or incitement directed at a person or group, typically on a protected characteristic |
| vulgar | Profanity, obscenity or crude language, with no requirement that it be aimed at anyone |
| sex | Sexual content, including solicitation and sexualisation of minors, which also carries crime |
| crime | Requests for, or instruction in, illegal activity, including violence, weapons, drugs and fraud |
| self-harm | Content describing, encouraging or seeking means of self-injury or suicide |
| Model | Hate | Vulgar | Sex | Crime | Self-harm | Macro F1 |
|---|---|---|---|---|---|---|
| Baszta 0.2 | 0.743 | 0.969 | 0.985 | 0.965 | 0.835 | 0.900 |
| Baszta 0.1 | 0.733 | 0.967 | 0.978 | 0.960 | 0.852 | 0.898 |
| Baszta 0.3 | 0.745 | 0.960 | 0.981 | 0.957 | 0.838 | 0.896 |
| Sójka 0.1B† | 0.628 | 0.742 | 0.889 | 0.707 | 0.886 | 0.770 |
| Model | Hate | Vulgar | Sex | Crime | Self-harm | Macro | Micro |
|---|---|---|---|---|---|---|---|
| Baszta 1.0 | 0.575 | 0.571 | 0.606 | 0.965 | 0.820 | 0.707 | 0.927 |
| Baszta 0.4 | 0.514 | 0.333 | 0.621 | 0.985 | 0.800 | 0.651 | - |
| Baszta 0.5 | 0.435 | 0.333 | 0.703 | 0.985 | 0.750 | 0.641 | - |
| Baszta 0.2 | 0.419 | 0.333 | 0.500 | 0.985 | 0.778 | 0.603 | 0.918 |
| Sweep top-1 (ablation baseline, §4.6) | 0.489 | 0.400 | 0.686 | 0.898 | 0.693 | 0.633 | 0.828 |
| Baszta 0.1 (clean, no synthetic) | 0.410 | 0.400 | 0.650 | 0.675 | 0.708 | 0.569 | 0.651 |
| Category | OOD AUC | OOD pos-prob mean |
|---|---|---|
| hate | 0.86 | 0.55 |
| vulgar | 0.49 | 0.13 |
| sex | 0.93 | 0.21 |
| crime | 0.72 | 0.19 |
| self-harm | 0.94 | 0.27 |
| Calibration | Gadzi Macro F1 |
|---|---|
| Raw ( ) | 0.178 |
| Platt scaling | 0.218 |
| Isotonic regression | 0.216 |
| Oracle per-category thresholds | 0.603 |
| Protocol (172/348 split) | Test Macro F1 |
|---|---|
| Raw ( ) | 0.493 |
| + per-category threshold | 0.603 |
| + temperature + threshold | 0.699 |
| Sójka reference | 0.619 |
| Metric (Test split, N = 348) | Estimate | 95% CI |
|---|---|---|
| Macro F1 (5 categories) | 0.699 | [0.493, 0.797] |
| Macro F1 (4 categories, excl. vulgar) | 0.674 | [0.569, 0.763] |
| Micro F1 | 0.927 | [0.903, 0.949] |
| Operating point | Balanced macro | Balanced micro | Safe-text FPR | Gadzi macro | Gadzi micro |
|---|---|---|---|---|---|
| Gadzi-fit (T_crime=4.3, t_crime=0.025) | 0.742 | 0.541 | 1.000 | 0.700 | 0.932 |
| Balanced re-fit | 0.952 | 0.959 | 0.055 | 0.496 | 0.778 |
| Metric (both systems tuned on same cal split) | Ours | Sójka | Diff (ours−Sójka) | 95% CI | P(ours not better) |
|---|---|---|---|---|---|
| Micro F1 | 0.929 | 0.903 | +0.026 | [+0.004, +0.049] | 0.011 |
| Macro F1 (5-cat) | 0.712 | 0.782 | −0.070 | [−0.263, +0.046] | 0.875 |
| Always-crime baseline (micro / macro) | 0.910 / 0.197 | - | - | - | - |
| Benchmark | Rows | In training | Status |
|---|---|---|---|
| PL-Guard (NASK) [ 12 ] | 900 | 899 | excluded - memorised |
| PL-Guard-adv (NASK) | 900 | 765 | excluded - memorised |
| KLEJ CBD (test) [ 25 ] | 1,000 | 149 | 149 dropped → 851 held-out |
| BAN-PL_1 [ 9 ] | 24,000 | 461 | 461 dropped → 23,539 held-out |
| HateCheck-PL [ 24 ] | 3,815 | 0 | fully held-out |
| PolyGuardPrompts-PL [ 13 ] | 1,725 | 0 | fully held-out |
| Benchmark (held-out) | N | Task / head | F1 (95% CI) | Safe/neg FPR |
|---|---|---|---|---|
| KLEJ CBD (clean) | 851 | cyberbully / hate-head | 0.671 [0.606, 0.735] | 0.102 |
| BAN-PL_1 (clean) | 23,539 | harmful / any-head | 0.667 [0.661, 0.673] | 0.855 |
| PolyGuardPrompts-PL | 1,725 | harmful / any-head | 0.562 [0.531, 0.591] | 0.357 |
| HateCheck-PL | 3,815 | hate / hate-head | 0.644 (diagnostic) | 0.334 |
| Failure mode (hate head) | Worst functionalities (accuracy) |
|---|---|
| False positives (fires on non-hateful) | counter-speech quoting/referencing hate (0.48–0.61), positive/neutral identity mentions ( ident_pos_nh 0.57, target_group_nh 0.51) |
| False negatives (misses hateful) | slur-only hate ( slur_h 0.30), implicit/emotive derogation (0.37), space-obfuscated hate ( spell_space_add_h 0.41) |
| Benchmark (N) | Head | Baszta 1.0 | Sójka (0.5) | Sójka (matched) | FPR ours / theirs |
|---|---|---|---|---|---|
| KLEJ CBD (851) | hate | 0.671 | 0.088 | 0.438 | 0.102 / 0.087 |
| BAN-PL (23,539) | any | 0.667 | 0.694 | 0.754 | 0.855 / 0.230 |
| PolyGuard-PL (1,725) | any | 0.562 | 0.290 | 0.475 | 0.357 / 0.197 |
| HateCheck-PL (3,815) | hate | 0.644 | 0.527 | 0.767 | 0.334 / 0.661 |
| System | KLEJ CBD | BAN-PL | PolyGuard-PL | HateCheck-PL | Gadzi detect. |
|---|---|---|---|---|---|
| Baszta 1.0 (124M) | 0.669 / 0.109 | 0.667 / 0.855 | 0.562 / 0.357 | 0.710 / 0.578 | 0.906 |
| Sójka 0.1B v1.1 (124M) | 0.433 / 0.131 | 0.754 / 0.230 | 0.475 / 0.197 | 0.784 / 0.745 | 0.675 |
| HerBERT-PL-Guard (124M) | 0.459 / 0.266 | 0.762 / 0.354 | 0.816 / 0.070 | 0.786 / 0.593 | 0.992 |
| Qwen3Guard-Gen (0.6B) | 0.262 / 0.028 | 0.263 / 0.078 | 0.766 / 0.085 | 0.636 / 0.384 | 0.985 |
| Model (Gadzi, macro over 5 cats) | ECE | Adaptive-ECE | MCE | Brier |
|---|---|---|---|---|
| Baszta 1.0 raw | 0.092 | 0.096 | 0.691 | 0.075 |
| Baszta 1.0 + temperature | 0.107 | 0.111 | 0.753 | 0.068 |
| Sójka 0.1B | 0.136 | 0.139 | 0.619 | 0.114 |
| Attack (Gadzi, balanced op-point) | No defence: macro / flip | Deobfuscated: macro / flip |
|---|---|---|
| homoglyph | 0.477 / 0.055 | 0.496 / 0.000 |
| leetspeak | 0.415 / 0.071 | 0.471 / 0.024 |
| space-insertion | 0.409 / 0.100 | 0.485 / 0.024 |
| char-repetition | 0.296 / 0.146 | 0.476 / 0.030 |
| diacritics-strip | 0.486 / 0.026 | 0.486 / 0.026 |
| combined-heavy | 0.343 / 0.124 | 0.426 / 0.078 |
| Baseline | Val macro / micro | Gadzi macro / micro |
|---|---|---|
| TF-IDF + LogReg | 0.868 / 0.883 | 0.462 / 0.642 |
| TF-IDF + Compl-NB | 0.717 / 0.763 | 0.393 / 0.690 |
| HerBERT (Baszta 1.0) | 0.9111 / 0.9585 (val) | 0.496 / 0.778 (balanced op-point) |
| Model | Val macro | Gadzi macro | Gadzi micro | OOD ECE | OOD Brier |
|---|---|---|---|---|---|
| Baszta 1.0 (no alpha) | 0.9111 | 0.699 | 0.927 | 0.092 | 0.075 |
| Baszta 1.0-α (inverse-freq alpha) | 0.9136 | 0.656 | 0.927 | 0.133 | 0.093 |
| Recipe | Val macro | OOD macro | OOD micro vs. Sójka | OOD ECE / Brier |
|---|---|---|---|---|
| Baszta 1.0 ([CLS], no alpha) | 0.9111 | 0.699 | 0.929 (p = 0.011 ✓) | 0.092 / 0.075 |
| Baszta 1.0-α (per-class alpha) | 0.9136 | 0.656 | 0.927 (p = 0.017 ✓) | 0.133 / 0.093 |
| Baszta 1.0-mp (mean pooling) | 0.9138 | 0.606 | 0.913 (p = 0.218 ✗) | 0.109 / 0.093 |
| Model / ensemble | OOD macro | OOD micro (vs. Sójka) | OOD Brier |
|---|---|---|---|
| Baszta 1.0 (single, best macro) | 0.712 | 0.929 (p = 0.011) | 0.075 |
| Baszta 1.0 + Baszta 1.0-α + Baszta 1.0-mp (mean-prob) | 0.632 | 0.932 (p = 0.005) | 0.071 |
| Baszta 1.0 + Baszta 1.0-α (mean-prob) | 0.660 | 0.933 (p = 0.004) | 0.070 |
| Clean held-out benchmark | Baszta 1.0 | Baszta 1.0-α | Baszta 1.0+Baszta 1.0-α ensemble |
|---|---|---|---|
| HateCheck-PL hate F1 | 0.644 | 0.613 | 0.645 |
| KLEJ CBD hate F1 | 0.671 | 0.687 | 0.671 |
| BAN-PL any-head F1 | 0.667 | 0.663 | 0.667 |
| PolyGuard-PL binary F1 | 0.562 | 0.511 | 0.517 |
| Dimension | Ours (Baszta 1.0+temp) | Sójka 0.1B | Δ / note |
|---|---|---|---|
| Encoder | HerBERT (124M) | MMLW-RoBERTa (100M) | similar |
| Loss | Focal (γ=1.5) + R-Drop | BCE | - |
| Training data | 26,248 (+synthetic) | ~6.9K | larger |
| Compute | 1× T4 and 2× GH200, FP16 | A100 cluster | constrained |
| Gadzi micro F1 | 0.927 [0.903, 0.949] | 0.582 | Sójka untuned, see below |
| Gadzi macro (cal split) | 0.699 [0.493, 0.797] | 0.619 | directional (81% boot.) |