Abstract
We study neural multi-label classification under severe label imbalance through sentence-level detection of the 19 refined Schwartz human values in 74k English news and manifesto sentences (ValueEval'24 corpus). Each sentence carries a roughly balanced moral-presence label and a 19-way value annotation. First, moral presence is learnable from single sentences: a DeBERTa-base classifier reaches positive-class F1≈0.73 at the default threshold, which calibration does not improve. Second, comparing direct multi-label detectors with presence-gated hierarchies under an 8 GB consumer-grade GPU budget, we find that gating does not improve over direct prediction, as gate recall becomes a bottleneck. Third, studying lightweight auxiliary signals and small ensembles, we isolate decision-threshold calibration as a decisive, often overlooked factor: a standard text-only baseline already matches the best official ValueEval'24 English run at the default threshold (macro-F1=0.282 vs. ≈0.28), and tuning the threshold on validation alone raises it to 0.315, most of our overall gain. Lightweight features do not survive a paired per-seed test; a soft-voting ensemble reaches our best macro-F1=0.332. Calibration is not architecture-specific: it reproduces on RoBERTa-base, where the gain is larger (+0.043). To our knowledge, this is the first systematic comparison of direct and presence-gated architectures, lightweight feature-augmented encoders, and instruction-tuned Large Language Models (LLMs) at sentence level; benchmarked 7-9B LLMs (zero-/few-shot and QLoRA) lag behind the supervised ensemble under the same budget. We provide empirical guidance for compute-efficient, value-aware NLP models.
Explore similar work
May 21, 2026cs.CL
Detecting Schwartz values in political texts is hard: cues are often implicit, and neighboring values differ by fine distinctions. Two remedies are widely assumed to help: more surrounding document text, and explicit moral knowledge. Knowledge-based retrieval has improved benchmarks elsewhere, but whether either transfers here is untested, because published systems vary context, knowledge, and model family at once. We separate these factors under matched conditions on the ValuesML/Touché ValueEval format. The input ranges from the target sentence to a local window to the full document. Retrieval is either absent or drawn from a curated moral knowledge base, injected by early, late, or cross-attention fusion. Supervised DeBERTa-v3 encoders are compared against zero-shot LLMs from 12B to 123B. More context is not uniformly better: full-document input improves the encoders by 2.5-3.8 macro-F1 points but does not consistently help the LLMs. Retrieved knowledge helps more reliably, improving every model family and context under early fusion. A control substituting random knowledge-base entries shows the families gain differently: LLMs from relevance, encoders from exposure to the value ontology. Neither larger encoders nor larger LLMs guarantee gains, and early fusion outperforms both trainable variants. Value-sensitive NLP should evaluate context, knowledge, and model family jointly.
Víctor Yeste, Paolo Rosso
Jan 10, 2026cs.CL
Building NLP systems for subjective tasks requires one to ensure their alignment to contrasting human values. We propose the MultiCalibrated Subjective Task Learner framework (MC-STL), which clusters annotations into identifiable human value clusters by three approaches (similarity of annotator rationales, expert-value taxonomies or rater's sociocultural descriptors) and calibrates predictions for each value cluster by learning cluster-specific embeddings. We demonstrate MC-STL on several subjective learning settings, including ordinal, binary, and preference learning predictions, and evaluate it on multiple datasets covering toxic chatbot conversations, offensive social media posts, and human preference alignment. The results show that MC-STL consistently outperforms the baselines that ignore the latent value structure of the annotations, delivering gains in discrimination, value-specific calibration, and disagreement-aware metrics.
Mohammed Fayiz Parappan, Ricardo Henao
Jul 22, 2026cs.CL
Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values. Our evaluation set contains 1,000 Russian situational texts, balanced across the ten values and independently labeled by two human annotators per item. We evaluate 21 instruction-tuned LLM runs under a fixed ranked-response protocol; 20 runs with reliable outputs form the semantic panel. Pooled Acc@1 is 0.683 and Acc@3 is 0.892, showing that models often locate the correct motivational region while ranking close alternatives unstably. Adjacent values account for 50.9% of semantic errors, compared with 24.4% under a checkpoint-specific null. Eight directed confusions recur across checkpoints and human-confirmed subsets. Several are strongly asymmetric, including Universalism to Benevolence, Tradition to Conformity, and Security to Power, whereas Stimulation-Hedonism forms a bidirectional boundary. Their severity is checkpoint-specific and can bias higher-order value profiles. The results motivate value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.
Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha +5