cs.CLOct 5, 2026

Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

Authors: Jiulin Li, Ping Huang

Organizations: Beijing Liuyi Guanhua Technology Co., Ltd. · State Key Laboratory of General Artificial Intelligence, BIGAI

Abstract

Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

    Jul 27, 2026Zongyou Yang, Yinghan HouMulti-Dimensional EvaluationMath Problems

  2. ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions

    Jun 16, 2026Peixian Zhou, Yuxu Chen, Chaorui Zhang +3Reasoning BenchmarkReasoning Skills

  3. Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks

    Apr 10, 2025Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar +4Reasoning BenchmarkHuman-Annotated Benchmark