Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi
Organizations: Beijing Liuyi Guanhua Technology Co., Ltd. · State Key Laboratory of General Artificial Intelligence, BIGAI
Abstract
Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.
Figures & tables
| System | Overall | 95% CI | Theory | Case | Invalid |
|---|---|---|---|---|---|
| Kimi-K3 | 75.56 | 73.84–77.21 | 86.04 | 60.89 | 4.98 |
| GLM-5.3 | 75.08 | 73.34–76.74 | 84.59 | 61.75 | 0.24 |
| Qwen3.8-Max | 74.84 | 73.10–76.50 | 84.59 | 61.18 | 5.98 |
| DS V4.1 Flash | 72.43 | 70.64–74.15 | 83.56 | 56.84 | 0.00 |
| DS V4 Pro | 72.35 | 70.56–74.07 | 84.66 | 55.11 | 0.00 |
| MiniMax-M3 | 63.20 | 61.29–65.07 | 71.25 | 51.93 | 1.08 |
| System | Fact ( ) | Theory ( ) | Gap (pp) |
|---|---|---|---|
| Kimi-K3 | 55.60 | 65.16 | +9.55 |
| GLM-5.3 | 56.03 | 66.38 | +10.34 |
| Qwen3.8-Max | 56.47 | 64.98 | +8.52 |
| DS V4.1 Flash | 51.29 | 61.32 | +10.03 |
| DS V4 Pro | 45.91 | 62.54 | +16.64 |
| MiniMax-M3 | 45.04 | 57.49 | +12.45 |
| System | Subset | Native | Disabled | Holm | |
|---|---|---|---|---|---|
| Flash | Overall | 76.97 | 75.53 | +1.43 | |
| Flash | Theory | 84.00 | 77.47 | +6.53 | |
| Flash | Case | 69.93 | 73.60 | -3.67 | |
| Pro | Overall | 76.83 | 70.10 | +6.73 | |
| Pro | Theory | 85.07 | 72.87 | +12.20 | |
| Pro | Case | 68.60 | 67.33 | +1.27 |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Statistic | Original | Final |
|---|---|---|
| Total items | 3,000 | 2,492 |
| Theory items | 1,500 | 1,454 |
| Case items | 1,500 | 1,038 |
| Theory categories | 14 | 14 |
| Case categories | 11 | 11 |
| Options per item | 4 | 4 |
| ID | Theory category | ID | Case category | ||
|---|---|---|---|---|---|
| T01 | Five Elements | 104 | C01 | Career | 96 |
| T02 | Stems and Branches | 104 | C02 | Education | 96 |
| T03 | Ten Gods | 104 | C03 | Wealth | 96 |
| T04 | Ten Gods Relations | 104 | C04 | Romance | 96 |
| T05 | Combination and Clash | 104 | C05 | Health | 96 |
| T06 | Strength and Weakness | 104 | C06 | Personality | 96 |
| Metric | Original | Final |
|---|---|---|
| Items | 3,000 | 2,492 |
| Mean empirical difficulty | 0.2326 | 0.2776 |
| Mean discrimination | 0.2182 | 0.2574 |
| Six-model all-correct | 1,595 | 1,118 |
| Six-model all-wrong | 205 | 204 |
| Holm-significant pairs | 11/15 | 11/15 |
| Item band | Removed |
|---|---|
| Six of six correct | 477 |
| Five of six correct | 30 |
| Two to four correct | 0 |
| One of six correct | 0 |
| Zero of six correct | 1 |
| Total | 508 |
| Model | Original | Final | Change |
|---|---|---|---|
| Kimi-K3 | 79.50 | 75.56 | -3.94 |
| GLM-5.3 | 79.13 | 75.08 | -4.05 |
| Qwen3.8-Max | 78.97 | 74.84 | -4.13 |
| DS V4.1 Flash | 76.97 | 72.43 | -4.54 |
| DS V4 Pro | 76.83 | 72.35 | -4.48 |
| MiniMax-M3 | 69.07 | 63.20 | -5.87 |
| Run family | Models | Reasoning | Source |
|---|---|---|---|
| V1 native | Six listed models | Native | Mixed routes |
| V2 Flash | DS V4.1 Flash | Disabled | DeepSeek official |
| V2 Pro | DS V4 Pro | Disabled | DeepSeek official |
| Endpoint configuration | Observed route(s) | Mode/max | UTC span | Raw | Set | Canonical | SHA prefix |
|---|---|---|---|---|---|---|---|
| DS V4 Pro (native) | DeepSeek official (3004) | native; 65536 | 13:32–16:21 | 3004 | final | 2492 | aa380761e51f |
| DS V4.1 Flash (native) | DeepSeek official (3002) | native; 32768/65536 | 09:07–10:02 | 3002 | final | 2492 | 2225eb6179df |
| MiniMax-M3 (native) | SCNet (1807); early/unknown (1340) | native; 32768/65536 | 07:45–13:24 | 3147 | final | 2492 | 981c54a982f4 |
| Kimi-K3 (native) | SCNet (1816); micuapi (1424); early/unknown (919) | native; 32768/65536 | 07:46–17:00 | 4159 | final | 2492 | c8b94ded91c9 |
| Qwen3.8-Max (native) | SCNet (1636); micuapi (1589); early/unknown (883) | native; 32768/65536 | 07:45–17:35 | 4108 | final | 2492 | 248c1ba7d807 |
| GLM-5.3 (native) | SCNet (1489); micuapi (1706); early/unknown (913) | native; 32768/65536 | 07:45–14:39 | 4108 | final | 2492 | ce25699f48ab |
| System | Micro | Task macro | Category macro |
|---|---|---|---|
| Kimi-K3 | 75.56 | 73.46 | 73.69 |
| GLM-5.3 | 75.08 | 73.17 | 73.38 |
| Qwen3.8-Max | 74.84 | 72.89 | 73.13 |
| DS V4.1 Flash | 72.43 | 70.20 | 70.41 |
| DS V4 Pro | 72.35 | 69.88 | 70.15 |
| MiniMax-M3 | 63.20 | 61.59 | 61.82 |
| Endpoint alias | Overall | 95% CI | Theory | Case |
|---|---|---|---|---|
| Kimi-K3 | 79.50 | 78.02–80.91 | 86.40 | 72.60 |
| GLM-5.3 | 79.13 | 77.64–80.55 | 85.00 | 73.27 |
| Qwen3.8-Max | 78.97 | 77.47–80.39 | 85.00 | 72.93 |
| DS V4.1 Flash | 76.97 | 75.43–78.44 | 84.00 | 69.93 |
| DS V4 Pro | 76.83 | 75.29–78.31 | 85.07 | 68.60 |
| MiniMax-M3 | 69.07 | 67.39–70.70 | 72.07 | 66.07 |
| Model | Theory | Case | Gap |
|---|---|---|---|
| Kimi-K3 | 86.04 | 60.89 | 25.15 |
| GLM-5.3 | 84.59 | 61.75 | 22.84 |
| Qwen3.8-Max | 84.59 | 61.18 | 23.42 |
| DS V4.1 Flash | 83.56 | 56.84 | 26.72 |
| DS V4 Pro | 84.66 | 55.11 | 29.56 |
| MiniMax-M3 | 71.25 | 51.93 | 19.32 |
| Model | Primary | Excl. invalid | Change |
|---|---|---|---|
| Kimi-K3 | 75.56 | 79.52 | +3.96 |
| GLM-5.3 | 75.08 | 75.26 | +0.18 |
| Qwen3.8-Max | 74.84 | 79.60 | +4.76 |
| DS V4.1 Flash | 72.43 | 72.43 | 0.00 |
| DS V4 Pro | 72.35 | 72.35 | 0.00 |
| MiniMax-M3 | 63.20 | 63.89 | +0.69 |
| Task | Category | Difficulty | |
|---|---|---|---|
| Theory | Shensha Basics | 104 | 0.240 |
| Theory | Pattern Basics | 104 | 0.220 |
| Theory | Strength and Weakness | 104 | 0.210 |
| Theory | Nayin | 104 | 0.114 |
| Theory | Twelve Stages | 104 | 0.109 |
| Case | Career | 96 | 0.630 |
| Release | Task | Cue (%) | Charts | Max reuse | |
|---|---|---|---|---|---|
| Original | Theory | 1,500 | 284 (18.93) | – | – |
| Original | Case | 1,500 | 481 (32.07) | 274 | 33 |
| Original | Total | 3,000 | 765 (25.50) | – | – |
| Final | Theory | 1,454 | 278 (19.12) | – | – |
| Final | Case | 1,038 | 286 (27.55) | 249 | 25 |
| Final | Total | 2,492 | 564 (22.63) | – | – |
| Model | Subset | Native | Disabled | |
|---|---|---|---|---|
| Flash | Overall | 72.43 | 71.79 | +0.64 |
| Flash | Theory | 83.56 | 77.03 | +6.53 |
| Flash | Case | 56.84 | 64.45 | -7.61 |
| Pro | Overall | 72.35 | 66.29 | +6.06 |
| Pro | Theory | 84.66 | 72.49 | +12.17 |
| Pro | Case | 55.11 | 57.61 | -2.50 |
| Alias | Tok./ | pp/1k | Lat. med. N/D | Ratio | |
|---|---|---|---|---|---|
| Flash | 43 | 353,933 | 0.283 | 10.68/1.43 | 7.44 |
| Pro | 202 | 71,180 | 1.405 | 20.95/2.07 | 10.13 |
| Model | Invalid | Invalid (%) |
|---|---|---|
| Qwen3.8-Max | 149 | 5.98 |
| Kimi-K3 | 124 | 4.98 |
| MiniMax-M3 | 27 | 1.08 |
| GLM-5.3 | 6 | 0.24 |
| DS V4.1 Flash | 0 | 0.00 |
| DS V4 Pro | 0 | 0.00 |