Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.
Figures & tables
System
Overall
95% CI
Theory
Case
Invalid
Kimi-K3
75.56
73.84–77.21
86.04
60.89
4.98
GLM-5.3
75.08
73.34–76.74
84.59
61.75
0.24
Qwen3.8-Max
74.84
73.10–76.50
84.59
61.18
5.98
DS V4.1 Flash
72.43
70.64–74.15
83.56
56.84
0.00
DS V4 Pro
72.35
70.56–74.07
84.66
55.11
0.00
MiniMax-M3
63.20
61.29–65.07
71.25
51.93
1.08
Table 1: Post-selection results on the final benchmark candidate (%). Accuracy includes invalid answers; intervals are nominal Wilson 95% descriptions.
Figure 3: Final-set Theory and Case accuracy. Every endpoint has a lower Case score; invalid-response sensitivity retains the gap.
Figure 4: Final-set category accuracy (%), shown to one decimal place. The outlined Mean column reports the six-model mean computed from unrounded accuracies. Categories are sorted by six-model mean; numbers in parentheses indicate the number of items evaluated per model. These are descriptive post-selection results.
System
Fact ( n=464 )
Theory ( n=574 )
Gap (pp)
Kimi-K3
55.60
65.16
+9.55
GLM-5.3
56.03
66.38
+10.34
Qwen3.8-Max
56.47
64.98
+8.52
DS V4.1 Flash
51.29
61.32
+10.03
DS V4 Pro
45.91
62.54
+16.64
MiniMax-M3
45.04
57.49
+12.45
Table 2: Accuracy by existing Case label (%). “Fact” and “theory” are dataset tags, not verified outcomes. Gap = theory minus fact, using unrounded values.
System
Subset
Native
Disabled
Δ
Holm p
Flash
Overall
76.97
75.53
+1.43
0.1795
Flash
Theory
84.00
77.47
+6.53
1.66×10−8
Flash
Case
69.93
73.60
-3.67
0.00935
Pro
Overall
76.83
70.10
+6.73
3.26×10−13
Pro
Theory
85.07
72.87
+12.20
4.92×10−24
Pro
Case
68.60
67.33
+1.27
0.3663
Table 3: DeepSeek native versus disabled provider-configuration contrast on the original 3,000 items. Delta is native minus disabled accuracy in percentage points; Holm correction covers the six displayed tests. Native Flash records used 32,768 and later 65,536 maximum output tokens; native Pro used 65,536 throughout.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Original
Final
Total items
3,000
2,492
Theory items
1,500
1,454
Case items
1,500
1,038
Theory categories
14
14
Case categories
11
11
Options per item
4
4
Appendix
Table 4: Dataset statistics for the original and final releases.
ID
Theory category
n
ID
Case category
n
T01
Five Elements
104
C01
Career
96
T02
Stems and Branches
104
C02
Education
96
T03
Ten Gods
104
C03
Wealth
96
T04
Ten Gods Relations
104
C04
Romance
96
T05
Combination and Clash
104
C05
Health
96
T06
Strength and Weakness
104
C06
Personality
96
Appendix
Table 5: Final taxonomy and category sizes.
Metric
Original
Final
Items
3,000
2,492
Mean empirical difficulty
0.2326
0.2776
Mean discrimination
0.2182
0.2574
Six-model all-correct
1,595
1,118
Six-model all-wrong
205
204
Holm-significant pairs
11/15
11/15
Appendix
Table 6: Observed changes under model-informed hardening. Final-set test counts are post-selection descriptions.
Item band
Removed
Six of six correct
477
Five of six correct
30
Two to four correct
0
One of six correct
0
Zero of six correct
1
Total
508
Appendix
Table 7: Actual deletion-band composition of the final hardening.
Model
Original
Final
Change
Kimi-K3
79.50
75.56
-3.94
GLM-5.3
79.13
75.08
-4.05
Qwen3.8-Max
78.97
74.84
-4.13
DS V4.1 Flash
76.97
72.43
-4.54
DS V4 Pro
76.83
72.35
-4.48
MiniMax-M3
69.07
63.20
-5.87
Appendix
Table 8: Observed endpoint-system accuracy before and after hardening. Original scores are recomputed from the outcome matrix used by hardening; final scores filter those responses by retained ID.
Run family
Models
Reasoning
Source
V1 native
Six listed models
Native
Mixed routes
V2 Flash
DS V4.1 Flash
Disabled
DeepSeek official
V2 Pro
DS V4 Pro
Disabled
DeepSeek official
Appendix
Table 9: Run configuration. “Native” means no thinking-disable parameter was sent.
Endpoint configuration
Observed route(s)
Mode/max
UTC span
Raw
Set
Canonical
SHA prefix
DS V4 Pro (native)
DeepSeek official (3004)
native; 65536
13:32–16:21
3004
final
2492
aa380761e51f
DS V4.1 Flash (native)
DeepSeek official (3002)
native; 32768/65536
09:07–10:02
3002
final
2492
2225eb6179df
MiniMax-M3 (native)
SCNet (1807); early/unknown (1340)
native; 32768/65536
07:45–13:24
3147
final
2492
981c54a982f4
Kimi-K3 (native)
SCNet (1816); micuapi (1424); early/unknown (919)
native; 32768/65536
07:46–17:00
4159
final
2492
c8b94ded91c9
Qwen3.8-Max (native)
SCNet (1636); micuapi (1589); early/unknown (883)
native; 32768/65536
07:45–17:35
4108
final
2492
248c1ba7d807
GLM-5.3 (native)
SCNet (1489); micuapi (1706); early/unknown (913)
native; 32768/65536
07:45–14:39
4108
final
2492
ce25699f48ab
Appendix
Table 10: Per-system run and route manifest. Raw counts include transport retries; canonical counts are the item records used by the analysis. Exact route segments, API model IDs, base-path classes, response-file paths, and complete hashes are retained in the recovered CSV manifest.
Figure 5: Reasoning-effort summaries: median and 90th percentile per item. The token panel includes only the two direct DeepSeek runs with reliable positive token reports. The character panel covers all six systems and treats absent reasoning content as zero; characters remain a route-dependent proxy.
System
Micro
Task macro
Category macro
Kimi-K3
75.56
73.46
73.69
GLM-5.3
75.08
73.17
73.38
Qwen3.8-Max
74.84
72.89
73.13
DS V4.1 Flash
72.43
70.20
70.41
DS V4 Pro
72.35
69.88
70.15
MiniMax-M3
63.20
61.59
61.82
Appendix
Table 11: Final-set micro and macro scores. Task macro is the equal-weight mean of Theory and Case accuracy; category macro is the equal-weight mean over 25 category accuracies.
Endpoint alias
Overall
95% CI
Theory
Case
Kimi-K3
79.50
78.02–80.91
86.40
72.60
GLM-5.3
79.13
77.64–80.55
85.00
73.27
Qwen3.8-Max
78.97
77.47–80.39
85.00
72.93
DS V4.1 Flash
76.97
75.43–78.44
84.00
69.93
DS V4 Pro
76.83
75.29–78.31
85.07
68.60
MiniMax-M3
69.07
67.39–70.70
72.07
66.07
Appendix
Table 12: Pre-selection reference results on the original 3,000 items. Intervals are nominal Wilson 95% intervals and do not account for source/chart clustering.
Model
Theory
Case
Gap
Kimi-K3
86.04
60.89
25.15
GLM-5.3
84.59
61.75
22.84
Qwen3.8-Max
84.59
61.18
23.42
DS V4.1 Flash
83.56
56.84
26.72
DS V4 Pro
84.66
55.11
29.56
MiniMax-M3
71.25
51.93
19.32
Appendix
Table 13: Theory and Case accuracy and primary gap (percentage points).
Model
Primary
Excl. invalid
Change
Kimi-K3
75.56
79.52
+3.96
GLM-5.3
75.08
75.26
+0.18
Qwen3.8-Max
74.84
79.60
+4.76
DS V4.1 Flash
72.43
72.43
0.00
DS V4 Pro
72.35
72.35
0.00
MiniMax-M3
63.20
63.89
+0.69
Appendix
Table 14: Sensitivity to excluding invalid responses from each model’s denominator.
Figure 6: Final-set micro accuracy with nominal 95% Wilson intervals. Values are percentages; intervals are descriptive after outcome-informed selection.
Figure 7: Final-set pairwise accuracy differences in percentage points (row minus column). Asterisks mark nominal exact McNemar p<.05 after Holm correction across 15 pairs; these post-selection markers are exploratory.
Task
Category
n
Difficulty
Theory
Shensha Basics
104
0.240
Theory
Pattern Basics
104
0.220
Theory
Strength and Weakness
104
0.210
Theory
Nayin
104
0.114
Theory
Twelve Stages
104
0.109
Case
Career
96
0.630
Appendix
Table 15: Most and least difficult categories by six-model empirical difficulty.
Figure 8: Theory-category accuracy across the six models. Vector labels show rounded percentages; full category names follow Table 5 .
Figure 9: Case-category accuracy across the six models. Vector labels show rounded percentages; the larger spread in several categories is consistent with the aggregate Theory-to-Case gap.
Figure 10: Final-set item counts by number of models matching the released key, pooled across Theory and Case (2,492 items).
Release
Task
n
Cue n (%)
Charts
Max reuse
Original
Theory
1,500
284 (18.93)
–
–
Original
Case
1,500
481 (32.07)
274
33
Original
Total
3,000
765 (25.50)
–
–
Final
Theory
1,454
278 (19.12)
–
–
Final
Case
1,038
286 (27.55)
249
25
Final
Total
2,492
564 (22.63)
–
–
Appendix
Table 16: Structural-cue and repeated-chart audit. The cue is present when the correct option is uniquely the longest. Chart statistics apply to Case items; “shared items” counts Case records whose chart occurs at least twice.
Figure 11: Native and disabled-configuration accuracy on the original 3,000 items. Connected dots show paired provider configurations; displayed deltas are native minus disabled.
Model
Subset
Native
Disabled
Δ
Flash
Overall
72.43
71.79
+0.64
Flash
Theory
83.56
77.03
+6.53
Flash
Case
56.84
64.45
-7.61
Pro
Overall
72.35
66.29
+6.06
Pro
Theory
84.66
72.49
+12.17
Pro
Case
55.11
57.61
-2.50
Appendix
Table 17: Selected-final-set native/disabled diagnostic. Because inclusion uses the native-arm outcomes, these deltas and nominal adjusted p -values are post-selection summaries, not confirmatory effects.
Alias
ΔC
Tok./ ΔC
pp/1k
Lat. med. N/D
Ratio
Flash
43
353,933
0.283
10.68/1.43
7.44 ×
Pro
202
71,180
1.405
20.95/2.07
10.13 ×
Appendix
Table 18: Token and request-latency efficiency for the paired DeepSeek configurations on the original 3,000 items. ΔC is the native-minus- disabled net key-match count; pp/1k is accuracy-point gain per 1,000 additional reasoning tokens per item. Latency is native/disabled seconds; the ratio divides the two medians.
Figure 12: Overall accuracy versus mean reasoning-content length over all final-set records, including zeros for absent content. Characters are used because reported reasoning tokens are not comparable across routes. This descriptive cross-system plot is confounded by route and task composition.
Model
Invalid n
Invalid (%)
Qwen3.8-Max
149
5.98
Kimi-K3
124
4.98
MiniMax-M3
27
1.08
GLM-5.3
6
0.24
DS V4.1 Flash
0
0.00
DS V4 Pro
0
0.00
Appendix
Table 19: Invalid responses in the six-model main evaluation.
Figure 13: Invalid response rates on the final set. Empty answers are counted as wrong in the primary accuracy metric.
Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framework that separates scorer-independent execution evidence, including termination, answer exposure, parseability, and completion length, from scorer-dependent correctness. Across 2,550 outputs from five fixed Qwen and DeepSeek configurations on MATH and ARC-Challenge, matched 2,048-token limits produce sharply different execution mixtures: 49 of 450 Qwen MATH outputs terminate without a final answer, compared with 5 of 300 DeepSeek MATH outputs and none of the 750 ARC outputs. Among the same 300 DeepSeek MATH question-model pairs, no missing-final length termination is observed at 8,192 tokens. A coverage-audited targeted verification study further shows that candidate-selection and aggregation policies can substantially alter comparative accuracy estimates. These results demonstrate that accuracy conflates execution case mix with verification policy. Evaluations of test-time methods should therefore report pre-intervention execution states, verification coverage, and scorer provenance alongside accuracy.
Zongyou Yang, Yinghan Hou
Dyson School of Design Engineering Imperial College London London, United Kingdom · Department of Electrical and Electronic Engineering Imperial College London London, United Kingdom
Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogic, an English--Chinese aligned benchmark that tests whether models preserve logical reasoning performance when the same latent logical structure is expressed in English and diverse Chinese surface realizations. Built from formal logical templates, the benchmark contains three data sets: (i) the General aligned set, derived from 60 General Propositions across nine template families; (ii) the Difficult aligned set, derived from 40 Difficult Problems; and (iii) the Chinese-only set, covering 15 language-specific phenomenon types. Each aligned item pairs one English reference expression with five Chinese realizations. Experiments on Qwen3, Ministral, and GLM models reveal a persistent English--Chinese performance gap. Back-translation from standard Chinese into English often improves performance on the General aligned set, but produces mixed effects on the Difficult aligned set, where Qwen3-32B and GLM-5.1 perform worse after translation. These results indicate that Chinese surface realization, translation artifacts, and model-specific behavior jointly affect multilingual logical reasoning. Overall, ChLogic provides a useful stress test for the robustness of multilingual reasoning.
Peixian Zhou, Yuxu Chen, Chaorui Zhang +3
College of Mathematics, Sichuan University, China · Theory Lab, 2012 Labs, Huawei Technologies Co., Ltd.
Commonsense reasoning is a key language model capability, as it is purportedly a prerequisite for many basic tasks, unlike specific factual knowledge. It is often measured with multiple-choice questions (MCQ) benchmarks, e.g. HellaSwag and PIQA. Some of these benchmarks, however, are outdated and contain numerous validity issues. We illustrate some typical validity issues with a case study on HellaSwag, one of the most popular and problematic benchmarks for commonsense reasoning. The issues we find range from basic ungrammaticality and numerous typos to misleading prompts or equally correct options. We show that if we remove question prompts or replace them with "Lorem ipsum dolor...", about 68% of model predictions do not change. We argue that this occurs due to inner flaws in the benchmark, not mere contamination that might be present in some models. Since benchmark scores are an essential part of model selection in both research and commercial applications, these issues can have severe consequences. Based on our findings, we propose BenCheck, a package for benchmark validity analysis that encapsulates the main checks performed in our case study and can be used to audit commonsense reasoning benchmarks. We apply these checks to PIQA, Global PIQA, and Winogrande.
Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar +4
1CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt · 2PleIAs, Paris, France