Recent advances in paraphrase detection reveal a fundamental trade-off: large language models achieve high accuracy but require high computation, while efficient Siamese-BERT variants offer practical scalability with reduced transparency in rationale generation. We present R-DEIM Net, a 76M-parameter dual-expert architecture exploring whether moderate-scale models can achieve competitive accuracy on paraphrase detection while enabling human-readable rationale generation. The architecture combines two specialized components: an Interaction Expert that captures token-level similarity patterns through multi-scale 2D convolutions and attention head allowing variable input length, and a Reasoning Expert that uses a Flan-T5-small decoder to generate rationales as auxiliary supervision. Rather than re-encoding generated text, we extract and pool decoder hidden states as complementary features for classification. On the Quora Question Pairs dataset, R-DEIM Net achieves 90.07% accuracy and 90.16% F1-score via 10-fold cross-validation. This represents competitive performance with strong transformer-based baselines (e.g., MFAE BERT: 90.54% accuracy) and recent large language model based approaches (LLaMA-70B) while using a substantially smaller parameter budget. The model generates rationales alongside predictions, providing potential for auxiliary human-readable descriptions.
Figures & tables
Figure 1: R-DEIM Net Architecture. A dual-expert multi-task framework. Branch A (Interaction) captures token-to-token patterns via multi-scale 2D-CNNs, while Branch B (Reasoning) distills latent decoder trajectories into a “reasoning vector” to ground classification in natural language rationales.
Model
Size
Train Acc.
Train F1
Test Acc.
Test F1
bert-base-uncased
( Devlin et al., 2018 )
110M
90.95±10.04
89.54±14.49
85.99±8.11
84.63±12.64
paraphrase-albert-base-v2
12M
88.29±6.26
87.87±6.74
85.00±3.99
84.49±4.43
paraphrase-MiniLM-L3-v2
17M
90.55±1.53
90.61±1.53
86.52±0.57
86.62±0.58
paraphrase-MiniLM-L6-v2
22M
91.16±1.69
91.25±1.67
87.05±0.71
87.20±0.66
paraphrase-MiniLM-L12-v2
33M
92.97±1.30
93.05±1.27
88.42±0.37
88.56±0.35
Table 1: Encoder benchmarking on QQP.
Model
Acc.
F1 -Score
ABCNN ( Yin et al., 2016 )
63.59%
-
Syn-tree ( Chen et al., 2017 )
75.50%
-
Siamese-CNN ( Wang et al., 2017 )
79.60%
-
MP-CNN ( Wang et al., 2017 )
81.38%
-
XGBoost ( Ansari and Sharma, 2020 )
82.44%
80.44%
Siamese-LSTM ( Wang et al., 2017 )
82.58%
-
Table 2: Comparative Performance. R-DEIM Net obtains strong accuracy and F1 on QQP using a compact architecture (76M params).
Zero-shot
Fine-tuning
Dataset
Acc.
F1
Epochs
Acc.
F1
SoDD ( Pašek et al., 2022 )
73.50
63.01
3
88.60
88.40
PIT ( Xu et al., 2015 )
68.97
70.17
3
81.26
81.04
PAWS ( Zhang et al., 2019 )
53.45
53.54
2
68.80
67.97
MRPC ( Dolan and Brockett, 2005 )
50.92
50.29
3
69.51
65.85
SciTail ( Khot et al., 2018 )
66.60
60.60
2
81.47
81.59
Table 3: Transfer learning results of R-DEIM Net.
Figure 2: Structural Interaction Distribution. ECDFs of diagonal interaction ratios. The significant rightward shift for duplicates validates that semantic identity is strongly correlated with positional interaction mass.
Interaction Regime
k5/k3
k7/k3
Low
1.14
1.15
Mid
1.11
1.09
High
1.09
1.04
Table 4: Relative activation of larger CNN kernels normalized by the 3×3 kernel.
Figure 3: Token-level perturbation analysis. Content words (Nouns, Verbs) exert the strongest influence on model decisions.
Figure 4: Group-level POS perturbation. The model is highly robust to the removal of syntactic function words but sensitive to the loss of semantic content.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train ( N )
Test ( N )
Ratio ( 0:1 )
SODD
804,576
94,514
14:5
SprintFAQ
80,800
20,200
100:1
PAWS
49,175
2,000
63:50
CQADupStack
37,924
9,482
1:1
SciTail
23,088
2,126
17:10
PIT (Twitter)
11,530
838
15:8
Appendix
Table 5: Dataset statistics and class ratios ( 0:1 ).
Parameter
Value
Shared Encoder
MiniLM-L12-v2 (384-dim)
SLM Decoder
Flan-T5-small (512-dim)
CNN Out Channels (per kernel)
128
CNN Kernel Sizes
[3, 5, 7]
ANN Hidden Size
512
ANN Dropout
0.3
Appendix
Table 6: Hyperparameters and configuration of R-DEIM Net.
Parallel reasoning, where a generator samples many candidate solutions and an aggregator selects the best, is one of the most effective forms of test-time scaling in large language models, and pairwise self-verification has become its strongest aggregation primitive. Yet pairwise verification carries a heavy cost: each judgment reads two complete solutions in full, and existing methods perform tens of such judgments per problem regardless of whether the comparison is informative. We introduce CAPS (Cascaded Adaptive Pairwise Selection), an inference-only framework that allocates verifier compute non-uniformly along two orthogonal axes: an evidence axis that adapts how much of each candidate the judge sees, and a distribution axis that adapts how comparisons are spread across the pool. CAPS instantiates these into a four-stage cascade with an optional rescue subroutine, and admits a closed-form verifier-token cost in which the per-candidate marginal cost is roughly halved relative to uniform full-evidence schedules. On four self-verifying models (Qwen3-14B, GPT-OSS-20B, Qwen3-4B-Instruct/Thinking) and five reasoning benchmarks spanning code (LiveCodeBench-v5/v6, CodeContests) and math (AIME 2025, HMMT 2025), CAPS outperforms the leading pairwise verifier on 14 of 20 suites while using 25.4% of its verifier-token budget on code, and outperforms pointwise self-verification on all 20. The trade-off suites admit an interpretable diagnostic in terms of the verifier's accuracy at partial versus full evidence, providing a concrete pre-deployment check for cascade suitability.
Fangzhou Lin, Shuo Xing, Peiran Li +6
Texas A&M University · Georgia Institute of Technology · Tohoku University +1
Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently sequential decoding incurs substantial latency, motivating parallelization of the generation process. However, existing parallel reasoning approaches suffer from performance degradation compared to their sequential counterparts, and often rely on specialized inference engines. We introduce ThreadWeaver, a framework for adaptive parallel reasoning that matches the accuracy of comparably sized sequential reasoning models while significantly reducing inference latency via three key innovations: 1) a two-stage parallel trajectory generator that produces high-quality parallel chain-of-thought data for supervised fine-tuning; 2) a trie-based rollout design that enables parallel reasoning on any off-the-shelf autoregressive inference engine; and 3) a parallelization-aware reinforcement learning framework that trains the model to balance reasoning accuracy with effective parallelization. Across six challenging math reasoning benchmarks, ThreadWeaver trained on top of Qwen3-8B achieves performance on par with cutting-edge sequential reasoning models (79.9% on AIME24 and 71.9% on average) while delivering up to 1.53x speedup in token latency, establishing a new Pareto frontier between accuracy and efficiency.
Long Lian, Sida Wang, Felix Juefei-Xu +7
Meta Superintelligence Labs (MSL), Menlo Park, CA, United States · UC Berkeley, Berkeley, CA, United States · UCSF, San Francisco, CA, United States
Parallel reasoning enhances Large Reasoning Models (LRMs) but incurs prohibitive costs due to futile paths caused by early errors. To mitigate this, path pruning at the prefix level is essential, yet existing research remains fragmented without a standardized framework. In this work, we propose the first systematic taxonomy of path pruning, categorizing methods by their signal source (internal vs. external) and learnability (learnable vs. non-learnable). This classification reveals the unexplored potential of learnable internal methods, motivating our proposal of STOP (Super TOken for Pruning). Extensive evaluations across LRMs ranging from 1.5B to 20B parameters demonstrate that STOP achieves superior effectiveness and efficiency compared to existing baselines. Furthermore, we rigorously validate the scalability of STOP under varying compute budgets - for instance, boosting GPT-OSS-20B accuracy on AIME25 from 84% to nearly 90% under fixed compute budgets. Finally, we distill our findings into formalized empirical guidelines to facilitate optimal real-world deployment. Code, data and models are available at https://bijiaxihh.github.io/STOP
Jiaxi Bi, Tongxu Luo, Wenyu Du +2
The Chinese University of Hong Kong, Shenzhen · Shenzhen Loop Area Institute · DualityRL