A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification
Authors: Aleesha Zainab, Asifullah Khan, Muhammad Ahmed Khalid, Faheem Ullah Khan
Organizations: Pattern Recognition Lab, Department of Computer and Information Sciences (DCIS), Pakistan Institute of Engineering and Applied Sciences (PIEAS), Nilore, Islamabad, Pakistan. · Deep Learning Lab, Center for Mathematical Sciences, Pakistan Institute of Engineering and Applied Sciences (PIEAS), Nilore, Islamabad, Pakistan. · PIEAS Artificial Intelligence Center (PAIC), Pakistan Institute of Engineering and Applied Sciences (PIEAS), Nilore, Islamabad, Pakistan.
Organizations in critical national infrastructure sectors must assess heterogeneous documents for sensitivity before routing or storage. Manual assessment is slow, inconsistent, and unscalable. Extending our prior leakage-controlled benchmark, BERT established the top single-encoder baseline (89.14% accuracy, 89.33% F1-score under 5-fold cross-validation on the Strategic 16K corpus). However, transformer baselines suffer from a structural limitation: fixed input length truncation discards evidence beyond the retained window-precisely where sensitive cables tend to be longest. We present Channel-Boosted MAS (CB-MAS) and instantiate it as IC-MAS (Iterative Consultation Multi-Agent System) to solve this without long-context computational costs. A Channel Critic Agent learns document-adaptive trust weights governing Gated Channel Boosting between two first-window encoders, while paired Consultation Agents iteratively exchange belief states to reconcile evidence from the beginning and end of long documents. IC-MAS holds computation constant regardless of document length by reconciling fixed windows in a compact representation space. Ablation studies show critic-controlled Channel Boosting provides the bulk of accuracy gains, while consultation recovers recall without precision collapse. Critic-Controlled Gated Channel Boosting with Max-Pool fusion and Blackboard Adaptive Consultation achieves 90.72% accuracy, 91.23% F1-score, 92.01% sensitive recall, and 90.46% sensitive precision, using about 54% less average computation than a fixed-round baseline. Gains over the single-encoder baseline are statistically significant (McNemar's test, p less than 0.000001; paired t-test). We include LIME/SHAP explainability, multi-agent evaluation, and an honest accounting of limitations.
Figures & tables
Study
Method
Dataset
Score
XAI
MAS
Alzhrani et al. [ 3 ]
CNN
PlusD
F1=0.91
No
No
Petrolini et al. [ 4 ]
BERT
Privacy
F1=0.95
No
No
McDonald [ 5 ]
Classif.
Gov. recs.
BA=0.70
No
No
Ahmad et al. [ 8 ]
SVM/RF/NB
Reuters
Acc=0.97
No
No
Hart et al. [ 14 ]
Text clf.
DLP
High
No
No
Zainab et al. [ 1 ]
BERT
Strategic 16K
F1=0.8933
–
–
Table 1: Comparison of Methodologies in Prior Sensitivity-Classification Studies
Dataset
Accuracy
F1
Raw PlusD (leaky)
99.0%
99.0%
Strategic 16K (cleaned)
86.26%
86.83%
Table 2: Impact of Label Leakage on TF-IDF + Logistic Regression Performance
Figure 1: The WikiLeaks PlusD interface showing the Cablegate collection of 251,287 diplomatic cables used in this study (left), and the seven original sensitivity classification categories in the PlusD database (right).
Model
Acc.
F1
S-Recall
S-Prec.
BERT
89.14%
89.33%
86.66%
92.21%
ELECTRA
88.57%
88.90%
87.24%
90.70%
RoBERTa
85.85%
86.51%
86.57%
86.45%
LR + TF-IDF
86.26%
86.83%
86.41%
87.25%
SVM
86.35%
86.95%
86.74%
87.16%
Naïve Bayes
83.91%
84.92%
86.40%
83.48%
Table 3: 5-Fold Cross-Validation Results on Strategic 16K
Figure 2: Illustration of the truncation problem: a fixed-length input window retains only a portion of the document, while content beyond the truncation boundary is never presented to the model.
Figure 3: High-level overview of the IC-MAS pipeline: document segmentation, asymmetric dual-window processing, iterative consultation, adaptive halting, and final decision, all coordinated through the shared Blackboard.
Figure 4: Overall IC-MAS architecture with the Domain Evidence and Decision & Verification Agent. Classification = Dual-Window Reader, Consultation and Consensus Agents.
Figure 5: Channel Critic Agent data flow: first-window text is encoded by a frozen DistilBERT, reduced to a CLS embedding, and passed through a trainable MLP to produce the two document-adaptive trust weights α and β .
Figure 6: Gated Channel Boosting: the BERT vector p1 and RoBERTa vector p2 cross-gate one another, modulated by the critic’s trust weights α and β , producing boosted vectors b1 and b2 .
Figure 7: Iterative Consultation and Adaptive Halting combined. Two agents exchange messages across up to three rounds (top); a shared confidence-check mechanism, reading the combined output after each round, decides whether to continue to the next round or halt and output the current prediction (bottom).
Model
Acc.
F1
S-Recall
S-Prec.
BERT (baseline)
89.14%
89.33%
86.66%
92.21%
RoBERTa
85.85%
86.51%
86.57%
86.45%
BERT+RoBERTa CB (Trans.)
89.86%
90.44%
91.52%
89.39%
CB + Dual-Window (OR)
88.97%
89.83%
92.93%
86.93%
CB + IC-MAS (Trans., fixed)
90.57%
91.08%
91.81%
90.36%
CB + IC-MAS (Trans., adapt.)
90.40%
90.95%
92.07%
89.86%
Table 4: Five-Fold Cross-Validation Results on Strategic 16K
Figure 8: ROC curve (left, AUC =0.958 ) and Precision–Recall curve (right, Sensitive class) for the core model, pooled across all 5 test folds. The marked point shows where the recall-targeted operating threshold ( t=0.295 , tuned to reach 92% Sensitive Recall) falls along each curve.
Figure 9: Confusion matrix for the core model at the recall-targeted decision threshold ( t=0.295 ), pooled across all 5 test folds ( n=16,000 ).
Figure 10: Per-metric mean and standard deviation across the 5 cross-validation folds for the core model, corresponding to the summary statistics in Table 4 .
Figure 11: Distribution of predicted confidence, separately for correctly and incorrectly classified documents, computed at the default decision threshold ( t=0.5 ).
Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructure, while rule-based tools often miss threats that depend on context. This study presents TorchSight, an open-source local system for security document classification built around a fine-tuned Qwen 3.5 27B model. The model was trained on 78,358 samples from 13 permissively licensed sources and GPT-4 synthetic data covering seven security categories and 51 subcategories. In the main evaluation on 1,000 documents, the model reached 95.0% category-level accuracy (95% confidence interval: 93.5-96.2). The tested commercial models scored 75.4-79.9% under the same prompting protocol. On a separate external set of 500 held-out samples, the model reached 93.8% accuracy, which suggests that performance extends beyond the main benchmark, although the margin depends on dataset composition and difficult boundary cases. The results show that a fine-tuned local model can support accurate security document classification while keeping document processing under local control.
Ivan Dobrovolskyi
MS in AI & ML Engineering, Independent Researcher, Sunnyvale, CA, USA
Multi-agent document assessment for retrieval-augmented generation is computationally expensive, driving practitioners toward smaller, deployable models whose assessment mechanisms remain poorly understood. We conduct a controlled study of training-free interventions on 7B-9B instruction-tuned models across diverse QA benchmarks, revealing a sharp dichotomy in how models benefit from assessment. For weaker baselines, the dominant mechanism is per-document isolation. Astoundingly, assessment-free isolation matches full multi-agent assessment, demonstrating that resolving multi-document context confusion, rather than scoring quality, drives outsized gains of up to 50 percentage points. Conversely, for strong baselines where scoring quality matters, we introduce Reasoning-Score Coupling, a label-free perturbation probe that classifies scoring behavior. Integrating these findings, we propose MADARA, a model-adaptive routing architecture. Crucially, MADARA's diagnostic thresholds derived from a single pilot model generalize zero-shot to four unseen model families, providing a robust, lightweight pipeline to eliminate computational overhead.
LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden threshold, rather than being determined by any single message. We introduce a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection (BCBI), which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence. To mitigate this threat, we propose LAtent Transition Test-time Evaluation (LATTE), a clean-only latent-transition defense that learns benign communication dynamics and quarantines anomalous agent updates before their responses propagate. Across several benchmarks, BCBI yields selective activation with little premature activation; without knowing the attack target or trigger, LATTE limits propagation with minimal disruption.
Jia-Hao Xiao, Lei Feng, Min-Ling Zhang
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University)